Local LLMsQwenOpen-weight modelsInferenceDeveloper Tools

Run Qwen3.8-Flash-Next Locally With Strata: Hardware, Speed, Trade-offs

person
LaunchBoosts Research Desk·AI-assisted research
8 min read

Researched and drafted with AI assistance from the 7 public sources listed at the end of this article, then published after automated editorial checks. Spotted an error? Tell us at support@launchboosts.com.

Run Qwen3.8-Flash-Next Locally With Strata: Hardware, Speed, Trade-offs — LaunchBoosts

Yes, you can run Qwen3.8-Flash-Next on a gaming PC. The model has 125B parameters but only 6B are active per token. Strata is a new open-source inference engine that runs it with a 12 GB or larger NVIDIA or AMD GPU and 32–64 GB of system RAM. The "100 tokens per second on an RTX 4090" headline is roughly what people are reporting, with two catches: it uses aggressive 2–3-bit quantization, and RAM matters more than GPU size.

Strata reached the Hacker News front page on October 4–5, 2026. Below: what the model is, how Strata fits it onto consumer hardware, which build to choose, what the speed and quality numbers actually show, and the license terms to check before you build a product on it.

What Qwen3.8-Flash-Next is

Alibaba's Qwen team released Qwen3.8-Flash-Next on August 26, 2026, with full weights and an official FP8 variant on Hugging Face and ModelScope, according to CellCog's release summary. The team calls it an experimental preview of the Qwen4 architecture.

The official model card lists these specs:

  • Parameters: 125B total, made up of 6B active per token, a 51B n-gram embedding table and a 4B multi-token-prediction (MTP) module
  • Layers: 48, with a hybrid of Gated DeltaNet and Qwen Sparse Attention
  • Context: 262,144 tokens natively, extendable to 1M with YaRN scaling
  • Multimodal: includes a vision encoder
  • License: Qwen Community License 1.0

CellCog's summary adds that there are 512 experts per layer, with 10 routed experts and 1 shared expert active for each token. That is why Strata's README talks about "24,576 specialists": 48 layers × 512 experts.

The model card's selected benchmark scores include 62.5 on SWE-bench Pro (Qwen3.7-Plus: 55.8), 58.7 on DeepSWE 1.1, 73.9 on CoWorkBench and 84.5 on AndroidWorld. These are vendor-reported numbers. Treat them as a reason to test the model, not as proof.

The full-precision repository is about 360 GB, according to CellCog. That's why the official serving options (SGLang, vLLM and TokenSpeed) assume datacenter GPUs. Strata targets the hardware most developers actually own.

How Strata fits a 125B model on a 12 GB card

Because the model is a mixture of experts, only a small fraction of its weights are used for any given token. Strata's README spreads the model across the whole machine instead of trying to fit it into VRAM:

  1. The most-used experts go on the GPU. The router keeps choosing these experts, so serving them from VRAM is where the speed comes from.
  2. Every expert also stays in system RAM. Experts that aren't on the GPU are read from RAM when the router selects them.
  3. The SSD is used for the least-used parts in low-RAM setups. Strata warns this is much slower.
  4. Speculative decoding ("guess and check"). A smaller helper drafts the next tokens and the full model verifies them. The README says this returns results "1.6–1.8x sooner."

This design also explains why the 51B n-gram table isn't a problem. According to CellCog, Qwen designed that table to be offloaded to host memory with asynchronous prefetch, so it doesn't have to sit in VRAM.

The installer is a single script: START-HERE.bat on Windows or ./setup.sh on Linux. It detects your hardware, recommends a build, downloads roughly 70 GB of weights and serves a browser UI at http://127.0.0.1:8080. It also exposes OpenAI- and Anthropic-compatible APIs on localhost, so existing SDK code can point at it by changing the base URL. Strata is MIT-licensed. The model weights keep their own license.

Hardware requirements

Component Strata's stated requirement
GPU NVIDIA RTX 20/30/40/50 series or AMD RX 7800/7900/9000 series, 12 GB VRAM minimum
System RAM 32 GB minimum, 64 GB recommended
Storage About 80 GB free, preferably on an SSD
OS Windows 10/11 or Linux with current drivers

The most important line in the docs is this one, quoted by VRAM Calculator's breakdown: "A bigger graphics card makes it faster but does not lower the RAM it needs." If you have 32 GB of RAM, your options are very limited, whatever GPU you own.

Which build to pick

Strata ships several quantized builds. The table below uses VRAM Calculator's figures from Strata's documentation. Speeds were measured on an RTX 5070 (12 GB) with 64 GB of RAM.

Build Bits/weight Min. system RAM Output tok/s at 4K context at 128K context
Coder 1.89 32 GB 50.6 44.0
Q2_0 2.40 48 GB 90.3 67.2
IQ2_XS 2.50 48 GB 73.8 59.8
IQ3_XXS 3.00 64 GB (for 128K context) 62.1 45.8
IQ3_S 3.50 64 GB 51.6 40.5

On that machine, Q2_0 reads prompts at about 2,650 tokens per second. The README puts it more practically: a long first prompt takes roughly a minute per 30,000 tokens.

Some rules of thumb for choosing:

  • 32 GB RAM: the Coder build is your only realistic choice. Use it for coding agents and accept the lower speed.
  • 48 GB RAM: Q2_0 is the fastest general build, but its quality loss shows up mostly in code tasks (see below).
  • 64 GB RAM or more: IQ3_S is the build to judge quality against. Its self-reported quality roughly matches full precision.

What the speed and quality numbers actually show

Speed: the headline is plausible, but it depends on your setup

Reports from the Hacker News thread and the VRAM Calculator roundup vary widely:

  • RTX 4090, Ryzen 7950X3D, 128 GB DDR5: 124 tok/s
  • RTX 4090, 96 GB RAM: 70–80 tok/s at 128K context
  • RTX 5090, 64 GB RAM, IQ2_XS: about 200 tok/s (one commenter); another user reported 120–150 tok/s
  • RX 7900 XTX: 64.5 tok/s at 200K context
  • Two RTX 3060s, IQ3_S: about 40 tok/s, and read a 160K-token prompt in 123 seconds

One user ran a side-by-side test on two RTX 3090s: Strata's IQ3_XXS build reached about 100 tok/s, compared with 38 tok/s for llama.cpp running a 2-bit variant.

For comparison, owners of NVIDIA's DGX Spark on the NVIDIA developer forums report roughly 43–47 tok/s on coding work and up to 64 tok/s with NVIDIA's NVFP4 quant. On this model, a mid-range gaming GPU with plenty of RAM can beat a dedicated AI box on single-user generation speed.

Quality: check it yourself, especially for vision

Strata's own evaluation, summarized by VRAM Calculator, puts the Q2_0 build about 4 points below full precision on its task average, with most of that gap in code benchmarks. IQ3_S scores roughly the same as full precision. These are the project's own figures, not an independent evaluation.

The Hacker News thread raised a more specific concern. One commenter tested object-coordinate grounding with identical GGUF weights and measured a median error of 154.8 pixels in Strata versus 46.5 pixels in llama.cpp. They described that as falling from a 35B-class model to a 9B-class one. If your product depends on screenshots, UI agents or document layout, run this kind of test yourself before relying on Strata's vision path.

Several commenters also pointed out that Strata is a young, fast-moving project and hasn't been through the code review that llama.cpp or vLLM get. That doesn't make it wrong. It does mean you should pin versions and re-test after every update.

Known limitations

From the README and VRAM Calculator's notes:

  • The first model load takes 1–3 minutes, and you may need to close other apps to free RAM.
  • Requests are handled one at a time by default. This is a single-user engine, not a multi-tenant server.
  • AMD cards can't process images on Windows. Use Linux if you need vision on AMD.
  • Contexts beyond 262K are experimental. Recall testing has only been confirmed up to 320K.

Licensing: read it before you ship

Qwen3.8-Flash-Next is not Apache-2.0. It uses the Qwen Community License 1.0. According to CellCog's reading of the license:

  • Commercial use is allowed, with conditions.
  • Products above 100M monthly active users or $20M in monthly revenue must display the model name prominently.
  • "Model-as-a-Service or AI Work Assistant businesses" need a separate license from Qwen. Internal use is exempt.

The practical result: running the model locally for your own coding, research or internal tools is the simple case. Selling hosted access to it, or building an AI assistant product directly on it, is the case where you need to read the full license text, or talk to Qwen, before launch. This is a summary, not legal advice.

Local or hosted: a quick cost check

If you're running a local model to save on API bills, compare it with the hosted option. CellCog reports that the managed Qwen3.8-Flash on QwenCloud costs $0.16 per million input tokens and $0.47 per million output tokens. At those prices, a single developer's coding-agent traffic costs very little. The case for running locally is usually about something other than per-token cost:

  • Privacy and data residency. Code and documents never leave your machine.
  • No rate limits or outages. It works offline and on flights.
  • Predictable cost if you already own the GPU and the RAM.
  • Experimenting with agents that burn through millions of tokens on long contexts.

The case against: the RAM you'd have to buy (64–128 GB of DDR5 isn't cheap), single-request throughput, and quality loss at low bit-widths. Put your real monthly volume into the LLM API Cost Calculator and estimate prompt sizes with the Token Counter before you buy hardware.

If you have a Mac with a lot of unified memory, there's another route. Ollama's library lists a qwen3.8-flash-next:125b-mlx build at 105 GB with 256K context and image input.

What to do now

  1. Check your RAM first. With 64 GB or more and any 12 GB+ GPU from the last few generations, try IQ3_S. With 48 GB, try Q2_0. With 32 GB, try the Coder build.
  2. Point your existing code at localhost. Strata exposes OpenAI- and Anthropic-compatible APIs, so you can A/B test it against your current hosted model by changing the base URL.
  3. Run your own evals, not just the benchmarks. Test with 20–50 real tasks from your product. If you use vision, include grounding and screenshot tasks, which is where the Hacker News test found the biggest drop.
  4. Pin a version. The project is changing quickly, so re-test after updates.
  5. Watch what's coming next: whether llama.cpp and other mainstream engines adopt similar expert placement, independent quality evaluations of the sub-4-bit builds, and whether Qwen4 keeps this architecture. If it does, "frontier-ish MoE on a gaming PC" becomes normal rather than a one-off trick.

If you're building tools around local inference, you can list them for free on the LaunchBoosts AI tools directory. Plenty of developers are setting up local models this month.

Frequently asked questions

Can an RTX 4090 really run Qwen3.8-Flash-Next at 100 tokens per second?

Some people say so, but results vary. On Hacker News one commenter reported 124 tok/s on an RTX 4090 with 128 GB of DDR5 and a Ryzen 7950X3D. Another 4090 owner with 96 GB of RAM reported 70–80 tok/s at 128K context. Your speed depends on the quant, the context length and how much system RAM you have.

How much RAM do I need to run Qwen3.8-Flash-Next with Strata?

Strata's README sets 32 GB as the minimum and recommends 64 GB. The 32 GB figure only covers the smallest 'Coder' build. The general 2-bit builds need about 48 GB, and the 3-bit builds need 64 GB. A bigger GPU makes the model faster, but it doesn't reduce how much RAM you need.

Does a 2-bit or 3-bit quant of Qwen3.8-Flash-Next lose much quality?

According to the project's own figures, the 2-bit Q2_0 build scores about 4 points below full precision, mostly on code tasks, while IQ3_S roughly matches it. These numbers are self-reported. One Hacker News test also found much worse vision grounding in Strata than in llama.cpp using the same weights, so test the tasks you actually care about.

Can I use Qwen3.8-Flash-Next commercially?

It ships under the Qwen Community License 1.0, not Apache-2.0. According to CellCog's summary, commercial use is allowed with conditions, but businesses selling the model as a service or building AI work assistants need a separate license from Qwen. Read the license text before you ship a product on it.

What is the difference between Strata and Ollama or llama.cpp for this model?

Strata is built for this one model. It puts the most-used experts in VRAM, keeps the rest in RAM, and adds speculative decoding, which lets a small mid-range GPU work. Ollama lists a 105 GB MLX build aimed at large unified-memory Macs. llama.cpp is more mature and more widely reviewed, but in one user's report it was slower on the same hardware.

Sources

  1. Strata inference engine for Qwen3.8-Flash-Next (GitHub README)— github.com
  2. Qwen/Qwen3.8-Flash-Next model card (Hugging Face)— huggingface.co
  3. Qwen3.8-Flash-Next Is Out: Confirmed Specs, License, and the Leak Scorecard (CellCog)— cellcog.ai
  4. Hacker News discussion: Run Qwen 3.8 Flash Next (125B) on consumer hardware— news.ycombinator.com
  5. Strata: Running Qwen 3.8 Flash Next on a 12 GB Card (VRAM Calculator)— vramcalculator.com
  6. Qwen3.8-Flash-Next on DGX Spark / GB10 (NVIDIA Developer Forums)— forums.developer.nvidia.com
  7. qwen3.8-flash-next (Ollama library)— ollama.com
person

LaunchBoosts Research Desk

AI-assisted research

Explainers on software and AI industry trends, drafted with AI assistance from the public sources cited in each article and published after automated editorial checks for length, independent sourcing and originality.