artificial intelligence

How a 125B Model Reaches 100 Tok/s on an RTX 4090

How a 125B Model Reaches 100 Tok/s on an RTX 4090

Your GeForce RTX 4090 has 24 GB of VRAM, the fast memory attached to its graphics processing unit (GPU). Qwen3.8-Flash-Next looks impossible beside it: 125 billion learned parameters, plus a 51-billion-parameter n-gram embedding table and a multi-step prediction layer. The naive plan is to load everything into the graphics card. Strata takes a different route, spreading the work across the GPU, system memory, processor, and solid-state drive (SSD). (nvidia.com)

As of October 4, 2026, the clearest published RTX 4090 result is a community benchmark from September 30. On a 4090 paired with an Intel Core i9-13900K and 64 GB of DDR5 memory, Strata reported 106.1 tokens per second with the IQ2_XS build and 98.1 tokens per second with IQ3_XXS. Those figures counted the complete request, including the model’s internal thinking; short-chat peaks reached roughly 115 and 120 tokens per second. That is impressive, but it is a measured configuration rather than a promise every 4090 will deliver. (github.com)

The model is large on paper, sparse in practice

A token is a small piece of text, such as a word fragment, punctuation mark, or space. A parameter is a learned number that helps the model decide what those tokens mean and what should come next.

Qwen3.8-Flash-Next uses a mixture-of-experts architecture, usually shortened to MoE. Instead of one giant network doing every calculation, an MoE model contains many specialist sub-networks called experts. A router selects the experts needed for each token. This model has 512 experts, but only 10 routed experts plus one shared expert are active for a token, working out to about 6 billion activated parameters. (huggingface.co)

That distinction explains the first part of the speed story. The model has enormous total capacity, but it does not perform the full 125-billion-parameter calculation for every token. The catch is that all those experts still need to be available somewhere, because the router may ask for a different specialist on the next token. Sparse computation reduces arithmetic; it does not make the weights vanish.

Qwen also adds a large n-gram embedding table. An n-gram is a short sequence of tokens, such as a pair or triple, stored as a lookup rather than processed through the whole network every time. These embeddings add capacity with relatively little per-token computation and are well suited to being kept outside the GPU.

The desktop becomes a memory system

Strata treats the whole PC as one coordinated machine. The GPU keeps the model’s shared components, the draft layer, the attention cache, and the hottest experts in VRAM. System RAM, the larger memory attached to the CPU, holds the full expert pool. The CPU computes experts that are not currently cached on the GPU, while the SSD supplies large tables and model data when needed.

The expert cache behaves like a working set. Frequently used experts stay on the GPU; less-used ones remain in RAM. When the conversation changes, the cache adapts. Data that must cross between CPU memory and the GPU travels over PCI Express, commonly called PCIe, so a cache hit saves both transfer time and CPU work. This is why extra VRAM can matter more than a small increase in raw GPU speed for this particular workload.

The 4090 is therefore not holding a 125B model in isolation. Its 24 GB is holding the most valuable working set while the rest of the computer fills in the gaps. A fast NVMe SSD and a capable CPU are part of the accelerator, even though they are not printed on the graphics card’s specification sheet.

Three tricks create the speed

The first is quantization. Quantization stores model weights with fewer bits, reducing the amount of memory and bandwidth they need. Strata offers several packed versions of Qwen3.8-Flash-Next. IQ2_XS uses less memory and is a strong daily-driver choice, while IQ3_XXS uses more memory and CPU work in exchange for higher quality.

Quantized build Reported RTX 4090 result Best fit
IQ2_XS 106.1 tok/s whole request; about 115 tok/s peak Faster everyday chats
IQ3_XXS 98.1 tok/s whole request; about 120 tok/s peak More quality and longer reasoning

The second trick is speculative decoding. Strata uses the model’s multi-token prediction, or MTP, layer as a small in-model draft system. It guesses several upcoming tokens, then the full model checks those guesses together. When the guesses are accepted, one expensive verification pass produces multiple tokens instead of calculating each token in isolation. The large model remains the authority; the draft layer is there to reduce repeated overhead.

The third trick is memory management for long context. A context window is the amount of text the model can keep available during a conversation. The key-value cache, usually called the KV cache, stores intermediate attention information about earlier tokens. Strata can keep the cache in 8-bit form and stream less frequently used portions from RAM, leaving more VRAM available for experts. The 4090 benchmark used a 128K context setting, 8-bit KV storage, MTP, and KV streaming.

A sensible 4090 setup

The installation is intentionally less intimidating than the architecture:

# Linux
./setup.sh

# Windows
# Double-click START-HERE.bat

The setup program asks which model, quantized size, context length, and image support to use. For a 4090, IQ2_XS is the practical speed choice; IQ3_XXS is the quality choice if the system has enough RAM. A 128K context matches the published benchmark configuration, while 64 GB of system RAM gives the engine room to keep the experts beside the operating system and other applications.

Plan for roughly 70–80 GB of model storage, additional space for the MTP layer, and more room if you install image support or another quantized build. An NVMe SSD is strongly preferred. Strata’s current notes list 32 GB as a possible lower floor, but 64 GB is far more comfortable for the larger quality packs. Keep the NVIDIA driver current and avoid filling the GPU with games, browser acceleration, or other AI programs while testing. (github.com)

What 100 tokens per second really means

The headline measures answer generation, not every part of an AI request. Reading a large prompt, called prefill, is a separate stage from generating new tokens, called decoding. Long prompts, cold caches, image input, internal reasoning, CPU speed, RAM bandwidth, and the number of expert requests that miss the GPU cache can all change the result.

A token is also not a complete word. Strata’s documentation estimates roughly three-quarters of a word per token, so 100 tokens per second describes an extremely fast stream of text rather than 100 finished words every second. In a short conversation, the reply can appear faster than most people can comfortably read. During the first launch, however, the machine may spend minutes loading tens of gigabytes into memory before the fast generation begins.

The practical lesson is that the RTX 4090 does not perform a miracle by itself. Qwen3.8-Flash-Next combines sparse expert routing with aggressive quantization, while Strata adds a tiered memory system and verified drafting. Put those pieces together, and a model that normally suggests a server rack can produce a remarkably quick local response on a well-equipped gaming PC.

ahsan

ahsan

Hello! I am Mr Ahsan, the writer of the Website. I am from Netherland. I like to write about technology and the news around it.

Comments (0)

No comments yet. Be the first to respond!

Leave a Comment

Your comment will be visible after review.