Javier Cacheiro

Back

Lately, I have been playing a lot with really large local models for different use cases where it is very important to keep full control of the data.

There are really good open-weight models right now that offer state-of-the-art performance, but to run them you need really powerful GPUs with a lot of VRAM.

One of the most versatile nodes for this task, as well as for training, is the B300 node: eight NVIDIA B300 GPUs, each one with 268 GiB of VRAM, which means 2.1 TB of VRAM in a single node.

This high amount of VRAM in just one node is one of the reasons why B300 nodes are in such high demand nowadays.

If you compare with an equivalent B200 node (the next GPU), it has 1.4 TB of VRAM, so roughly 700 GB less.

The great advantage of having 2 TB of VRAM is that with just one node you can run some of the frontier open-weight models right now without having to enter the wonderful world of multi-node inference and the fun of debugging the interconnection when you do not get the expected throughput (I can talk about that in a different post).

To measure how much one of these nodes can offer, I benchmarked five large models to see how much performance we can get for local inference.

Models tested#

Parameter counts complicate this comparison because several of the models that I benchmarked are mixture-of-experts, and published counts mix total and active parameters. So instead of the number from the model card, I use the measured VRAM consumption from what vLLM reported during loading, summed across the eight GPUs:

500 GiB 1,000 GiB 1,500 GiB Memory usage: Weight footprint Kimi-K3 MXFP4 · 1,532 GiB 1,532 GiB GLM-5.2 BF16 BF16 · 1,412 GiB 1,412 GiB Qwen3.8-2.4T NVFP4 · 1,340 GiB 1,340 GiB DeepSeek-V4-Pro FP8 · 850 GiB 850 GiB GLM-5.2 FP8 FP8 · 714 GiB 714 GiB
Memory usage for the model weights calculated from vLLM's log. It already accounts for quantization, and it is what determines how much space is left for the KV cache.

Workloads#

Using the great GuideLLM ↗ framework I created seven different workloads, representing different types of usage. They differ only in how many tokens go in and how many come out, but those numbers are very important because they affect performance a lot.

workloadin → outuse case description
CHAT-S2,048 → 512Ordinary interactive chat: short assistant turns. Latency-dominated.
CHAT-L32,768 → 2,048A long technical conversation with accumulated history and detailed answers.
CODE-I16,384 → 2,048Interactive coding assistant: explain this, fix this, review this.
CODE-A65,536 → 4,096A coding-agent call with a repository in context.
API-S2,048 → 256Short structured request: extraction, routing, JSON generation, metadata tagging. High request rate, no reasoning.
DOC-L131,072 → 2,048Long-document analysis and RAG: reports, logs, retrieved context. Prefill-dominated.
BATCH-D4,096 → 8,192Decode-heavy batch generation. Approximates maximum sustained output capacity.

BATCH-D runs with reasoning disabled and forces ignore_eos so the model cannot stop early, which makes it a maximum throughput probe. API-S also runs with reasoning disabled; the other five have reasoning enabled.

Results#

The metric reported here is server_output_tps, as reported by vLLM.

Let’s start with the raw ranking using the BATCH-D workload (the one where we can get the most throughput):

2,500 5,000 7,500 Peak batch throughput DeepSeek-V4-Pro FP8 · 850 GiB 9,183 GLM-5.2 FP8 FP8 · 714 GiB 7,204 Qwen3.8-2.4T NVFP4 · 1,340 GiB 6,168 Kimi-K3 MXFP4 · 1,532 GiB 4,422 GLM-5.2 BF16 BF16 · 1,412 GiB 4,360
Peak sustained output on BATCH-D, each model at the concurrency that maximised it. Headline metric is server_output_tps as reported by vLLM — what an operator sees on their own dashboard.

DeepSeek-V4-Pro leads at 9,183 tok/s, 27% ahead of GLM-5.2 FP8 even though it is 19% larger.

Kimi-K3 and GLM-5.2 BF16 are very close, but the great surprise is Qwen3.8-2.4T at 6,168 tok/s, which is of a similar size to Kimi-K3 and GLM-5.2 BF16 but performs much faster.

Comparing within size tiers#

Even though they are all state-of-the-art models, the difference in size is important, so it is fairer to compare them by establishing two tiers, mid and large, depending on the size:

tiermodelfootprintBATCH-DCHAT-S
midDeepSeek-V4-Pro850 GiB9,1835,493
GLM-5.2 FP8714 GiB7,2043,989
largeQwen3.8-2.4T1,340 GiB6,1684,115
Kimi-K31,532 GiB4,4222,880
GLM-5.2 BF161,412 GiB4,3603,658

Read this way:

  • In the mid tier, DeepSeek-V4-Pro beats GLM-5.2 FP8 by 27% on batch throughput while being 19% larger.
  • In the large tier, Qwen3.8-2.4T leads by ~40% over both Kimi-K3 and GLM-5.2 BF16, despite sitting between them in footprint.

Individual workload results#

Here are the individual results for the seven workloads, with the concurrency numbers (how many clients run at the same time) reported to the right of each bar:

2,000 4,000 API-S · 2,048 → 256 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 3,759 c512 GLM-5.2 FP8 FP8 · 714 GiB 2,684 c256 GLM-5.2 BF16 BF16 · 1,412 GiB 2,617 c256 Qwen3.8-2.4T NVFP4 · 1,340 GiB 2,443 c512 Kimi-K3 MXFP4 · 1,532 GiB 1,651 c256
API-S — 2,048 in → 256 out. Short structured requests: extraction, routing, JSON generation, metadata. Reasoning off. DeepSeek-V4-Pro and Qwen3.8-2.4T both peak at 512 concurrent requests, the other three at 256.
2,000 4,000 6,000 CHAT-S · 2,048 → 512 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 5,493 c512 Qwen3.8-2.4T NVFP4 · 1,340 GiB 4,115 c512 GLM-5.2 FP8 FP8 · 714 GiB 3,989 c256 GLM-5.2 BF16 BF16 · 1,412 GiB 3,658 c256 Kimi-K3 MXFP4 · 1,532 GiB 2,880 c512
CHAT-S — 2,048 in → 512 out. Ordinary interactive chat. The most common use case.
1,000 2,000 CODE-I · 16,384 → 2,048 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 2,335 c256 Qwen3.8-2.4T NVFP4 · 1,340 GiB 2,004 c256 GLM-5.2 FP8 FP8 · 714 GiB 1,814 c256 GLM-5.2 BF16 BF16 · 1,412 GiB 1,458 c192 Kimi-K3 MXFP4 · 1,532 GiB 1,390 c256
CODE-I — 16,384 in → 2,048 out. An interactive coding assistant: explain this function, make this edit, review this diff.
500 1,000 1,500 CHAT-L · 32,768 → 2,048 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 1,535 c256 GLM-5.2 FP8 FP8 · 714 GiB 1,071 c64 Qwen3.8-2.4T NVFP4 · 1,340 GiB 1,039 c256 GLM-5.2 BF16 BF16 · 1,412 GiB 855 c64 Kimi-K3 MXFP4 · 1,532 GiB 774 c256
CHAT-L — 32,768 in → 2,048 out. A long technical conversation carrying accumulated history.
500 1,000 CODE-A · 65,536 → 4,096 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 1,328 c256 Qwen3.8-2.4T NVFP4 · 1,340 GiB 997 c256 GLM-5.2 FP8 FP8 · 714 GiB 839 c64 GLM-5.2 BF16 BF16 · 1,412 GiB 719 c64 Kimi-K3 MXFP4 · 1,532 GiB 691 c64
CODE-A — 65,536 in → 4,096 out. A coding-agent model call with a repository in context. Real agents will issue many of these requests.
200 400 DOC-L · 131,072 → 2,048 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 384 c256 Qwen3.8-2.4T NVFP4 · 1,340 GiB 293 c256 GLM-5.2 BF16 BF16 · 1,412 GiB 260 c96 GLM-5.2 FP8 FP8 · 714 GiB 258 c16 Kimi-K3 MXFP4 · 1,532 GiB 212 c256
DOC-L — 131,072 in → 2,048 out. Long-document analysis and RAG.
5,000 10,000 BATCH-D · 4,096 → 8,192 · peak output tok/s DeepSeek-V4-Pro FP8 · 850 GiB 9,183 c512 GLM-5.2 FP8 FP8 · 714 GiB 7,204 c256 Qwen3.8-2.4T NVFP4 · 1,340 GiB 6,168 c512 Kimi-K3 MXFP4 · 1,532 GiB 4,422 c1024 GLM-5.2 BF16 BF16 · 1,412 GiB 4,360 c256
BATCH-D — 4,096 in → 8,192 out. Decode-heavy batch generation with EOS suppressed, so the model cannot stop early. It allows us to measure peak performance.

Latency#

In this benchmark we focus mainly on throughput, but there is another important aspect to measure: latency (how long you have to wait for the first token to appear).

So I measured time-to-first-token p95 at a fixed load of 256 concurrent requests. It spreads over a factor of 2.4 across the five models, against a factor of 1.7 in throughput at the same level of concurrency.

modelTTFT p95 @ c256tok/s @ c256
DeepSeek-V4-Pro9,852 ms4,237
GLM-5.2 BF1613,155 ms3,658
GLM-5.2 FP813,214 ms3,989
Qwen3.8-2.4T16,394 ms3,241
Kimi-K323,918 ms2,460

Kimi-K3 at 23.9 seconds to first token is by far the slowest, so you have to balance throughput against latency.

DeepSeek-V4-Flash#

Everything above compares five models between 714 GiB and 1,532 GiB. But just for measuring how much performance we can get on this server with a more modest model, I also measured DeepSeek-V4-Flash, which is much smaller: 170 GiB.

5,000 10,000 15,000 DeepSeek-V4-Flash · peak output tok/s by workload BATCH-D 4,096→8,192 15,484 c512 CHAT-S 2,048→512 9,010 c512 API-S 2,048→256 6,404 c1024 CODE-I 16,384→2,048 4,825 c256 CHAT-L 32,768→2,048 2,482 c256 CODE-A 65,536→4,096 2,471 c256 DOC-L 131,072→2,048 634 c64
DeepSeek-V4-Flash across all seven workloads, peak output tok/s with the concurrency that produced it. MXFP4 weights, speculative decoding disabled.

Conclusions#

With just one B300 node we can serve lots of concurrent requests, even for state-of-the-art models.

Serving LLMs on a single 8×B300 server
https://javicacheiro.com/blog/throughput-benchmarking-on-b300
Author Javier Cacheiro López
Published at September 10, 2026