Serving LLMs on a single 8×H200 server
Benchmarking six open-weight models, including MiMo-V2.6, on one 8×H200 server.
Comparison of the performance of six open-weight models — GLM-5.3, GLM-5.3-Flash, DeepSeek-V4.1-Flash, Qwen3.8-Flash-Next and Xiaomi’s MiMo-V2.6 in both its Flash-RL and Pro-RL sizes — on a H200 node with 8 GPUs, measured on the same seven workload shapes I used for the B300 comparison.
They are not all the same weight class: MiMo-V2.6-Pro-RL is 1.02T parameters (42B active) and GLM-5.3 is 743B, against 309B for MiMo-V2.6-Flash-RL. What they have in common is that each one fits on this single node.
Method#
One node with 8× NVIDIA H200 (SM90, 141 GiB each). TP8 throughout except Qwen’s
SGLang cells, which are TP4. The flags used correspond to the ones from each model’s published
recipe in the vLLM and SGLang docs. There are only three forced deviations: --max-model-len 262144 for
GLM-5.3 under vLLM (the published command will not boot otherwise),
--max-num-seqs 512 for GLM-5.3-Flash under data parallelism, and a raised
engine-start timeout.
The two MiMo-V2.6 checkpoints store their expert weights in MXFP4, which this
hardware has no native path for: H200 is SM90, and FP4 tensor cores arrive with
Blackwell. Their published recipe therefore pins --moe-runner-backend marlin
on H200 where it uses deep_gemm on B300, so every MiMo number below is MXFP4
dequantised through Marlin rather than computed in FP4. That is a property of
running these models on Hopper, not something to tune away.
Both MiMo checkpoints also ship a DFlash speculative drafter inside the
repository under dflash/, which turns out to matter more than any other single
flag here — see Speculative decoding below.
GuideLLM 0.7.1, with seven workload shapes from 2,048 to 131,072 input tokens, runs bounded by duration rather than prompt count. Same workloads as the ones used for the B300 comparison.
The 256-stream runs are provisional due to some issues in the benchmarking procedure.
The exact commands#
Every configuration below, verbatim. Some boilerplate is common to all of them:
the HuggingFace cache is bind-mounted from /fsx because the root disk on this
node is too small to hold the weights, and --ulimit memlock=-1 is there
because several of these models pin host memory for their state tables. Flags
otherwise come straight from each model’s published hw=h200 recipe cell.
zai-org/GLM-5.3#
SGLang latency — sgl-glm53-lowlat
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:latest \
sglang serve --model-path zai-org/GLM-5.3 \
--tp 8 --speculative-algorithm EAGLE --speculative-num-steps 5 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 6 \
--mem-fraction-static 0.8 \
--host 0.0.0.0 --port 30000bashSGLang throughput — sgl-glm53-hithru
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:latest \
sglang serve --model-path zai-org/GLM-5.3 \
--tp 8 --dp 8 --enable-dp-attention --moe-a2a-backend deepep \
--speculative-algorithm EAGLE --speculative-num-steps 1 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 2 \
--mem-fraction-static 0.85 --chunked-prefill-size 32768 \
--max-running-requests 256 \
--host 0.0.0.0 --port 30000bashzai-org/GLM-5.3-Flash#
SGLang latency — sgl-flash-lowlat
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:glm-5.3-flash \
sglang serve --model-path zai-org/GLM-5.3-Flash \
--tp-size 8 --ep-size 8 --mem-fraction-static 0.75 \
--dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
--kv-cache-dtype bfloat16 --moe-runner-backend deep_gemm \
--speculative-algorithm EAGLE --speculative-num-steps 5 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 6 \
--reasoning-parser glm45 --tool-call-parser glm47 \
--host 0.0.0.0 --port 30000bashSGLang throughput — sgl-flash-hithru
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:glm-5.3-flash \
sglang serve --model-path zai-org/GLM-5.3-Flash \
--tp-size 8 --ep-size 8 --dsa-prefill-backend tilelang \
--dsa-decode-backend tilelang --kv-cache-dtype bfloat16 \
--moe-runner-backend deep_gemm --reasoning-parser glm45 \
--tool-call-parser glm47 \
--host 0.0.0.0 --port 30000bashzai-org/GLM-5.3#
vLLM latency — vllm-glm53
VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
--max-model-len 262144 --kv-cache-dtype fp8 --tensor-parallel-size 8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
--reasoning-parser glm47 --enable-auto-tool-choice \
--served-model-name glm-5.3 \
--host 0.0.0.0 --port 8000bashzai-org/GLM-5.3-Flash#
vLLM latency — vllm-flash
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:nightly \
--model zai-org/GLM-5.3-Flash \
--tensor-parallel-size 8 \
--speculative-config {"method":"mtp","num_speculative_tokens":5} \
--tool-call-parser glm47 --reasoning-parser glm47 \
--enable-auto-tool-choice --served-model-name glm-5.3-flash \
--host 0.0.0.0 --port 8000bashvLLM balanced — vllm-flash-tep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:nightly \
--model zai-org/GLM-5.3-Flash \
--tensor-parallel-size 8 --enable-expert-parallel \
--speculative-config {"method":"mtp","num_speculative_tokens":5} \
--tool-call-parser glm47 --reasoning-parser glm47 \
--enable-auto-tool-choice --served-model-name glm-5.3-flash \
--host 0.0.0.0 --port 8000bashvLLM throughput — vllm-flash-dep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:nightly \
--model zai-org/GLM-5.3-Flash \
--max-num-seqs 512 --data-parallel-size 8 --enable-expert-parallel \
--speculative-config {"method":"mtp","num_speculative_tokens":5} \
--tool-call-parser glm47 --reasoning-parser glm47 \
--enable-auto-tool-choice --served-model-name glm-5.3-flash \
--host 0.0.0.0 --port 8000bashzai-org/GLM-5.3#
vLLM balanced — vllm-glm53-tep
VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
--tensor-parallel-size 8 --enable-expert-parallel --max-model-len 262144 \
--kv-cache-dtype fp8 --speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
--reasoning-parser glm47 --enable-auto-tool-choice \
--served-model-name glm-5.3 \
--host 0.0.0.0 --port 8000bashvLLM throughput — vllm-glm53-dep
VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
--data-parallel-size 8 --enable-expert-parallel --max-model-len 262144 \
--kv-cache-dtype fp8 --speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
--reasoning-parser glm47 --enable-auto-tool-choice \
--served-model-name glm-5.3 \
--host 0.0.0.0 --port 8000bashdeepseek-ai/DeepSeek-V4.1-Flash#
SGLang latency — sgl-dsv41-lowlat
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:dev-dsv41 \
sglang serve --model-path deepseek-ai/DeepSeek-V4.1-Flash \
--trust-remote-code --tp 8 --ep-size 8 --mem-fraction-static 0.8 \
--attention-backend dsv4 --moe-runner-backend flashinfer_mxfp4 \
--cuda-graph-max-bs-decode 64 --reasoning-parser auto \
--tool-call-parser auto --enable-decoder-swa-bounded-replay \
--host 0.0.0.0 --port 30000bashSGLang throughput — sgl-dsv41-hithru
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:dev-dsv41 \
sglang serve --model-path deepseek-ai/DeepSeek-V4.1-Flash \
--trust-remote-code --tp 8 --ep-size 8 --mem-fraction-static 0.8 \
--attention-backend dsv4 --moe-runner-backend flashinfer_mxfp4 \
--cuda-graph-max-bs-decode 64 --reasoning-parser auto \
--tool-call-parser auto --max-running-requests 256 \
--host 0.0.0.0 --port 30000bashQwen/Qwen3.8-Flash-Next#
SGLang latency — sgl-qwen38-lowlat
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:qwen38flashnext \
sglang serve --model-path Qwen/Qwen3.8-Flash-Next \
--tp 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
--linear-attn-prefill-backend flashinfer \
--linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
--reasoning-parser auto --linear-attn-verify-backend triton \
--speculative-algorithm NEXTN --speculative-num-steps 3 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
--max-running-requests 96 \
--host 0.0.0.0 --port 30000bashSGLang throughput — sgl-qwen38-hithru
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:qwen38flashnext \
sglang serve --model-path Qwen/Qwen3.8-Flash-Next \
--tp 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
--linear-attn-prefill-backend flashinfer \
--linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
--reasoning-parser auto --ep 4 \
--host 0.0.0.0 --port 30000bashQwen/Qwen3.8-Flash-Next-FP8#
SGLang latency — sgl-qwen38fp8-lowlat
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:qwen38flashnext \
sglang serve --model-path Qwen/Qwen3.8-Flash-Next-FP8 \
--tp 4 --ep 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
--linear-attn-prefill-backend flashinfer \
--linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
--reasoning-parser auto --linear-attn-verify-backend triton \
--speculative-algorithm NEXTN --speculative-num-steps 3 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
--host 0.0.0.0 --port 30000bashSGLang throughput — sgl-qwen38fp8-hithru
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:qwen38flashnext \
sglang serve --model-path Qwen/Qwen3.8-Flash-Next-FP8 \
--tp 4 --ep 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
--linear-attn-prefill-backend flashinfer \
--linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
--reasoning-parser auto \
--host 0.0.0.0 --port 30000bashdeepseek-ai/DeepSeek-V4.1-Flash#
vLLM latency — vllm-dsv41-tp
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:nightly \
--model deepseek-ai/DeepSeek-V4.1-Flash \
--tensor-parallel-size 8 --language-model-only \
--tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
--enable-auto-tool-choice --reasoning-parser deepseek_v41 \
--gpu-memory-utilization 0.9 \
--speculative-config {"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic","rejection_sample_method":"block","enable_adaptive_verification":false} \
--max-model-len 262144 --max-num-seqs 128 --max-num-batched-tokens 16384 \
--served-model-name dsv41 \
--host 0.0.0.0 --port 8000bashvLLM throughput — vllm-dsv41-dep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:nightly \
--model deepseek-ai/DeepSeek-V4.1-Flash \
--data-parallel-size 8 --enable-expert-parallel --language-model-only \
--tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
--enable-auto-tool-choice --reasoning-parser deepseek_v41 \
--gpu-memory-utilization 0.9 \
--speculative-config {"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic","rejection_sample_method":"block","enable_adaptive_verification":false} \
--max-model-len 262144 --max-num-seqs 128 --max-num-batched-tokens 16384 \
--served-model-name dsv41 \
--host 0.0.0.0 --port 8000bashQwen/Qwen3.8-Flash-Next-FP8#
vLLM balanced — vllm-qwen38fp8-tep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:qwen38-flash-next \
--model Qwen/Qwen3.8-Flash-Next-FP8 \
--tensor-parallel-size 8 --enable-expert-parallel --moe-backend triton \
--gpu-memory-utilization 0.85 --max-num-seqs 256 --enable-prefix-caching \
--no-enable-flashinfer-autotune --enable-auto-tool-choice \
--tool-call-parser qwen3_coder --reasoning-parser qwen3 \
--served-model-name qwen38fp8 \
--host 0.0.0.0 --port 8000bashvLLM throughput — vllm-qwen38fp8-dep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
-e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm/vllm-openai:qwen38-flash-next \
--model Qwen/Qwen3.8-Flash-Next-FP8 \
--data-parallel-size 8 --enable-expert-parallel --moe-backend triton \
--gpu-memory-utilization 0.85 --max-num-seqs 256 --enable-prefix-caching \
--no-enable-flashinfer-autotune --enable-auto-tool-choice \
--tool-call-parser qwen3_coder --reasoning-parser qwen3 \
--served-model-name qwen38fp8 \
--host 0.0.0.0 --port 8000bashXiaomiMiMo/MiMo-V2.6-Flash-RL#
Stable vLLM cannot load the MXFP4-stored weights at all, so these use the image published for the series rather than a release tag.
vLLM latency — vllm-flash-tp (the published command, verbatim)
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
vllm/vllm-openai:mimo-v26 \
--model XiaomiMiMo/MiMo-V2.6-Flash-RL \
--tensor-parallel-size 4 --trust-remote-code --gpu-memory-utilization 0.95 \
--max-model-len auto --reasoning-parser mimo --tool-call-parser mimo \
--enable-auto-tool-choice --generation-config vllm \
--host 0.0.0.0 --port 8000bashvLLM latency + DFlash — vllm-flash-dflash-r3
$DFLASH is the dflash/ directory inside the downloaded snapshot. Note the
--gpu-memory-utilization 0.90: at the published 0.95 the drafter has nowhere
to allocate and the engine dies with a CUDA OOM.
DFLASH=$(ls -d /fsx/hf-cache/hub/models--XiaomiMiMo--MiMo-V2.6-Flash-RL/snapshots/*/dflash)
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
vllm/vllm-openai:mimo-v26 \
--model XiaomiMiMo/MiMo-V2.6-Flash-RL \
--tensor-parallel-size 4 --trust-remote-code --gpu-memory-utilization 0.90 \
--max-model-len auto \
--speculative-config '{"method":"dflash","model":"'"$DFLASH"'","num_speculative_tokens":7}' \
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
--reasoning-parser mimo --tool-call-parser mimo \
--enable-auto-tool-choice --generation-config vllm \
--host 0.0.0.0 --port 8000bashvLLM balanced — vllm-flash-tep
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
vllm/vllm-openai:mimo-v26 \
--model XiaomiMiMo/MiMo-V2.6-Flash-RL \
--tensor-parallel-size 8 --enable-expert-parallel \
--trust-remote-code --gpu-memory-utilization 0.95 --max-model-len auto \
--reasoning-parser mimo --tool-call-parser mimo \
--enable-auto-tool-choice --generation-config vllm \
--host 0.0.0.0 --port 8000bashSGLang — sgl-flash-cell (the published H200 cell, verbatim)
docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:dev \
sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Flash-RL \
--tp 4 --moe-runner-backend marlin --trust-remote-code \
--reasoning-parser mimo --tool-call-parser mimo \
--host 0.0.0.0 --port 30000bashXiaomiMiMo/MiMo-V2.6-Pro-RL#
The 1T checkpoint. Identical flags, TP8, and the Pro repository.
vLLM latency — vllm-pro-tp
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
vllm/vllm-openai:mimo-v26 \
--model XiaomiMiMo/MiMo-V2.6-Pro-RL \
--tensor-parallel-size 8 --trust-remote-code --gpu-memory-utilization 0.95 \
--max-model-len auto --reasoning-parser mimo --tool-call-parser mimo \
--enable-auto-tool-choice --generation-config vllm \
--host 0.0.0.0 --port 8000bashvLLM latency + DFlash — vllm-pro-dflash-r3
DFLASH=$(ls -d /fsx/hf-cache/hub/models--XiaomiMiMo--MiMo-V2.6-Pro-RL/snapshots/*/dflash)
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
vllm/vllm-openai:mimo-v26 \
--model XiaomiMiMo/MiMo-V2.6-Pro-RL \
--tensor-parallel-size 8 --trust-remote-code --gpu-memory-utilization 0.90 \
--max-model-len auto \
--speculative-config '{"method":"dflash","model":"'"$DFLASH"'","num_speculative_tokens":7}' \
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
--reasoning-parser mimo --tool-call-parser mimo \
--enable-auto-tool-choice --generation-config vllm \
--host 0.0.0.0 --port 8000bashSGLang — sgl-pro-cell
docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
-p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
lmsysorg/sglang:dev \
sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Pro-RL \
--tp 8 --moe-runner-backend marlin --trust-remote-code \
--reasoning-parser mimo --tool-call-parser mimo \
--host 0.0.0.0 --port 30000bashSingle request performance#
MiMo-V2.6-Flash-RL with DFlash is the fastest single stream in the set at 345 tok/s, and MiMo-V2.6-Pro-RL with DFlash is third at 273 — a 1T model sitting above every 300–700B configuration here, on the strength of one speculative-decoding flag. Without it the same two checkpoints fall to 232 and 140.
Behind them, DeepSeek-V4.1-Flash under vLLM reaches 294 tok/s at 3.3 ms between tokens — and the same model under SGLang’s published H200 cell is the slowest in the chart at 68 tok/s. That 4.3× gap is the one result here that is not a tuning trade-off: it is the same checkpoint on the same GPUs, and part of it is speculative decoding, which vLLM runs and the SGLang H200 cell does not. SGLang’s two published DeepSeek cells also return near-identical numbers at every point, so on this hardware they are effectively one configuration.
Otherwise the pattern is the one you would expect: low-latency recipes on top, high-throughput recipes at the bottom, and the largest model (GLM-5.3, 743B) holding its own at 195 tok/s because its EAGLE 5-step draft is nearly free at batch 1.
Concurrent request performance#
The order inverts almost completely: the low-latency recipes fall to the bottom and the throughput-oriented ones rise. Qwen3.8-Flash-Next FP8 under vLLM leads at 6,475 tok/s, and the model that was fastest single-stream — DeepSeek under vLLM’s latency strategy — is now mid-table.
Complete results#
Aggregate output tok/s at 64 concurrent requests — the operating point where a second repetition agreed within 6%, so these are the soundest numbers here.
| configuration | API-S | CHAT-S | CODE-I | CHAT-L | CODE-A | DOC-L | BATCH-D |
|---|---|---|---|---|---|---|---|
| GLM-5.3-Flash | |||||||
| SGLang latency | 759 | 1,675 | 1,608 | 964 | 957 | 264 | 4,093 |
| SGLang throughput | 1,529 | 2,840 | 1,873 | 1,248 | 868 | 222 | 4,369 |
| vLLM latency | 1,193 | 2,038 | 2,003 | 1,305 | 1,225 | 386 | 5,025 |
| vLLM balanced | 1,257 | 2,238 | 2,097 | 1,308 | 1,216 | 388 | 5,433 |
| vLLM throughput | 1,478 | 2,091 | 1,845 | 1,317 | 1,699 | 587 | 4,490 |
| GLM-5.3 | |||||||
| SGLang latency | 891 | 1,139 | 616 | 354 | 289 | 94 | 1,516 |
| SGLang throughput | 1,090 | 1,521 | 930 | 468 | 392 | rej | 4,009 |
| vLLM latency | 760 | 1,238 | 569 | 320 | 300 | 95 | 2,014 |
| vLLM balanced | 801 | 1,239 | 599 | 328 | 287 | 99 | 2,051 |
| vLLM throughput | — | — | — | — | — | — | — |
| DeepSeek-V4.1-Flash | |||||||
| SGLang latency | 437 | 655 | 624 | 624 | n/c | n/c | 2,184 |
| SGLang throughput | 437 | 655 | 624 | n/c | n/c | n/c | 2,184 |
| vLLM latency | 1,620 | 2,927 | 1,731 | 1,006 | 965 | n/c | 5,343 |
| vLLM throughput | 1,331 | 2,095 | 1,333 | 746 | 992 | rej | 5,700 |
| Qwen3.8-Flash-Next bf16 | |||||||
| SGLang latency | 1,032 | 1,685 | 1,936 | 1,189 | 1,155 | 353 | 4,446 |
| SGLang throughput | 1,857 | 2,403 | 1,873 | 1,100 | 1,100 | 359 | 4,327 |
| Qwen3.8-Flash-Next FP8 | |||||||
| SGLang latency | 1,170 | 1,681 | 1,917 | 1,198 | 1,390 | 342 | 4,453 |
| SGLang throughput | 2,075 | 2,840 | 2,496 | 1,249 | 1,039 | 289 | 4,276 |
| vLLM balanced | 2,435 | 3,398 | 3,102 | 1,838 | 2,082 | 378 | 6,475 |
| vLLM throughput | — | — | — | — | — | — | — |
| MiMo-V2.6-Flash-RL | |||||||
| SGLang | 2,075 | 2,838 | 1,986 | 1,043 | 825 | 266 | 4,358 |
| vLLM latency | 1,856 | 3,263 | 2,026 | 1,277 | 1,123 | 335 | 2,924 |
| vLLM latency + DFlash | 1,909 | 3,285 | 2,410 | 1,507 | 1,313 | 370 | 4,154 |
| vLLM balanced | 2,511 | 4,384 | 2,476 | 1,989 | 1,587 | 482 | 4,950 |
| vLLM throughput | 1,527 | 1,960 | 1,735 | 1,238 | 1,278 | 386 | 2,020 |
| MiMo-V2.6-Pro-RL | |||||||
| SGLang | 1,311 | 1,960 | 1,285 | 664 | 571 | 195 | 2,869 |
| vLLM latency | 1,295 | 2,393 | 1,438 | 909 | 739 | 231 | 2,359 |
| vLLM latency + DFlash | 1,335 | 2,679 | 1,648 | 987 | 892 | 255 | 3,198 |
| vLLM balanced | 1,205 | 2,392 | 1,317 | 861 | 712 | 224 | 2,096 |
| vLLM throughput | 763 | 1,091 | 663 | 525 | 514 | 193 | 1,103 |
rej — the server refused every request. n/c — no request finished inside the window.
Adding the MiMo models turns what used to be a clean sweep into a three-way split. Before they were measured, Qwen3.8-Flash-Next FP8 under vLLM won six of the seven benchmarks. It now wins three:
| benchmark | best configuration | output tok/s |
|---|---|---|
| API-S | MiMo-V2.6-Flash-RL, vLLM balanced | 2,511 |
| CHAT-S | MiMo-V2.6-Flash-RL, vLLM balanced | 4,384 |
| CODE-I | Qwen3.8-Flash-Next FP8, vLLM balanced | 3,102 |
| CHAT-L | MiMo-V2.6-Flash-RL, vLLM balanced | 1,989 |
| CODE-A | Qwen3.8-Flash-Next FP8, vLLM balanced | 2,082 |
| DOC-L | GLM-5.3-Flash, vLLM throughput | 587 |
| BATCH-D | Qwen3.8-Flash-Next FP8, vLLM balanced | 6,475 |
MiMo-V2.6-Flash-RL takes the conversational shapes and the 32k one; Qwen keeps the code shapes and the decode-heavy batch shape; GLM-5.3-Flash still owns the 128k DOC-L column it led before. Three different models, and for five of the seven the same strategy — vLLM with expert parallelism.
No configuration is good at everything. The best API-S cell is mid-table on DOC-L; the best DOC-L cell is mid-table on API-S. If your traffic is one shape, benchmark that shape.
Results per benchmark#
The same 64-stream measurement, one panel per workload shape, each model shown with whichever of its configurations was fastest on that shape. The point of splitting it out is that the ordering genuinely changes between panels — no model wins everywhere, and two of them win nothing.
Two things are worth pulling out. DeepSeek-V4.1-Flash is missing from the DOC-L panel — not because it was slow, but because no request finished inside the window on any of its configurations, so there is no rate to plot. And MiMo-V2.6-Pro-RL, the 1T model, places mid-table or last on every shape except the single-stream chart further up: parameter count buys quality, not tokens per second.
Speculative decoding: DFlash#
The largest single improvement in this whole comparison is not a parallelism strategy or an engine choice. It is one flag.
Both MiMo-V2.6 checkpoints ship a DFlash drafter inside the model repository.
Pointing --speculative-config at it, changing nothing else except dropping
--gpu-memory-utilization from 0.95 to 0.90 to leave the drafter room:
| output tok/s | MiMo-Flash | MiMo-Pro | ||
|---|---|---|---|---|
| baseline | +DFlash | baseline | +DFlash | |
| CHAT-S, 1 stream | 232 | 345 (+48%) | 140 | 273 (+95%) |
| CHAT-S, 64 streams | 3,263 | 3,285 (+1%) | 2,393 | 2,679 (+12%) |
| BATCH-D, 64 streams | 2,924 | 4,154 (+42%) | 2,359 | 3,198 (+36%) |
| DOC-L, 64 streams | 335 | 370 (+10%) | 231 | 255 (+10%) |
DFlash nearly doubles the 1T model’s single-stream throughput and improves every shape measured on both models. The shape of the gain is what you would expect from speculation: largest where the GPU is idle waiting on a sequential decode (one stream: +48% and +95%), smallest where 64 concurrent requests already keep it busy (+1% and +12%). It does not disappear under load, though — BATCH-D, which is decode-heavy at 8,192 output tokens, still gains 36–42%.
If you serve either of these models, this is the first thing to turn on.
What would not run on H200#
Two configurations from the published recipes cannot run on this hardware, and both trace back to MXFP4 having no native SM90 path:
--moe-a2a-backend deepeptogether with Marlin. SGLang rejects the combination outright:Runner backend MoeRunnerBackend.MARLIN requires a fused func for a2a backend deepep, but none is registered. B300 avoids this because it runsdeep_gemm; H200 is forced onto Marlin, which has no DeepEP kernel.--attention-backend fa4. Fails during startup withValueError: Expected size in shape to be strictly positive, but got 0, raised from CUTLASS’s layout builder. Isolated by changing one variable at a time: removing expert parallelism and keeping fa4 still fails; removing fa4 and keeping expert parallelism serves normally. This is a narrow claim — this model, this image, this node — not a general statement about FA4 on Hopper.
The MiMo cookbook page’s prose recommends keeping both of those on H200. Its machine-readable H200 cell omits both. The cell is right and the prose is wrong, which is a good argument for reading the config rather than the page. And the reduced version of the prose configuration that H200 can run still loses to the published cell: 167 against 201 tok/s single-stream, 2,151 against 2,838 at 64 streams.
Throughput vs concurrency#
Every model shows the same shape, and none of them is close to saturated at 16 streams. Qwen3.8-Flash-Next FP8 starts second-slowest of the seven at one stream and finishes first at 64; MiMo-V2.6-Pro-RL does the opposite, leading at c1 and ending last. Ranking a model on single-stream numbers tells you very little about how it will serve a loaded endpoint.
Within a model, the low-latency and throughput recipes cross somewhere — and where they cross is not portable either:
| model | crossover |
|---|---|
| DeepSeek-V4.1-Flash | below 16 concurrent requests |
| GLM-5.3 | between 16 and 64 |
| GLM-5.3-Flash | between 16 and 64 |
| Qwen3.8-Flash-Next (bf16 and FP8) | above 64 |
| MiMo-V2.6-Flash-RL | between 1 and 16 — but only against the balanced recipe |
| MiMo-V2.6-Pro-RL | never — the latency recipe leads at every concurrency measured |
A rule of thumb learned on one model picks the wrong recipe for another. If you serve Qwen at 32 concurrent requests, the low-latency recipe is still the faster choice; at the same load DeepSeek has already crossed over.
Qwen3.8-Flash-Next Quantization comparison#
Comparison of Qwen3.8-Flash-Next BF16 checkpoint (336 GB) and FP8 (173 GB) checkpoints:
| Qwen3.8-Flash-Next | API-S | CHAT-S | CODE-I | BATCH-D c64 |
|---|---|---|---|---|
| BF16, SGLang throughput | 1,857 | 2,403 | 1,873 | 4,327 |
| FP8, SGLang throughput | 2,075 | 2,840 | 2,496 | 4,276 |
On the decode-heavy shape the two are within 1.2% — half the weight memory for no measurable throughput change. On the shorter shapes FP8 is ahead by 12–33%, which is the more useful result if your traffic looks like API-S or CHAT-S.
(The 256-stream peaks are left out of this comparison on purpose: the only sound figure for either checkpoint at that concurrency is a re-measurement of the FP8 one, so a BF16-vs-FP8 comparison there would be mixing two different measurement windows.)
Conclusion#
Pick the tuning for the load you actually expect. Within a single model the low-latency and throughput recipes are different operating points, not better and worse versions of each other, and the crossover between them sits somewhere different for every model — so the choice cannot be carried over from one deployment to the next.
The engine matters less than the tuning for GLM-5.3-Flash, where SGLang and vLLM land within a few percent of each other at 64 concurrent requests. It matters a great deal elsewhere: vLLM is 17–51% ahead on Qwen3.8-Flash-Next FP8 and 2.5–4.5× ahead on DeepSeek-V4.1-Flash, on the same checkpoint and the same GPUs. Which way it goes is model-specific.
Look for a speculative drafter before tuning anything else. Enabling DFlash on MiMo-V2.6 beat every parallelism change tried on that model — +95% on single-stream Pro — and it is one flag pointing at a directory that is already inside the checkpoint. That is a better return than any strategy choice in this comparison.
Read the machine-readable recipe, not the page around it. Both MiMo configurations that refused to start on H200 were things the cookbook’s prose recommends and its own H200 cell omits, and both traced back to the same cause: MXFP4 has no native path on SM90, so Marlin is forced, and Marlin rules out the DeepEP kernel the prose asks for. The published cell had already accounted for the hardware.
Which is the practical point: none of this transfers. Benchmark your model, your engine and your workload shape before production, because every one of those three changed the answer here.