Javier Cacheiro

Back

Comparison of the performance of six open-weight models — GLM-5.3, GLM-5.3-Flash, DeepSeek-V4.1-Flash, Qwen3.8-Flash-Next and Xiaomi’s MiMo-V2.6 in both its Flash-RL and Pro-RL sizes — on a H200 node with 8 GPUs, measured on the same seven workload shapes I used for the B300 comparison.

They are not all the same weight class: MiMo-V2.6-Pro-RL is 1.02T parameters (42B active) and GLM-5.3 is 743B, against 309B for MiMo-V2.6-Flash-RL. What they have in common is that each one fits on this single node.

Method#

One node with 8× NVIDIA H200 (SM90, 141 GiB each). TP8 throughout except Qwen’s SGLang cells, which are TP4. The flags used correspond to the ones from each model’s published recipe in the vLLM and SGLang docs. There are only three forced deviations: --max-model-len 262144 for GLM-5.3 under vLLM (the published command will not boot otherwise), --max-num-seqs 512 for GLM-5.3-Flash under data parallelism, and a raised engine-start timeout.

The two MiMo-V2.6 checkpoints store their expert weights in MXFP4, which this hardware has no native path for: H200 is SM90, and FP4 tensor cores arrive with Blackwell. Their published recipe therefore pins --moe-runner-backend marlin on H200 where it uses deep_gemm on B300, so every MiMo number below is MXFP4 dequantised through Marlin rather than computed in FP4. That is a property of running these models on Hopper, not something to tune away.

Both MiMo checkpoints also ship a DFlash speculative drafter inside the repository under dflash/, which turns out to matter more than any other single flag here — see Speculative decoding below.

GuideLLM 0.7.1, with seven workload shapes from 2,048 to 131,072 input tokens, runs bounded by duration rather than prompt count. Same workloads as the ones used for the B300 comparison.

The 256-stream runs are provisional due to some issues in the benchmarking procedure.

The exact commands#

Every configuration below, verbatim. Some boilerplate is common to all of them: the HuggingFace cache is bind-mounted from /fsx because the root disk on this node is too small to hold the weights, and --ulimit memlock=-1 is there because several of these models pin host memory for their state tables. Flags otherwise come straight from each model’s published hw=h200 recipe cell.

zai-org/GLM-5.3#

SGLang latency — sgl-glm53-lowlat

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:latest \
  sglang serve --model-path zai-org/GLM-5.3 \
  --tp 8 --speculative-algorithm EAGLE --speculative-num-steps 5 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 6 \
  --mem-fraction-static 0.8 \
  --host 0.0.0.0 --port 30000
bash

SGLang throughput — sgl-glm53-hithru

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:latest \
  sglang serve --model-path zai-org/GLM-5.3 \
  --tp 8 --dp 8 --enable-dp-attention --moe-a2a-backend deepep \
  --speculative-algorithm EAGLE --speculative-num-steps 1 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 2 \
  --mem-fraction-static 0.85 --chunked-prefill-size 32768 \
  --max-running-requests 256 \
  --host 0.0.0.0 --port 30000
bash

zai-org/GLM-5.3-Flash#

SGLang latency — sgl-flash-lowlat

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:glm-5.3-flash \
  sglang serve --model-path zai-org/GLM-5.3-Flash \
  --tp-size 8 --ep-size 8 --mem-fraction-static 0.75 \
  --dsa-prefill-backend tilelang --dsa-decode-backend tilelang \
  --kv-cache-dtype bfloat16 --moe-runner-backend deep_gemm \
  --speculative-algorithm EAGLE --speculative-num-steps 5 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 6 \
  --reasoning-parser glm45 --tool-call-parser glm47 \
  --host 0.0.0.0 --port 30000
bash

SGLang throughput — sgl-flash-hithru

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:glm-5.3-flash \
  sglang serve --model-path zai-org/GLM-5.3-Flash \
  --tp-size 8 --ep-size 8 --dsa-prefill-backend tilelang \
  --dsa-decode-backend tilelang --kv-cache-dtype bfloat16 \
  --moe-runner-backend deep_gemm --reasoning-parser glm45 \
  --tool-call-parser glm47 \
  --host 0.0.0.0 --port 30000
bash

zai-org/GLM-5.3#

vLLM latency — vllm-glm53

VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
  --max-model-len 262144 --kv-cache-dtype fp8 --tensor-parallel-size 8 \
  --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
  --reasoning-parser glm47 --enable-auto-tool-choice \
  --served-model-name glm-5.3 \
  --host 0.0.0.0 --port 8000
bash

zai-org/GLM-5.3-Flash#

vLLM latency — vllm-flash

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:nightly \
  --model zai-org/GLM-5.3-Flash \
  --tensor-parallel-size 8 \
  --speculative-config {"method":"mtp","num_speculative_tokens":5} \
  --tool-call-parser glm47 --reasoning-parser glm47 \
  --enable-auto-tool-choice --served-model-name glm-5.3-flash \
  --host 0.0.0.0 --port 8000
bash

vLLM balanced — vllm-flash-tep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:nightly \
  --model zai-org/GLM-5.3-Flash \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --speculative-config {"method":"mtp","num_speculative_tokens":5} \
  --tool-call-parser glm47 --reasoning-parser glm47 \
  --enable-auto-tool-choice --served-model-name glm-5.3-flash \
  --host 0.0.0.0 --port 8000
bash

vLLM throughput — vllm-flash-dep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:nightly \
  --model zai-org/GLM-5.3-Flash \
  --max-num-seqs 512 --data-parallel-size 8 --enable-expert-parallel \
  --speculative-config {"method":"mtp","num_speculative_tokens":5} \
  --tool-call-parser glm47 --reasoning-parser glm47 \
  --enable-auto-tool-choice --served-model-name glm-5.3-flash \
  --host 0.0.0.0 --port 8000
bash

zai-org/GLM-5.3#

vLLM balanced — vllm-glm53-tep

VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
  --tensor-parallel-size 8 --enable-expert-parallel --max-model-len 262144 \
  --kv-cache-dtype fp8 --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
  --reasoning-parser glm47 --enable-auto-tool-choice \
  --served-model-name glm-5.3 \
  --host 0.0.0.0 --port 8000
bash

vLLM throughput — vllm-glm53-dep

VLLM_ENGINE_READY_TIMEOUT_S=3600 \
vllm serve zai-org/GLM-5.3 \
  --data-parallel-size 8 --enable-expert-parallel --max-model-len 262144 \
  --kv-cache-dtype fp8 --speculative-config.method mtp \
  --speculative-config.num_speculative_tokens 5 --tool-call-parser glm47 \
  --reasoning-parser glm47 --enable-auto-tool-choice \
  --served-model-name glm-5.3 \
  --host 0.0.0.0 --port 8000
bash

deepseek-ai/DeepSeek-V4.1-Flash#

SGLang latency — sgl-dsv41-lowlat

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:dev-dsv41 \
  sglang serve --model-path deepseek-ai/DeepSeek-V4.1-Flash \
  --trust-remote-code --tp 8 --ep-size 8 --mem-fraction-static 0.8 \
  --attention-backend dsv4 --moe-runner-backend flashinfer_mxfp4 \
  --cuda-graph-max-bs-decode 64 --reasoning-parser auto \
  --tool-call-parser auto --enable-decoder-swa-bounded-replay \
  --host 0.0.0.0 --port 30000
bash

SGLang throughput — sgl-dsv41-hithru

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:dev-dsv41 \
  sglang serve --model-path deepseek-ai/DeepSeek-V4.1-Flash \
  --trust-remote-code --tp 8 --ep-size 8 --mem-fraction-static 0.8 \
  --attention-backend dsv4 --moe-runner-backend flashinfer_mxfp4 \
  --cuda-graph-max-bs-decode 64 --reasoning-parser auto \
  --tool-call-parser auto --max-running-requests 256 \
  --host 0.0.0.0 --port 30000
bash

Qwen/Qwen3.8-Flash-Next#

SGLang latency — sgl-qwen38-lowlat

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:qwen38flashnext \
  sglang serve --model-path Qwen/Qwen3.8-Flash-Next \
  --tp 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
  --linear-attn-prefill-backend flashinfer \
  --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
  --reasoning-parser auto --linear-attn-verify-backend triton \
  --speculative-algorithm NEXTN --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
  --max-running-requests 96 \
  --host 0.0.0.0 --port 30000
bash

SGLang throughput — sgl-qwen38-hithru

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:qwen38flashnext \
  sglang serve --model-path Qwen/Qwen3.8-Flash-Next \
  --tp 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
  --linear-attn-prefill-backend flashinfer \
  --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
  --reasoning-parser auto --ep 4 \
  --host 0.0.0.0 --port 30000
bash

Qwen/Qwen3.8-Flash-Next-FP8#

SGLang latency — sgl-qwen38fp8-lowlat

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:qwen38flashnext \
  sglang serve --model-path Qwen/Qwen3.8-Flash-Next-FP8 \
  --tp 4 --ep 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
  --linear-attn-prefill-backend flashinfer \
  --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
  --reasoning-parser auto --linear-attn-verify-backend triton \
  --speculative-algorithm NEXTN --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
  --host 0.0.0.0 --port 30000
bash

SGLang throughput — sgl-qwen38fp8-hithru

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:qwen38flashnext \
  sglang serve --model-path Qwen/Qwen3.8-Flash-Next-FP8 \
  --tp 4 --ep 4 --mem-fraction-static 0.85 --chunked-prefill-size 8192 \
  --linear-attn-prefill-backend flashinfer \
  --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 \
  --reasoning-parser auto \
  --host 0.0.0.0 --port 30000
bash

deepseek-ai/DeepSeek-V4.1-Flash#

vLLM latency — vllm-dsv41-tp

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:nightly \
  --model deepseek-ai/DeepSeek-V4.1-Flash \
  --tensor-parallel-size 8 --language-model-only \
  --tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
  --enable-auto-tool-choice --reasoning-parser deepseek_v41 \
  --gpu-memory-utilization 0.9 \
  --speculative-config {"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic","rejection_sample_method":"block","enable_adaptive_verification":false} \
  --max-model-len 262144 --max-num-seqs 128 --max-num-batched-tokens 16384 \
  --served-model-name dsv41 \
  --host 0.0.0.0 --port 8000
bash

vLLM throughput — vllm-dsv41-dep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:nightly \
  --model deepseek-ai/DeepSeek-V4.1-Flash \
  --data-parallel-size 8 --enable-expert-parallel --language-model-only \
  --tokenizer-mode deepseek_v41 --tool-call-parser deepseek_v41 \
  --enable-auto-tool-choice --reasoning-parser deepseek_v41 \
  --gpu-memory-utilization 0.9 \
  --speculative-config {"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic","rejection_sample_method":"block","enable_adaptive_verification":false} \
  --max-model-len 262144 --max-num-seqs 128 --max-num-batched-tokens 16384 \
  --served-model-name dsv41 \
  --host 0.0.0.0 --port 8000
bash

Qwen/Qwen3.8-Flash-Next-FP8#

vLLM balanced — vllm-qwen38fp8-tep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:qwen38-flash-next \
  --model Qwen/Qwen3.8-Flash-Next-FP8 \
  --tensor-parallel-size 8 --enable-expert-parallel --moe-backend triton \
  --gpu-memory-utilization 0.85 --max-num-seqs 256 --enable-prefix-caching \
  --no-enable-flashinfer-autotune --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --served-model-name qwen38fp8 \
  --host 0.0.0.0 --port 8000
bash

vLLM throughput — vllm-qwen38fp8-dep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  -e VLLM_ENGINE_READY_TIMEOUT_S=3600 \
  vllm/vllm-openai:qwen38-flash-next \
  --model Qwen/Qwen3.8-Flash-Next-FP8 \
  --data-parallel-size 8 --enable-expert-parallel --moe-backend triton \
  --gpu-memory-utilization 0.85 --max-num-seqs 256 --enable-prefix-caching \
  --no-enable-flashinfer-autotune --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder --reasoning-parser qwen3 \
  --served-model-name qwen38fp8 \
  --host 0.0.0.0 --port 8000
bash

XiaomiMiMo/MiMo-V2.6-Flash-RL#

Stable vLLM cannot load the MXFP4-stored weights at all, so these use the image published for the series rather than a release tag.

vLLM latency — vllm-flash-tp (the published command, verbatim)

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  vllm/vllm-openai:mimo-v26 \
  --model XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tensor-parallel-size 4 --trust-remote-code --gpu-memory-utilization 0.95 \
  --max-model-len auto --reasoning-parser mimo --tool-call-parser mimo \
  --enable-auto-tool-choice --generation-config vllm \
  --host 0.0.0.0 --port 8000
bash

vLLM latency + DFlash — vllm-flash-dflash-r3

$DFLASH is the dflash/ directory inside the downloaded snapshot. Note the --gpu-memory-utilization 0.90: at the published 0.95 the drafter has nowhere to allocate and the engine dies with a CUDA OOM.

DFLASH=$(ls -d /fsx/hf-cache/hub/models--XiaomiMiMo--MiMo-V2.6-Flash-RL/snapshots/*/dflash)

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  vllm/vllm-openai:mimo-v26 \
  --model XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tensor-parallel-size 4 --trust-remote-code --gpu-memory-utilization 0.90 \
  --max-model-len auto \
  --speculative-config '{"method":"dflash","model":"'"$DFLASH"'","num_speculative_tokens":7}' \
  --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
  --reasoning-parser mimo --tool-call-parser mimo \
  --enable-auto-tool-choice --generation-config vllm \
  --host 0.0.0.0 --port 8000
bash

vLLM balanced — vllm-flash-tep

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  vllm/vllm-openai:mimo-v26 \
  --model XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --trust-remote-code --gpu-memory-utilization 0.95 --max-model-len auto \
  --reasoning-parser mimo --tool-call-parser mimo \
  --enable-auto-tool-choice --generation-config vllm \
  --host 0.0.0.0 --port 8000
bash

SGLang — sgl-flash-cell (the published H200 cell, verbatim)

docker run --rm --gpus '"device=0,1,2,3"' --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:dev \
  sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Flash-RL \
  --tp 4 --moe-runner-backend marlin --trust-remote-code \
  --reasoning-parser mimo --tool-call-parser mimo \
  --host 0.0.0.0 --port 30000
bash

XiaomiMiMo/MiMo-V2.6-Pro-RL#

The 1T checkpoint. Identical flags, TP8, and the Pro repository.

vLLM latency — vllm-pro-tp

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  vllm/vllm-openai:mimo-v26 \
  --model XiaomiMiMo/MiMo-V2.6-Pro-RL \
  --tensor-parallel-size 8 --trust-remote-code --gpu-memory-utilization 0.95 \
  --max-model-len auto --reasoning-parser mimo --tool-call-parser mimo \
  --enable-auto-tool-choice --generation-config vllm \
  --host 0.0.0.0 --port 8000
bash

vLLM latency + DFlash — vllm-pro-dflash-r3

DFLASH=$(ls -d /fsx/hf-cache/hub/models--XiaomiMiMo--MiMo-V2.6-Pro-RL/snapshots/*/dflash)

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 8000:8000 -v /fsx/hf-cache:/root/.cache/huggingface \
  vllm/vllm-openai:mimo-v26 \
  --model XiaomiMiMo/MiMo-V2.6-Pro-RL \
  --tensor-parallel-size 8 --trust-remote-code --gpu-memory-utilization 0.90 \
  --max-model-len auto \
  --speculative-config '{"method":"dflash","model":"'"$DFLASH"'","num_speculative_tokens":7}' \
  --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' \
  --reasoning-parser mimo --tool-call-parser mimo \
  --enable-auto-tool-choice --generation-config vllm \
  --host 0.0.0.0 --port 8000
bash

SGLang — sgl-pro-cell

docker run --rm --gpus all --shm-size 32g --ulimit memlock=-1 --ipc=host \
  -p 30000:30000 -v /fsx/hf-cache:/root/.cache/huggingface \
  lmsysorg/sglang:dev \
  sglang serve --model-path XiaomiMiMo/MiMo-V2.6-Pro-RL \
  --tp 8 --moe-runner-backend marlin --trust-remote-code \
  --reasoning-parser mimo --tool-call-parser mimo \
  --host 0.0.0.0 --port 30000
bash

Single request performance#

100 200 300 One request at a time · CHAT-S · output tok/s MiMo-V2.6-Flash-RL vLLM latency + DFlash · 2.70 ms / token 345 DeepSeek-V4.1-Flash vLLM latency · 3.32 ms / token 294 MiMo-V2.6-Pro-RL vLLM latency + DFlash · 3.38 ms / token 273 MiMo-V2.6-Flash-RL vLLM latency · 4.10 ms / token 232 MiMo-V2.6-Flash-RL SGLang · 4.83 ms / token 201 GLM-5.3 SGLang latency · 4.71 ms / token 195 Qwen3.8-Flash-Next FP8 SGLang latency · 5.03 ms / token 184 GLM-5.3-Flash SGLang latency · 5.31 ms / token 174 MiMo-V2.6-Flash-RL vLLM balanced · 5.60 ms / token 174 DeepSeek-V4.1-Flash vLLM throughput · 5.68 ms / token 167 Qwen3.8-Flash-Next bf16 SGLang latency · 5.76 ms / token 164 GLM-5.3-Flash vLLM latency · 6.10 ms / token 150 GLM-5.3 vLLM latency · 6.41 ms / token 147 GLM-5.3-Flash vLLM balanced · 6.44 ms / token 143 Qwen3.8-Flash-Next FP8 vLLM balanced · 7.02 ms / token 140 MiMo-V2.6-Pro-RL vLLM latency · 6.84 ms / token 140 GLM-5.3 vLLM balanced · 7.02 ms / token 137 Qwen3.8-Flash-Next bf16 SGLang throughput · 7.16 ms / token 137 Qwen3.8-Flash-Next FP8 SGLang throughput · 7.07 ms / token 137 MiMo-V2.6-Pro-RL SGLang · 7.28 ms / token 123 GLM-5.3-Flash SGLang throughput · 8.64 ms / token 113 MiMo-V2.6-Pro-RL vLLM balanced · 8.51 ms / token 113 MiMo-V2.6-Flash-RL vLLM throughput · 9.05 ms / token 109 GLM-5.3-Flash vLLM throughput · 9.95 ms / token 99 GLM-5.3 SGLang throughput · 10.21 ms / token 96 DeepSeek-V4.1-Flash SGLang latency · 14.97 ms / token 68 DeepSeek-V4.1-Flash SGLang throughput · 15.00 ms / token 68 MiMo-V2.6-Pro-RL vLLM throughput · 19.15 ms / token 55
A single request, CHAT-S shape (2,048 in → 512 out). The sublabel is the inter-token latency — how long a reader waits between words.

MiMo-V2.6-Flash-RL with DFlash is the fastest single stream in the set at 345 tok/s, and MiMo-V2.6-Pro-RL with DFlash is third at 273 — a 1T model sitting above every 300–700B configuration here, on the strength of one speculative-decoding flag. Without it the same two checkpoints fall to 232 and 140.

Behind them, DeepSeek-V4.1-Flash under vLLM reaches 294 tok/s at 3.3 ms between tokens — and the same model under SGLang’s published H200 cell is the slowest in the chart at 68 tok/s. That 4.3× gap is the one result here that is not a tuning trade-off: it is the same checkpoint on the same GPUs, and part of it is speculative decoding, which vLLM runs and the SGLang H200 cell does not. SGLang’s two published DeepSeek cells also return near-identical numbers at every point, so on this hardware they are effectively one configuration.

Otherwise the pattern is the one you would expect: low-latency recipes on top, high-throughput recipes at the bottom, and the largest model (GLM-5.3, 743B) holding its own at 195 tok/s because its EAGLE 5-step draft is nearly free at batch 1.

Concurrent request performance#

2,000 4,000 6,000 Under load · BATCH-D · 64 concurrent · output tok/s Qwen3.8-Flash-Next FP8 vLLM balanced · 64 concurrent 6,475 DeepSeek-V4.1-Flash vLLM throughput · 64 concurrent 5,700 GLM-5.3-Flash vLLM balanced · 64 concurrent 5,433 DeepSeek-V4.1-Flash vLLM latency · 64 concurrent 5,343 GLM-5.3-Flash vLLM latency · 64 concurrent 5,025 MiMo-V2.6-Flash-RL vLLM balanced · 64 concurrent 4,950 GLM-5.3-Flash vLLM throughput · 64 concurrent 4,490 Qwen3.8-Flash-Next FP8 SGLang latency · 64 concurrent 4,453 Qwen3.8-Flash-Next bf16 SGLang latency · 64 concurrent 4,446 GLM-5.3-Flash SGLang throughput · 64 concurrent 4,369 MiMo-V2.6-Flash-RL SGLang · 64 concurrent 4,358 Qwen3.8-Flash-Next bf16 SGLang throughput · 64 concurrent 4,327 Qwen3.8-Flash-Next FP8 SGLang throughput · 64 concurrent 4,276 MiMo-V2.6-Flash-RL vLLM latency + DFlash · 64 concurrent 4,154 GLM-5.3-Flash SGLang latency · 64 concurrent 4,093 GLM-5.3 SGLang throughput · 64 concurrent 4,009 MiMo-V2.6-Pro-RL vLLM latency + DFlash · 64 concurrent 3,198 MiMo-V2.6-Flash-RL vLLM latency · 64 concurrent 2,924 MiMo-V2.6-Pro-RL SGLang · 64 concurrent 2,869 MiMo-V2.6-Pro-RL vLLM latency · 64 concurrent 2,359 DeepSeek-V4.1-Flash SGLang throughput · 64 concurrent 2,184 DeepSeek-V4.1-Flash SGLang latency · 64 concurrent 2,184 MiMo-V2.6-Pro-RL vLLM balanced · 64 concurrent 2,096 GLM-5.3 vLLM balanced · 64 concurrent 2,051 MiMo-V2.6-Flash-RL vLLM throughput · 64 concurrent 2,020 GLM-5.3 vLLM latency · 64 concurrent 2,014 GLM-5.3 SGLang latency · 64 concurrent 1,516 MiMo-V2.6-Pro-RL vLLM throughput · 64 concurrent 1,103
Sustained output on BATCH-D (4,096 in → 8,192 out) at 64 concurrent requests. Not the peak across the sweep — peak falls in the 256-stream column, which is the one that did not reproduce.

The order inverts almost completely: the low-latency recipes fall to the bottom and the throughput-oriented ones rise. Qwen3.8-Flash-Next FP8 under vLLM leads at 6,475 tok/s, and the model that was fastest single-stream — DeepSeek under vLLM’s latency strategy — is now mid-table.

Complete results#

Aggregate output tok/s at 64 concurrent requests — the operating point where a second repetition agreed within 6%, so these are the soundest numbers here.

configurationAPI-SCHAT-SCODE-ICHAT-LCODE-ADOC-LBATCH-D
GLM-5.3-Flash
SGLang latency7591,6751,6089649572644,093
SGLang throughput1,5292,8401,8731,2488682224,369
vLLM latency1,1932,0382,0031,3051,2253865,025
vLLM balanced1,2572,2382,0971,3081,2163885,433
vLLM throughput1,4782,0911,8451,3171,6995874,490
GLM-5.3
SGLang latency8911,139616354289941,516
SGLang throughput1,0901,521930468392rej4,009
vLLM latency7601,238569320300952,014
vLLM balanced8011,239599328287992,051
vLLM throughput———————
DeepSeek-V4.1-Flash
SGLang latency437655624624n/cn/c2,184
SGLang throughput437655624n/cn/cn/c2,184
vLLM latency1,6202,9271,7311,006965n/c5,343
vLLM throughput1,3312,0951,333746992rej5,700
Qwen3.8-Flash-Next bf16
SGLang latency1,0321,6851,9361,1891,1553534,446
SGLang throughput1,8572,4031,8731,1001,1003594,327
Qwen3.8-Flash-Next FP8
SGLang latency1,1701,6811,9171,1981,3903424,453
SGLang throughput2,0752,8402,4961,2491,0392894,276
vLLM balanced2,4353,3983,1021,8382,0823786,475
vLLM throughput———————
MiMo-V2.6-Flash-RL
SGLang2,0752,8381,9861,0438252664,358
vLLM latency1,8563,2632,0261,2771,1233352,924
vLLM latency + DFlash1,9093,2852,4101,5071,3133704,154
vLLM balanced2,5114,3842,4761,9891,5874824,950
vLLM throughput1,5271,9601,7351,2381,2783862,020
MiMo-V2.6-Pro-RL
SGLang1,3111,9601,2856645711952,869
vLLM latency1,2952,3931,4389097392312,359
vLLM latency + DFlash1,3352,6791,6489878922553,198
vLLM balanced1,2052,3921,3178617122242,096
vLLM throughput7631,0916635255141931,103

rej — the server refused every request. n/c — no request finished inside the window.

Adding the MiMo models turns what used to be a clean sweep into a three-way split. Before they were measured, Qwen3.8-Flash-Next FP8 under vLLM won six of the seven benchmarks. It now wins three:

benchmarkbest configurationoutput tok/s
API-SMiMo-V2.6-Flash-RL, vLLM balanced2,511
CHAT-SMiMo-V2.6-Flash-RL, vLLM balanced4,384
CODE-IQwen3.8-Flash-Next FP8, vLLM balanced3,102
CHAT-LMiMo-V2.6-Flash-RL, vLLM balanced1,989
CODE-AQwen3.8-Flash-Next FP8, vLLM balanced2,082
DOC-LGLM-5.3-Flash, vLLM throughput587
BATCH-DQwen3.8-Flash-Next FP8, vLLM balanced6,475

MiMo-V2.6-Flash-RL takes the conversational shapes and the 32k one; Qwen keeps the code shapes and the decode-heavy batch shape; GLM-5.3-Flash still owns the 128k DOC-L column it led before. Three different models, and for five of the seven the same strategy — vLLM with expert parallelism.

No configuration is good at everything. The best API-S cell is mid-table on DOC-L; the best DOC-L cell is mid-table on API-S. If your traffic is one shape, benchmark that shape.

Results per benchmark#

The same 64-stream measurement, one panel per workload shape, each model shown with whichever of its configurations was fastest on that shape. The point of splitting it out is that the ordering genuinely changes between panels — no model wins everywhere, and two of them win nothing.

1,000 2,000 API-S · 2k in / 256 out · 64 concurrent · output tok/s MiMo-V2.6-Flash-RL vLLM balanced 2,511 Qwen3.8-Flash-Next FP8 vLLM balanced 2,435 Qwen3.8-Flash-Next bf16 SGLang throughput 1,857 DeepSeek-V4.1-Flash vLLM latency 1,620 GLM-5.3-Flash SGLang throughput 1,529 MiMo-V2.6-Pro-RL vLLM latency + DFlash 1,335 GLM-5.3 SGLang throughput 1,090
API-S · 2,048 in → 256 out. Short API calls. The shape where prefill dominates and the engine has least room to hide.
2,000 4,000 CHAT-S · 2k in / 512 out · 64 concurrent · output tok/s MiMo-V2.6-Flash-RL vLLM balanced 4,384 Qwen3.8-Flash-Next FP8 vLLM balanced 3,398 DeepSeek-V4.1-Flash vLLM latency 2,927 GLM-5.3-Flash SGLang throughput 2,840 MiMo-V2.6-Pro-RL vLLM latency + DFlash 2,679 Qwen3.8-Flash-Next bf16 SGLang throughput 2,403 GLM-5.3 SGLang throughput 1,521
CHAT-S · 2,048 in → 512 out. A short chat turn — the most common interactive shape.
1,000 2,000 3,000 CODE-I · 16k in / 2k out · 64 concurrent · output tok/s Qwen3.8-Flash-Next FP8 vLLM balanced 3,102 MiMo-V2.6-Flash-RL vLLM balanced 2,476 GLM-5.3-Flash vLLM balanced 2,097 Qwen3.8-Flash-Next bf16 SGLang latency 1,936 DeepSeek-V4.1-Flash vLLM latency 1,731 MiMo-V2.6-Pro-RL vLLM latency + DFlash 1,648 GLM-5.3 SGLang throughput 930
CODE-I · 16,384 in → 2,048 out. Code completion with a file of context in the prompt.
1,000 2,000 CHAT-L · 32k in / 2k out · 64 concurrent · output tok/s MiMo-V2.6-Flash-RL vLLM balanced 1,989 Qwen3.8-Flash-Next FP8 vLLM balanced 1,838 GLM-5.3-Flash vLLM throughput 1,317 Qwen3.8-Flash-Next bf16 SGLang latency 1,189 DeepSeek-V4.1-Flash vLLM latency 1,006 MiMo-V2.6-Pro-RL vLLM latency + DFlash 987 GLM-5.3 SGLang throughput 468
CHAT-L · 32,768 in → 2,048 out. A long conversation carrying its own history forward.
1,000 2,000 CODE-A · 64k in / 4k out · 64 concurrent · output tok/s Qwen3.8-Flash-Next FP8 vLLM balanced 2,082 GLM-5.3-Flash vLLM throughput 1,699 MiMo-V2.6-Flash-RL vLLM balanced 1,587 Qwen3.8-Flash-Next bf16 SGLang latency 1,155 DeepSeek-V4.1-Flash vLLM throughput 992 MiMo-V2.6-Pro-RL vLLM latency + DFlash 892 GLM-5.3 SGLang throughput 392
CODE-A · 65,536 in → 4,096 out. Whole-repository analysis: a large prompt and a substantial answer.
200 400 600 DOC-L · 128k in / 2k out · 64 concurrent · output tok/s GLM-5.3-Flash vLLM throughput 587 MiMo-V2.6-Flash-RL vLLM balanced 482 Qwen3.8-Flash-Next FP8 vLLM balanced 378 Qwen3.8-Flash-Next bf16 SGLang throughput 359 MiMo-V2.6-Pro-RL vLLM latency + DFlash 255 GLM-5.3 vLLM balanced 99
DOC-L · 131,072 in → 2,048 out. The 128k document shape. Prefill dominates completely, and the ranking here looks unlike every other panel.
2,000 4,000 6,000 BATCH-D · 4k in / 8k out · 64 concurrent · output tok/s Qwen3.8-Flash-Next FP8 vLLM balanced 6,475 DeepSeek-V4.1-Flash vLLM throughput 5,700 GLM-5.3-Flash vLLM balanced 5,433 MiMo-V2.6-Flash-RL vLLM balanced 4,950 Qwen3.8-Flash-Next bf16 SGLang latency 4,446 GLM-5.3 SGLang throughput 4,009 MiMo-V2.6-Pro-RL vLLM latency + DFlash 3,198
BATCH-D · 4,096 in → 8,192 out. Decode-heavy batch generation — the shape that rewards raw token production.

Two things are worth pulling out. DeepSeek-V4.1-Flash is missing from the DOC-L panel — not because it was slow, but because no request finished inside the window on any of its configurations, so there is no rate to plot. And MiMo-V2.6-Pro-RL, the 1T model, places mid-table or last on every shape except the single-stream chart further up: parameter count buys quality, not tokens per second.

Speculative decoding: DFlash#

The largest single improvement in this whole comparison is not a parallelism strategy or an engine choice. It is one flag.

Both MiMo-V2.6 checkpoints ship a DFlash drafter inside the model repository. Pointing --speculative-config at it, changing nothing else except dropping --gpu-memory-utilization from 0.95 to 0.90 to leave the drafter room:

output tok/sMiMo-FlashMiMo-Pro
baseline+DFlashbaseline+DFlash
CHAT-S, 1 stream232345 (+48%)140273 (+95%)
CHAT-S, 64 streams3,2633,285 (+1%)2,3932,679 (+12%)
BATCH-D, 64 streams2,9244,154 (+42%)2,3593,198 (+36%)
DOC-L, 64 streams335370 (+10%)231255 (+10%)

DFlash nearly doubles the 1T model’s single-stream throughput and improves every shape measured on both models. The shape of the gain is what you would expect from speculation: largest where the GPU is idle waiting on a sequential decode (one stream: +48% and +95%), smallest where 64 concurrent requests already keep it busy (+1% and +12%). It does not disappear under load, though — BATCH-D, which is decode-heavy at 8,192 output tokens, still gains 36–42%.

If you serve either of these models, this is the first thing to turn on.

What would not run on H200#

Two configurations from the published recipes cannot run on this hardware, and both trace back to MXFP4 having no native SM90 path:

  • --moe-a2a-backend deepep together with Marlin. SGLang rejects the combination outright: Runner backend MoeRunnerBackend.MARLIN requires a fused func for a2a backend deepep, but none is registered. B300 avoids this because it runs deep_gemm; H200 is forced onto Marlin, which has no DeepEP kernel.
  • --attention-backend fa4. Fails during startup with ValueError: Expected size in shape to be strictly positive, but got 0, raised from CUTLASS’s layout builder. Isolated by changing one variable at a time: removing expert parallelism and keeping fa4 still fails; removing fa4 and keeping expert parallelism serves normally. This is a narrow claim — this model, this image, this node — not a general statement about FA4 on Hopper.

The MiMo cookbook page’s prose recommends keeping both of those on H200. Its machine-readable H200 cell omits both. The cell is right and the prose is wrong, which is a good argument for reading the config rather than the page. And the reduced version of the prose configuration that H200 can run still loses to the published cell: 167 against 201 tok/s single-stream, 2,151 against 2,838 at 64 streams.

Throughput vs concurrency#

GLM-5.3 GLM-5.3-Flash DeepSeek-V4.1-Flash Qwen3.8-Flash-Next bf16 Qwen3.8-Flash-Next FP8 MiMo-V2.6-Flash-RL MiMo-V2.6-Pro-RL 1k 2k 3k 4k 5k 6k 7k 1 16 64 concurrent requests 6,475 5,700 5,433 4,950 4,446 4,009 3,198
BATCH-D (4,096 in → 8,192 out), each model shown with its best configuration at 64 concurrent requests. Every model gains an order of magnitude from 1 to 64 streams; they do not gain it at the same rate, and the ordering at c1 is not the ordering at c64.

Every model shows the same shape, and none of them is close to saturated at 16 streams. Qwen3.8-Flash-Next FP8 starts second-slowest of the seven at one stream and finishes first at 64; MiMo-V2.6-Pro-RL does the opposite, leading at c1 and ending last. Ranking a model on single-stream numbers tells you very little about how it will serve a loaded endpoint.

Within a model, the low-latency and throughput recipes cross somewhere — and where they cross is not portable either:

modelcrossover
DeepSeek-V4.1-Flashbelow 16 concurrent requests
GLM-5.3between 16 and 64
GLM-5.3-Flashbetween 16 and 64
Qwen3.8-Flash-Next (bf16 and FP8)above 64
MiMo-V2.6-Flash-RLbetween 1 and 16 — but only against the balanced recipe
MiMo-V2.6-Pro-RLnever — the latency recipe leads at every concurrency measured

A rule of thumb learned on one model picks the wrong recipe for another. If you serve Qwen at 32 concurrent requests, the low-latency recipe is still the faster choice; at the same load DeepSeek has already crossed over.

Qwen3.8-Flash-Next Quantization comparison#

Comparison of Qwen3.8-Flash-Next BF16 checkpoint (336 GB) and FP8 (173 GB) checkpoints:

Qwen3.8-Flash-NextAPI-SCHAT-SCODE-IBATCH-D c64
BF16, SGLang throughput1,8572,4031,8734,327
FP8, SGLang throughput2,0752,8402,4964,276

On the decode-heavy shape the two are within 1.2% — half the weight memory for no measurable throughput change. On the shorter shapes FP8 is ahead by 12–33%, which is the more useful result if your traffic looks like API-S or CHAT-S.

(The 256-stream peaks are left out of this comparison on purpose: the only sound figure for either checkpoint at that concurrency is a re-measurement of the FP8 one, so a BF16-vs-FP8 comparison there would be mixing two different measurement windows.)

Conclusion#

Pick the tuning for the load you actually expect. Within a single model the low-latency and throughput recipes are different operating points, not better and worse versions of each other, and the crossover between them sits somewhere different for every model — so the choice cannot be carried over from one deployment to the next.

The engine matters less than the tuning for GLM-5.3-Flash, where SGLang and vLLM land within a few percent of each other at 64 concurrent requests. It matters a great deal elsewhere: vLLM is 17–51% ahead on Qwen3.8-Flash-Next FP8 and 2.5–4.5× ahead on DeepSeek-V4.1-Flash, on the same checkpoint and the same GPUs. Which way it goes is model-specific.

Look for a speculative drafter before tuning anything else. Enabling DFlash on MiMo-V2.6 beat every parallelism change tried on that model — +95% on single-stream Pro — and it is one flag pointing at a directory that is already inside the checkpoint. That is a better return than any strategy choice in this comparison.

Read the machine-readable recipe, not the page around it. Both MiMo configurations that refused to start on H200 were things the cookbook’s prose recommends and its own H200 cell omits, and both traced back to the same cause: MXFP4 has no native path on SM90, so Marlin is forced, and Marlin rules out the DeepEP kernel the prose asks for. The published cell had already accounted for the hardware.

Which is the practical point: none of this transfers. Benchmark your model, your engine and your workload shape before production, because every one of those three changed the answer here.

Serving LLMs on a single 8×H200 server
https://javicacheiro.com/blog/llm-benchmarking-on-h200
Author Javier Cacheiro López
Published at September 24, 2026