- Serving LLMs on a single 8×H200 server
Benchmarking six open-weight models, including MiMo-V2.6, on one 8×H200 server.
38 min - Serving LLMs on a single 8×B300 server
Measuring the throughput of five large models on one 8xB300 server.
16 min - Measuring how much quantization affects model quality
I benchmark different variants of Qwen3.8-27B to measure the effect of quantization
24 min
Back