Models
Loading community benchmarks…
Loading community benchmarks…
Explore the models our community is running. Open a model to compare its submitted benchmarks across chips, formats, and backends.
Peak throughput is the highest reported value and may come from different runs. Compare model, quantisation, backend, and token counts before drawing conclusions.
Models are grouped and displayed by name, independent of the package provider. BaseRT quantisations published under basecompute/ are compared with a matching upstream model name—for example, basecompute/Qwen3-4B and Qwen/Qwen3-4B. Individual results retain the submitted package name and format.
Filters find models with a matching submitted configuration. Card totals and decode peaks cover all their runs; prefill uses the selected PP size.
Run a benchmark with BaseRT, then submit the saved report to help build the community dataset.