Community submitted benchmarks for AI models running on edge devices.

Every result comes from a signed report produced on community hardware. Browse by model, quantisation and chip.

Current best runs
Runtime comparisons coming soon

We’re working on standardised benchmark support for llama.cpp, vLLM, and MLX so equivalent workloads can be compared across runtimes. Current community rankings reflect BaseRT submissions.

How it works
01

Download the CLI

BaseRT runs locally on macOS, Linux, Windows and Android.

02

Run offline

No account is required. Signed reports remain on your machine until you submit them.

03

Submit when ready

Choose one or more saved reports. Sign in only if you want results attributed to you.

Leaderboard

Model × quantisation × chip

Headline rankings compare compatible PP512 prefill and TG128 decode workloads. The measured workload is shown beside every result.

Rank by
0 of 0 configurations · showing 0
No configurations match these filters