A fusion is one or more models scored on a public benchmark.
- Reproducible
- Ran on shared compute and stored on the global cache. Anyone can re-run it and get the same
score.
- Unverified
- Self-reported, or imported from a third-party board. Not yet reproduced on the global cache.
- Solo model
- A single model, shown as a reference baseline (no submitter).
- Reports
- Extra scores for a fusion or model, imported from third-party platforms (LMArena, OpenRouter,
…) and dated. Fusions themselves are only ever submitted by users.
Read this first
Treat this as a public scoreboard, not a trust oracle. Per benchmark, the top reproducible
result is the current SOTA (the one place gold appears), and an unverified claim can still rank
above it. The chart shows what accuracy each dollar of run-cost buys.
Benchmarks
Pick a benchmark to open its leaderboard. Inside, you can tab across all benchmarks.
Loading…
| Benchmark |
Focus |
Fusions |
Reproducible |
Best reproducible |
Open |