๐Ÿ˜ฑ screamingface/ benchmarks github docs
Leaderboard

Fusions, ranked and reproducible.

A fusion is one or more models scored on a public benchmark.

Reproducible
Ran on shared compute and stored on the global cache. Anyone can re-run it and get the same score.
Unverified
Self-reported, or imported from a third-party board. Not yet reproduced on the global cache.
Solo model
A single model, shown as a reference baseline (no submitter).
Reports
Extra scores for a fusion or model, imported from third-party platforms (LMArena, OpenRouter, …) and dated. Fusions themselves are only ever submitted by users.
Read this first Treat this as a public scoreboard, not a trust oracle. Per benchmark, the top reproducible result is the current SOTA (the one place gold appears), and an unverified claim can still rank above it. The chart shows what accuracy each dollar of run-cost buys.

Benchmarks

Pick a benchmark to open its leaderboard. Inside, you can tab across all benchmarks.

Loading…
๐Ÿ˜ฑ screamingface leaderboard.screamingface.ai ยท MVP preview ยท mock data