😱 ScreamingFace
draft Fusion scores are internal results (Aug 12) pending publication; solo baselines are illustrative until the methodology page ships. For review only.
the verified fusion leaderboard

Fusion models are the real frontier

No single model wins every benchmark. Combine the models you already pay for and the combined system scores higher β€” we measured it, and you can re-run the measurement.

See the board β†’ Fusion Monsters β€” apply β†’
DRACO Β· open-context news benchmark Β· internal results, Aug 12
five-model fusion Β· ours68.6%
open-source fusion Β· ours61.9%
best single model *58.2%
budget flagship *52.4%

* solo bars illustrative pending the methodology page Β· verified = we re-ran the recipe Β· the full board β†’

β‰₯ $1.78/task Above that price, the fusion frontier sits above the solo frontier β€” more accuracy for the same spend. Below it, a single small model is still the right call. The board shows both, so you can pick by budget instead of by brand.
Every published score ships with its method. The recipe, the models, the cost, and a notebook that reproduces the run on free Colab compute β€” before your coffee cools. Read the method β†’ Reproduce it in Colab
the community program

Fusion Monsters

The hunt for the next verified score. Selected applicants get $50 of challenge credits on open models, a baseline we set with our own strong attempt, and a deadline. Beat it and you're in: a monthly allowance, the supported benchmarks, and a board where your win carries your name β€” until someone forks your recipe and beats you. 😱

allowance $300/mo admissions rolling Β· no rejections winning recipes published, under your name
Apply β€” takes 5 minutes β†’

Get the next one

One email per verified result. Nothing else.

😱 ScreamingFace β€” the verified fusion leaderboard