47 MRI-001 1
Same public model field. Different definition of work. Different order.
MRI-001 began with a larger question about AI capacity. The public ranking became the first bounded place to test whether changing the work could change what gets selected.
The question came first
MRI-001 did not begin with Qwen, Claude, or a leaderboard. It began with a larger question about AI demand: before models become tokens, compute, data centers and power, what work is actually being selected for?
The public ranking became the first bounded place where that question could be tested.
I had the question. I designed the test. AI gave me leverage to run it. The evidence was allowed to disagree with me.
THE PUBLIC BENCHMARK
Berkeley’s Function Calling Leaderboard combines different kinds of tool-use capability into one ranking. Its largest weights reward autonomy and sustained multi-turn work. Under that definition, Claude ranked 1 and Qwen ranked 47.
CHANGE THE WORK
MRI-001 defined a different job: short, repeatable, reviewable tool calls performed at scale. Because the job was already defined, autonomy received no weight. Correctness, cost, relevance, and speed became the definition of fitness.
Same public model field. Qwen ranked 1. Claude ranked 28.
The model did not change. The question did.
MRI-001 is not evidence that Qwen is better than Claude in general. It is evidence that “best model” is incomplete without “best for what?”
If different kinds of work require different levels of capability, model choice can affect cost, routing, compute, and eventually the infrastructure required to run AI.