MRI-001 / RESEARCH NOTE

What the Work Selects

A public model hierarchy moved when the work became more specific. A stronger test for magnetic pull failed. The question then widened into training, compute and infrastructure.

What the Work Selects

MRI-001 began with a question about artificial intelligence before it becomes visible as infrastructure. Models, tokens, compute, data centers, electricity and capital usually appear near the end of the chain. By then, earlier decisions have already become demand: the work has been categorized, a model has been selected, and the workload has begun to determine the compute and infrastructure that follow.

The first question was not whether one model could be made to beat another. It was whether Magnetic Regression could move that discussion upstream. Before asking how much capacity AI needs, what happens if we ask what the capacity is for? The model-ranking experiment became the first bounded place where that question could be tested with public evidence.

MRI-001 therefore asked a narrow question inside a larger one. If the same public model field is evaluated against a different definition of work, does the selection change? If it does, the result does not prove that the new winner is universally better. It shows that the hierarchy depended on what the comparison had been built to select.

The Public Hierarchy

MRI-001 used the Berkeley Function Calling Leaderboard because it offered a public model field, public component evidence and a visible scoring structure. The frozen material reconstructed for this project contained 109 model and model-variant entries. The reconstruction report reproduced the published BFCL aggregate formula to the official two-decimal precision: 98 rows matched exactly at two decimals, and all 109 matched within 0.01 percentage points.

The published BFCL V4 aggregate was reconstructed as:

Agentic = (Web Search Acc + Memory Acc) / 2

Overall Acc = 0.40 * Agentic + 0.30 * Multi Turn Acc + 0.10 * Live Acc + 0.10 * Non-Live AST Acc + 0.10 * Irrelevance Detection

That composition placed 40 percent of the aggregate on agentic behavior, 30 percent on multi-turn behavior, 10 percent on live tasks, 10 percent on non-live tasks and 10 percent on irrelevance detection. In the frozen processed baseline, the published aggregate comparator was Claude-Opus-4-5-20251101 (FC) at official rank 1. Qwen3-14B (Prompt) was official rank 47.

Nothing in MRI-001 requires saying Berkeley was wrong. A benchmark can be carefully calculated, publicly documented and useful for the kind of work it defines. The calculation may be exact. The frame still came first.

First For What?

The first MRI move is not to ask which model is best. It is to ask: best for what?

The work profile in MRI-001 was high-volume, low-risk automation: routine CRM, ticket triage, back-office updates and other bounded tool calls where individual failures are reviewable and low consequence, but volume is high. The relevant action is narrow. The task is already specified. The result can be inspected. The operational risk is not that the model fails to conduct a long autonomous investigation. The risk is that repeated simple actions become wrong, slow, expensive, noisy or unnecessarily interventionist.

That distinction matters. A model may be excellent at frontier reasoning, long-horizon planning and broad agentic behavior while being an unnecessarily expensive default for repeated work that is already structured. A different model may be less impressive in a general hierarchy and still be more fit for a particular repeated workload.

MRI-001 does not collapse those claims. It separates them.

The Experiment

The Profile A definition used for MRI-001 was frozen in the provenance files before analysis. It required non-live AST simple, multiple, parallel and parallel-multiple categories; live AST simple, multiple, parallel and parallel-multiple categories; relevance and irrelevance detection; and cost and latency fields.

Its weights were:

Single-turn tool-call accuracy carried the largest weight because the work being modeled was a bounded action. Cost efficiency mattered because small differences compound at scale. Relevance and irrelevance behavior mattered because routine automation requires restraint as well as action. Latency mattered because repeated delay compounds inside workflows.

The hard disqualification constraints were also frozen: required single-turn categories could not be missing, irrelevance detection below 70 percent disqualified the model, P95 latency above 20 seconds disqualified the model, and missing total cost disqualified the model. In the recomputed Profile A field, 52 entries were eligible and 57 were disqualified.

Why no agentic weight? Because the tested work did not require the model to choose goals, sustain a long plan or perform broad autonomous search. MRI did not penalize capabilities the work did not require. It declined to reward them. That is not a claim against frontier capability. It is a claim against treating every workload as though it requires the same capability frontier.

The public BFCL data were reconstructed first. The profile definitions, weights, eligibility rules and pass/fail criteria were preserved in provenance files. The pass rule required at least one defensible buyer profile to select a different model than the published aggregate winner because of a material operational constraint rather than arbitrary weighting. The calculation and materiality audit then recomputed every Profile A-D score, eligibility decision, disqualification reason and rank from the processed baseline, reporting zero eligibility mismatches, zero disqualification-reason mismatches and zero rank mismatches.

How this was done

MRI-001 is AI-assisted research. I framed the question, designed the test and decided what the result would have to survive. AI models, Python and Codex gave me leverage to reconstruct evidence, implement the calculations, preserve outputs, run sensitivity tests and challenge the interpretation.

I had the question. I designed the test. AI gave me leverage to run it. The evidence was allowed to disagree with me.

That last sentence became literal when the stronger Magnetic Regression test failed. The failure remains published with the positive result.

47 → 1

Under Berkeley's published aggregate, Claude-Opus-4-5-20251101 (FC) ranked 1 and Qwen3-14B (Prompt) ranked 47.

Under the MRI-001 Profile A work definition, Qwen3-14B (Prompt) ranked 1 with a profile score of 90.93. Claude-Opus-4-5-20251101 (FC) ranked 28 with a profile score of 81.70.

The models did not change. The public evidence did not change. The work definition changed. Once the work was narrowed to frequent, reviewable tool calls performed repeatedly at scale, the rank order changed.

That is the central finding. It does not say Qwen is better than Claude in general. It does not say the Berkeley hierarchy is invalid for the work Berkeley is aggregating. It says a serious model-selection decision can move when the task definition becomes more specific.

Not Only Price

One immediate objection is that MRI-001 merely replaced a capability benchmark with a price screen. The evidence does not support that as the whole explanation. The calculation and materiality audit ran leave-one-factor-out checks for Profile A, removing one registered factor at a time and renormalizing the remaining factors. These were diagnostics; they did not redefine the frozen profile.

CheckReported winnerPublished rankSelected scorePublished-winner scoreFlip survives
Full registered Profile AQwen3-14B (Prompt)4790.9381.70Yes
Cost removedGPT-4.1-2025-04-14 (Prompt)4589.8385.06Yes
Latency removedQwen3-14B (Prompt)4789.3378.47Yes
Capability-only weightingGPT-4.1-2025-04-14 (Prompt)4587.2981.32Yes

Removing cost did not return the published aggregate winner to first. Removing latency did not return the published aggregate winner to first. A capability-only weighting also selected a different model. The exact selected model changed under some diagnostics, but the aggregate winner was not restored once the workflow frame was narrowed to Profile A.

The result is not that one model is the true winner independent of work. It is the opposite: no single winner appeared independent of the work definition and operational constraints.

The Stronger Test Failed

MRI-001 began with a question about how AI demand is formed. Before demand becomes compute, data centers and power, decisions have already been made about what work AI is expected to do and which models count as fit for that work.

The first test showed that this choice was not fixed. Change the work, and the model ranking changed.

But that was not yet Magnetic Regression.

For the stronger claim, the existing ranking would have to do more than describe the field. It would have to influence what was selected next.

MRI-001 therefore tested whether Berkeley's published Rank 1 could pull later AI decisions toward Claude even when more specific evidence about the work favored Qwen. The same decision was presented in different ways: once with the work-specific evidence first, and once with the established Rank 1 made salient first.

If the recognized winner changed the later decision, that would be evidence that the center was beginning to reproduce itself.

It didn't.

Across the valid Profile A decisions, Qwen was selected whether the work-specific evidence came first or Berkeley's Rank 1 came first. The measured center pull was 0.00 percentage points.

That is why the preregistered verdict was:

PILOT FAIL

The pilot had planned 288 decision records across OpenAI, Anthropic, Google and xAI judge families. Of those, 149 completed without parsing or API errors; 139 produced parsing or compliance errors. All 288 row-level results remain in the evidence package.

What That Failure Means

The distinction is simple.

The first test asked whether the center could move. It could.

The second asked whether the existing center could pull the next decision back toward itself. Under these conditions, it could not.

The ranking was therefore contingent, but not shown to be magnetic.

That matters because Magnetic Regression is not the claim that every category, ranking or model automatically becomes self-reinforcing. A representation can simplify reality without gaining enough authority to govern what happens next. A center can exist without becoming magnetic.

MRI-001 therefore produced two linked findings. The model hierarchy changed when the work changed. The stronger test for self-reinforcing center pull failed.

Both results stay in the record.

Upstream: What Selects Capability?

Once the model-selection hierarchy moved, the question widened upstream. If a benchmark ranking depends on the work definition, then training and evaluation decisions may also depend on hidden definitions of work. What data domains were sampled? How were those domains weighted? Which tasks were treated as central? Which losses mattered? Which abilities were evaluated early enough to steer model development?

MRI-001 does not answer those questions. The public evidence package does not include the operational training data needed to answer them. A stronger upstream analysis would require domain inventories, sampling weights, token allocations, checkpoint evaluations, transfer results, loss curves, compute measurements and ablation results.

That limit is part of the result. Public rankings can reveal how a public comparison behaves under alternate work definitions. They cannot, by themselves, reveal every training decision that produced the models being compared.

Downstream: What Becomes Compute And Power?

The same logic also widens downstream. A task selects a model. The model and task create a workload. The workload begins determining compute, accelerators, memory, storage, networking, architecture, cooling, infrastructure and electricity.

AI demand is not one kind of demand. Frontier scientific reasoning, coding agents, medical systems, repetitive tool calls, private inference, synthetic media and disposable output can all consume AI capacity. They do not necessarily require the same model, hardware, architecture, latency, reliability or amount of centralized compute.

MRI-001 began with model selection, not energy measurement. It did not measure electricity use, data-center load, accelerator utilization, cooling demand or infrastructure buildout. It should not be cited as though it did. But the ranking movement opens the downstream research question: if different kinds of work are routed to different kinds of intelligence, the resulting compute demand may change.

The infrastructure should follow the work, not the other way around.

Where The Public Evidence Ends

MRI-001 is strongest where public evidence is sufficient: reconstructing a public ranking, applying a frozen work profile, reporting the rank movement, auditing materiality and preserving the failed center-pull pilot.

It is weakest where public evidence stops. Operational deployment would require current pricing, current model availability, buyer-specific prompts, actual tool schemas, routing constraints, batching behavior, output lengths, hidden reasoning controls, latency under production load, review costs, error costs, escalation rates, hardware deployment paths and energy measurements. None of those should be invented to make the result feel complete.

Public evidence can reveal that a public hierarchy depends on a prior definition of work. It cannot replace the operational evidence needed to decide a real deployment, infrastructure plan or power forecast.

What MRI-001 Found

MRI-001 did not prove that Qwen is generally better than Claude. It did not prove that Berkeley's leaderboard is wrong. It did not prove that every public center creates magnetic pull. It did not measure electricity use, data-center demand, hardware utilization or cooling. It did not reconstruct model training histories or establish a universal routing rule for all AI products.

Methods & Evidence →Download public evidence →Back to MRI →