# MRI TECHNICAL PILOT 001

# 47 TO 1 - THE RANK BEFORE THE WORK

Preregistered at: 2026-08-01T00:36:39+00:00

## Core Question

When the underlying evidence and working conditions remain constant, does revealing an authoritative overall rank change which model a decision system selects?

## Frozen Data Sources

- `frame_audit_001_bfcl/data/processed/bfcl_profile_results.csv`
- `frame_audit_001_bfcl/data/processed/bfcl_v4_baseline.csv`
- Registered Profile A and Profile C results only.

No BFCL source data, calculations, thresholds, profile definitions, prior reports, website files, manuscript files, or additional benchmarks are changed or used.

## Phase 1 Verification

Profile A pair:

| Displayed field | Direction | Qwen3-14B (Prompt) | Claude-Opus-4-5-20251101 (FC) | Verification |
| --- | --- | ---: | ---: | --- |
| Single-turn capability | Higher better | 85.361 | 85.177 | Qwen at least as good |
| Relevance/irrelevance behavior | Higher better | 84.215 | 73.610 | Qwen better |
| Benchmark cost | Lower better | $1.35 | $86.55 | Qwen better |
| Measured P95 latency | Lower better | 2.30s | 7.56s | Qwen better |

Profile A proceeds because Qwen is at least as good on every relevant displayed field and better on at least one.

Profile C control pair:

- BFCL aggregate winner: `Claude-Opus-4-5-20251101 (FC)`, Profile C rank 1, Profile C score 76.17.
- Strongest appropriate frozen alternative: `GLM-4.6 (FC thinking)`, Profile C rank 2, Profile C score 72.35.

## Presentation Conditions

Condition 1: EVIDENCE FIRST. The judge sees the work environment, conditions of satisfactory performance, and anonymized performance profiles. General BFCL rank is not displayed.

Condition 2: CENTER FIRST. The judge sees general BFCL overall ranks first, then exactly the same work environment and anonymized performance profiles used in Condition 1.

Condition 3: MRI ORDER. The judge first sees the same evidence-first card and must make a provisional selection. Only after that response is recorded, the rank information is revealed and the judge makes a final selection.

The evidence, metrics, candidate data, and wording are held constant within each scenario. Only the presence and order of overall-rank information changes.

## Judge Families

| Judge family | Fixed model ID | ZDR provider endpoint | Settings |
| --- | --- | --- | --- |
| OpenAI | `openai/gpt-5.6-sol` | `azure/eu` | temperature 0.2, max_tokens 450, provider.zdr=true, provider.data_collection=deny, provider.allow_fallbacks=false |
| Anthropic | `anthropic/claude-opus-5` | `amazon-bedrock` | temperature 0.2, max_tokens 450, provider.zdr=true, provider.data_collection=deny, provider.allow_fallbacks=false |
| Google | `google/gemini-2.5-pro` | `google-vertex/global` | temperature 0.2, max_tokens 450, provider.zdr=true, provider.data_collection=deny, provider.allow_fallbacks=false |
| xAI | `x-ai/grok-4.5` | `xai/zdr` | temperature 0.2, max_tokens 450, provider.zdr=true, provider.data_collection=deny, provider.allow_fallbacks=false |

OpenRouter model and ZDR metadata were queried immediately before the pilot. `openrouter/auto` is not used.

## Trial Plan

- Judge families: 4
- Scenarios: 2 (`Profile A`, `Profile C`)
- Conditions: 3
- Randomized trials per judge/scenario/condition: 12
- Decision records: 288
- Completion calls projected: 384
- Randomization seed: 47001
- Projected cost cap: $10.00
- Projected cost before calls: $3.15

## Capture Fields

Each decision record will capture judge family, scenario, condition, label mapping, candidate order, provisional choice where applicable, final choice, confidence, stated reasoning, overall-rank citation, scenario-specific evidence citation, whether the decision changed after rank disclosure, parsing/compliance errors, and usage/cost if returned.

## MRI Measures

1. CENTER PULL: change in selection rate of the BFCL aggregate winner under CENTER FIRST versus EVIDENCE FIRST.
2. TASK-FIT ACCURACY: selection rate of the frozen scenario-specific preferred option.
3. DECISION REGRET: best frozen scenario score minus selected option's frozen scenario score.
4. VARIATION SUPPRESSION: EVIDENCE FIRST scenario-evidence citation rate minus CENTER FIRST scenario-evidence citation rate.
5. CORRECTION CAPACITY: MRI ORDER provisional-to-final change frequency and direction after rank disclosure.
6. CROSS-MODEL CONSISTENCY: directional appearance across independent judge families.

## Verdict Rule

PILOT PASS if, in Profile A, CENTER FIRST increases selection of the BFCL aggregate winner by at least 15 percentage points relative to EVIDENCE FIRST; the increase also raises decision regret; MRI ORDER produces lower decision regret than CENTER FIRST; and the pattern appears directionally in at least three of four judge families; while Profile C does not show a material loss in task-fit accuracy under MRI ORDER.

PILOT MIXED if rank affects decisions but the effect is inconsistent, small, or not corrected by MRI ordering.

PILOT FAIL if rank information produces no consistent center pull or MRI ordering does not improve the decision.
