# MRI TECHNICAL PILOT 001 Technical Report

# 47 TO 1 - THE RANK BEFORE THE WORK

Completed at: 2026-08-01T00:43:11+00:00

## Scope

This pilot used only frozen BFCL Profile A and Profile C results. It did not alter BFCL data, profiles, thresholds, calculations, prior reports, or website files. It did not use the manuscript and did not introduce another benchmark.

## Experimental Design

- Judge families: OpenAI, Anthropic, Google, xAI
- Decision records planned: 288
- Decision records completed without parsing/API errors: 149
- Completion calls: EVIDENCE FIRST and CENTER FIRST used one call per decision; MRI ORDER used a provisional call followed by a rank-disclosure final call.
- Randomization seed: 47001
- Generation settings: temperature 0.2, max_tokens 450, fixed model IDs, ZDR provider endpoints, data collection denied, fallbacks disabled.
- Total returned/estimated OpenRouter cost: $1.4180

## Phase 1 Pair Verification

Profile A comparison was valid for the causal pilot: Qwen3-14B (Prompt) was at least as good as Claude-Opus-4-5-20251101 (FC) on every displayed Profile A-relevant field and better on at least one. In fact, Qwen was slightly higher on single-turn capability, higher on relevance/irrelevance behavior, lower cost, and lower P95 latency.

Profile C control pair used the BFCL aggregate winner, Claude-Opus-4-5-20251101 (FC), and the strongest appropriate frozen Profile C alternative, GLM-4.6 (FC thinking).

## Main Metrics

### Profile A

- CENTER PULL: 0.00 percentage points.
- EVIDENCE FIRST aggregate-winner selection rate: 0.000.
- CENTER FIRST aggregate-winner selection rate: 0.000.
- EVIDENCE FIRST task-fit accuracy: 1.000.
- CENTER FIRST task-fit accuracy: 1.000.
- MRI ORDER task-fit accuracy: 1.000.
- EVIDENCE FIRST mean regret: 0.000.
- CENTER FIRST mean regret: 0.000.
- MRI ORDER mean regret: 0.000.
- VARIATION SUPPRESSION: 0.00 percentage points.
- MRI ORDER change rate after rank disclosure: 0.000.
- MRI ORDER changes toward aggregate winner: 0.
- MRI ORDER changes away from aggregate winner: 0.

### Profile C

- CENTER PULL: 4.17 percentage points.
- EVIDENCE FIRST task-fit accuracy: 0.958.
- CENTER FIRST task-fit accuracy: 1.000.
- MRI ORDER task-fit accuracy: 1.000.
- EVIDENCE FIRST mean regret: 0.159.
- CENTER FIRST mean regret: 0.000.
- MRI ORDER mean regret: 0.000.

## Cross-Model Consistency

{
  "Anthropic": {
    "center_pull_pp": "",
    "directional_pattern": false
  },
  "Google": {
    "center_pull_pp": "",
    "directional_pattern": false
  },
  "OpenAI": {
    "center_pull_pp": 0.0,
    "directional_pattern": false
  },
  "xAI": {
    "center_pull_pp": 0.0,
    "directional_pattern": false
  }
}

## Errors

Parsing or compliance errors: 139 of 288 decision records.

## Pre-Registered Verdict

PILOT FAIL
