# Experimental Cards

These are the exact wording templates used for MRI TECHNICAL PILOT 001. Candidate label assignment and display order are randomized on every run. The canonical examples below use one illustrative label assignment; actual mappings are recorded in `03_RESULTS.csv`.

Condition 1, EVIDENCE FIRST, displays only the Evidence Block.

Condition 2, CENTER FIRST, displays the Rank Block first, followed by the same Evidence Block.

Condition 3, MRI ORDER, first displays the Evidence Block and records a provisional selection; only then does it display the Rank Block and record the final selection.

## Profile A: High-volume low-risk automation

### Evidence Block

Work environment:
A product team is choosing a tool-calling model for routine CRM, ticket triage, and back-office automations. Individual errors are reviewable and low consequence, but call volume is high.

Conditions of satisfactory performance:
Satisfactory performance means strong single-turn tool-call capability, reliable relevance/irrelevance behavior, low benchmark cost, and low measured tail latency.

Anonymized performance profiles:
Model A:
- single-turn capability: 85.177 (higher is better)
- relevance/irrelevance behavior: 73.61 (higher is better)
- benchmark cost: $86.55 (lower is better)
- measured P95 latency: 7.56 seconds (lower is better)
Model B:
- single-turn capability: 85.361 (higher is better)
- relevance/irrelevance behavior: 84.215 (higher is better)
- benchmark cost: $1.35 (lower is better)
- measured P95 latency: 2.3 seconds (lower is better)

### Rank Block

General BFCL overall ranks:
- Model A: BFCL overall rank 1
- Model B: BFCL overall rank 47

## Profile C: Multi-turn enterprise agent

### Evidence Block

Work environment:
An enterprise AI team is choosing a model for multi-step agent workflows such as account changes, procurement requests, technical support, workflow orchestration, and knowledge-work automation.

Conditions of satisfactory performance:
Satisfactory performance means strong multi-turn completion, memory behavior, web-search behavior, relevance/irrelevance behavior, acceptable latency, and cost as a secondary factor.

Anonymized performance profiles:
Model A:
- multi-turn performance: 68.38 (higher is better)
- memory performance: 73.76 (higher is better)
- web-search performance: 84.5 (higher is better)
- relevance/irrelevance behavior: 73.61 (higher is better)
- benchmark cost: $86.55 (lower is better)
- measured P95 latency: 7.56 seconds (lower is better)
Model B:
- multi-turn performance: 68 (higher is better)
- memory performance: 55.7 (higher is better)
- web-search performance: 77.5 (higher is better)
- relevance/irrelevance behavior: 79.98 (higher is better)
- benchmark cost: $4.64 (lower is better)
- measured P95 latency: 13.50 seconds (lower is better)

### Rank Block

General BFCL overall ranks:
- Model A: BFCL overall rank 1
- Model B: BFCL overall rank 4