# Plain Language Result

The pilot tested whether showing the general BFCL overall rank before the work-specific evidence would pull independent decision systems toward the aggregate winner.

In Profile A, the work-specific evidence favored Qwen3-14B (Prompt): it was at least as good on the displayed task-fit fields and better on cost, latency, and relevance behavior. The experiment then compared evidence-first decisions against decisions where the overall BFCL rank was shown first.

The measured center pull for Profile A was 0.00 percentage points. Mean decision regret was 0.000 under CENTER FIRST and 0.000 under MRI ORDER.

Profile C served as a control case where the aggregate winner was also the frozen scenario-specific preferred option. MRI ORDER task-fit accuracy in Profile C was 1.000.

This pilot does not establish commercial value, does not create a new benchmark, and does not revise BFCL. It only tests whether authoritative rank placement can change model-selection behavior when the underlying work-specific evidence remains constant.

PILOT FAIL
