# BFCL V4 Pass/Fail Rules

These rules are frozen before analysis.

## Main Pass Rule

PASS if at least one defensible buyer profile selects a different model than the published aggregate winner, and the difference is driven by a material operational constraint rather than arbitrary weighting.

MIXED if rankings change but the financial or operational consequence is weak.

FAIL if reasonable profiles do not change the decision or merely reproduce information already obvious from the BFCL leaderboard.

## What Counts As A Material Operational Constraint

A rank reversal can count toward PASS only if at least one of the following is true:

- The published winner fails a pre-registered hard constraint for a buyer profile.
- A buyer-specific winner materially lowers expected failed workflow cost.
- A buyer-specific winner materially lowers human-review burden.
- A buyer-specific winner materially lowers excess inference cost at similar success.
- A buyer-specific winner materially lowers procurement, operational, or compliance exposure.
- A buyer-specific winner materially improves latency-adjusted success for a workflow where latency directly affects throughput or user experience.

## What Does Not Count As A Pass

- A rank change produced only by arbitrary or post-hoc weights.
- A rank change between models that are operationally indistinguishable for the buyer.
- A rank change that depends on imputed missing data.
- A rank change that simply repeats a fact already obvious from the BFCL table, such as a cheaper model being cheaper.
- A profile that excludes most models for no buyer-relevant reason.

## Required Reporting Before Conclusion

The final dossier must report:

- The published BFCL aggregate winner.
- The selected model for each buyer profile.
- Whether each selection differs from the published aggregate winner.
- Which hard constraints, if any, changed the model set.
- Which cost, latency, hallucination, multi-turn, or completeness factor drove the result.
- Whether the rank change would plausibly alter a procurement, deployment, investment, or product decision.

## Financial Consequence Translation

The audit will translate benchmark-frame differences into plain-language buyer risk:

- Failed workflow cost: a wrong or missed tool call can force rework, break automation, or cause customer-facing failure.
- Human-review burden: lower reliability pushes tasks back to staff, turning nominal automation into supervised work.
- Excess inference cost: a stronger but more expensive model can be financially dominated if cheaper models achieve similar workflow success.
- Procurement risk: a single aggregate leaderboard can justify buying the wrong default model for the actual workload.
- Operational or compliance exposure: in regulated settings, tool misuse may trigger audit, remediation, or reporting obligations.

## Pass/Fail Examples

PASS example:

- BFCL aggregate winner ranks first overall.
- Regulated workflow profile disqualifies it for insufficient irrelevance detection or missing-function behavior.
- Another model meets the safety thresholds with comparable cost/latency.
- The dossier can explain the financial consequence as avoided compliance exposure and reduced human review.

MIXED example:

- Buyer weighting changes the top model, but the score difference is tiny and no hard constraint or cost/latency issue is materially different.

FAIL example:

- All four buyer profiles select the same model as the published aggregate, or every reversal depends on weights that do not map to a real workflow.
