Leaderboard

Overall pass rate

Consulted policy Hybrid tools

Models ranked by overall pass rate, with pass rates for each conflict type. Consulted policy, hybrid tools, 102 cases, two trials per case. All scores are percentages.
Rank Model Overall Semantic conflict Intent ambiguity Execution conflict
01
Claude Opus 5.5
34.3%
SC: Semantic conflict 17.4% IA: Intent ambiguity 32.9% EC: Execution conflict 80.6%
02
ChatGPT GPT-5.6 Sol
33.3%
SC: Semantic conflict 17.4% IA: Intent ambiguity 25.0% EC: Execution conflict 91.7%
03
ChatGPT GPT-5.5
28.9%
SC: Semantic conflict 8.7% IA: Intent ambiguity 26.3% EC: Execution conflict 86.1%
04
Claude Fable 5
28.4%
SC: Semantic conflict 4.3% IA: Intent ambiguity 30.3% EC: Execution conflict 86.1%
05
Claude Haiku 4.5
18.6%
SC: Semantic conflict 4.3% IA: Intent ambiguity 9.2% EC: Execution conflict 75.0%

Pass rate by task category

Hybrid tools

  • Semantic conflict
  • Intent ambiguity
  • Execution conflict
PASS RATE (%)
0.0% 25.0% 50.0% 75.0% 100.0% 17.4% 32.9% 80.6% Opus 5.5 17.4% 25.0% 91.7% GPT-5.6 Sol 8.7% 26.3% 86.1% GPT-5.5 4.3% 30.3% 86.1% Fable 5 4.3% 9.2% 75.0% Haiku 4.5
Pass rate by task category: Consulted policy, hybrid tools.
Model Semantic conflictIntent ambiguityExecution conflict
Opus 5.5 17.4%32.9%80.6%
GPT-5.6 Sol 17.4%25.0%91.7%
GPT-5.5 8.7%26.3%86.1%
Fable 5 4.3%30.3%86.1%
Haiku 4.5 4.3%9.2%75.0%

Inference cost

Hybrid tools

Cost

USD PER CASE
$0.00 $0.25 $0.50 $0.75 $1.00 $0.16 Opus 5.5 $0.27 GPT-5.6 Sol $0.35 GPT-5.5 $0.98 Fable 5 $0.11 Haiku 4.5
Inference cost: Consulted policy, hybrid tools.
Model Inference cost
Opus 5.5 $0.16
GPT-5.6 Sol $0.27
GPT-5.5 $0.35
Fable 5 $0.98
Haiku 4.5 $0.11

Tokens consumed

TOKENS PER CASE (THOUSANDS)
0.0K 200.0K 400.0K 600.0K 800.0K 282.7K Opus 5.5 264.4K GPT-5.6 Sol 273.6K GPT-5.5 406.8K Fable 5 592.9K Haiku 4.5
Tokens consumed: Consulted policy, hybrid tools.
Model Tokens consumed
Opus 5.5 282.7K
GPT-5.6 Sol 264.4K
GPT-5.5 273.6K
Fable 5 406.8K
Haiku 4.5 592.9K