Models ranked by overall pass rate, with pass rates for each conflict
type. Consulted policy, hybrid tools, 102 cases, two trials per case. All
scores are percentages.
Rank
Model
Overall ↓
Semantic conflict
Intent ambiguity
Execution conflict
01
Opus 5.5
34.3%
SC: Semantic conflict 17.4%
IA: Intent ambiguity 32.9%
EC: Execution conflict 80.6%
02
GPT-5.6 Sol
33.3%
SC: Semantic conflict 17.4%
IA: Intent ambiguity 25.0%
EC: Execution conflict 91.7%
03
GPT-5.5
28.9%
SC: Semantic conflict 8.7%
IA: Intent ambiguity 26.3%
EC: Execution conflict 86.1%
04
Fable 5
28.4%
SC: Semantic conflict 4.3%
IA: Intent ambiguity 30.3%
EC: Execution conflict 86.1%
05
Haiku 4.5
18.6%
SC: Semantic conflict 4.3%
IA: Intent ambiguity 9.2%
EC: Execution conflict 75.0%
Pass rate by task category
Hybrid tools
Semantic conflict
Intent ambiguity
Execution conflict
PASS RATE (%)
Pass rate by task category: Consulted policy, hybrid tools.
Model
Semantic conflict
Intent ambiguity
Execution conflict
Opus 5.5
17.4%
32.9%
80.6%
GPT-5.6 Sol
17.4%
25.0%
91.7%
GPT-5.5
8.7%
26.3%
86.1%
Fable 5
4.3%
30.3%
86.1%
Haiku 4.5
4.3%
9.2%
75.0%
Pass rate by task category: Prompted policy, hybrid tools.
Model
Semantic conflict
Intent ambiguity
Execution conflict
Opus 5.5
40.2%
57.9%
88.9%
GPT-5.6 Sol
27.2%
47.4%
88.9%
GPT-5.5
19.6%
38.2%
88.9%
Fable 5
38.0%
50.0%
91.7%
Haiku 4.5
5.4%
13.2%
69.4%
Pass rate by task category: Isolated control, hybrid tools.
Model
Semantic conflict
Intent ambiguity
Execution conflict
Opus 5.5
98.9%
97.4%
91.7%
GPT-5.6 Sol
98.9%
93.4%
97.2%
GPT-5.5
97.8%
93.4%
100.0%
Fable 5
95.7%
94.7%
100.0%
Haiku 4.5
85.9%
77.6%
86.1%
Inference cost
Hybrid tools
Cost
USD PER CASE
Inference cost: Consulted policy, hybrid tools.
Model
Inference cost
Opus 5.5
$0.16
GPT-5.6 Sol
$0.27
GPT-5.5
$0.35
Fable 5
$0.98
Haiku 4.5
$0.11
Inference cost: Prompted policy, hybrid tools.
Model
Inference cost
Opus 5.5
$0.16
GPT-5.6 Sol
$0.25
GPT-5.5
$0.38
Fable 5
$0.91
Haiku 4.5
$0.11
Inference cost: Isolated control, hybrid tools.
Model
Inference cost
Opus 5.5
$0.12
GPT-5.6 Sol
$0.21
GPT-5.5
$0.27
Fable 5
$0.66
Haiku 4.5
$0.09
Tokens consumed
TOKENS PER CASE (THOUSANDS)
Tokens consumed: Consulted policy, hybrid tools.
Model
Tokens consumed
Opus 5.5
282.7K
GPT-5.6 Sol
264.4K
GPT-5.5
273.6K
Fable 5
406.8K
Haiku 4.5
592.9K
Tokens consumed: Prompted policy, hybrid tools.
Model
Tokens consumed
Opus 5.5
255.1K
GPT-5.6 Sol
234.5K
GPT-5.5
290.3K
Fable 5
360.0K
Haiku 4.5
540.3K
Tokens consumed: Isolated control, hybrid tools.
Model
Tokens consumed
Opus 5.5
207.3K
GPT-5.6 Sol
189.5K
GPT-5.5
200.3K
Fable 5
279.7K
Haiku 4.5
461.4K
Other participants change the infrastructure. The agent must consult them to discover how to resolve conflicts.
Other participants change the infrastructure. Rules for resolving conflicts are provided in the agent’s prompt.
The agent works alone, with no concurrent participants or interference.
The agent can freely combine Python scripts with the AWS SDK, command-line tools, and Terraform.
Semantic conflict: two participants require incompatible outcomes. The agent must follow the resolution policy to decide which requirement takes precedence.
Intent ambiguity: concurrent changes make the intended target or value unclear. The agent must clarify which outcome is intended.
Execution conflict: participants have compatible goals, but overlapping operations temporarily block progress. The agent must wait, retry, or adapt.