CloudGYM

Can AI agents manage
shared cloud infrastructure?

GitHub
102 Live AWS cases
19 AWS services
65 Resource types
3 Task categories

Leaderboard

Consulted policy Hybrid tools

Models ranked by overall pass rate, with pass rates for each conflict type. Consulted policy, hybrid tools, 102 cases, two trials per case. All scores are percentages.
Rank Model Overall Semantic conflict Intent ambiguity Execution conflict
01
Claude Opus 5.5
34.3%
SC: Semantic conflict 17.4% IA: Intent ambiguity 32.9% EC: Execution conflict 80.6%
02
ChatGPT GPT-5.6 Sol
33.3%
SC: Semantic conflict 17.4% IA: Intent ambiguity 25.0% EC: Execution conflict 91.7%
03
ChatGPT GPT-5.5
28.9%
SC: Semantic conflict 8.7% IA: Intent ambiguity 26.3% EC: Execution conflict 86.1%
04
Claude Fable 5
28.4%
SC: Semantic conflict 4.3% IA: Intent ambiguity 30.3% EC: Execution conflict 86.1%
05
Claude Haiku 4.5
18.6%
SC: Semantic conflict 4.3% IA: Intent ambiguity 9.2% EC: Execution conflict 75.0%
View All