Leading LLMs for Coding & Agent Benchmarks

23 Models · 9 Benchmarks · AA Overall vs. Model Capability vs. Coding Agent + Harness
Loading latest data…
Imported from the original HTML. Price is Input / Output USD per 1M tokens.
正在加载模型数据…
Loading… 0 selected
Model InfoAA OverallModel CapabilityCoding Agent + Harness
ModelProviderPriceUSD / 1M tokens
Input / Output
AA Intelligence(v4.3.2) HLE(%) GDPval-AA v2.1(Elo) AutomationBench-AA(%) SciCodeCoding · %Terminal-Bench 4.0Model (%) Agent Harness DeepSWE v1.1(%) Terminal-Bench 4.0Harness (%) SWE-Atlas-QnA(%) Tokens / TaskAvg totalCost / TaskAvg API cost
No models match the current filters.
Note. Qwen3.8 Max (0902)* uses AA's “Claude Code + Qwen3.8 Max” agent result; AA does not explicitly label that agent run as “0902”.