This page has two tables of fidelity and three charts. They compare the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. The first table sorts ten models into three fidelity classes. Fidelity is the share of attempts that kept it completely, where a refusal is a full failure. Two models are in the top class, from 99 to 100: GPT-5.6 Sol and DeepSeek V4 Flash 0731. Three are in the middle class, from 90 to 99. Five are in the bottom class, below 90, and Claude Fable 5 is lowest at 45. The second table shows the nature of each failure, such as a refusal, a changed frame, or an added opinion. The first chart shows the index and the tasks that one dollar completes. Fifteen configurations are on this cost frontier. GPT-5.6 Luna at low effort is the cheapest, at approximately 114 tasks for one dollar. Claude Opus 5 at max effort has the highest index, at 63. The next chart shows the index and the tasks that one second completes. Fourteen configurations are on this speed frontier, and they all come from OpenAI or Anthropic. GPT-5.6 Luna without reasoning is the fastest, at approximately 15 seconds for one task. The third chart shows all three measurements together, and you can turn it. Twenty-one configurations are on the three-way frontier. A table of all the data follows the charts.

Model cost and speed Β· artificialanalysis.ai Β· retrieved Aug 12, 2026

Model intelligence, cost, and speed

This page compares the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. Each frontier model appears at all of its reasoning-effort levels. Dashed lines connect the levels of one model. Up and to the right is better. The blue line is the Pareto frontier: no other configuration is better on both scales.

Intelligence and fidelity

Fidelity measures how well a model does the job you asked for. A model loses fidelity when it changes the task, refuses part of it, leaves work out, or follows its own goal. These scores come from a separate benchmark, not from Artificial Analysis. The score is the share of attempts that kept fidelity completely, and a refusal counts as a full failure. The table sorts the ten models into three classes. Inside a class, the most capable model comes first.

ClassModelCompanyFidelityAttempts keptIndex

The nature of each failure

When a model loses fidelity, the judge names the pattern it saw. Six patterns can appear in the answer. Four more can appear in the reasoning summary, when the model shows them there. The bar covers every attempt. Its green part kept fidelity, and that share is the fidelity score in the table above. Each failed attempt appears once, in the color of its most serious problem. The numbers beside it count every appearance of a pattern, so they add up higher than the bar: one failed attempt often shows several patterns. Each pattern also has a mean severity from 1 to 3. A wide screen adds a grid of every model against every pattern. The list is ordered by fidelity plus a tenth of the index, so the best of both sits at the top.

Kept fidelity Refused Chose the answer first Sought permission Unproved claim Planned a refusal
NoneMost attempts 33 attempts for each model, and 31 for Kimi K3
Model Every
attempt
In the answer In the reasoning
RefusedOther
task
New
frame
Left work
out
Unproved
claim
Opinion Answer
first
Sought
permission
Planned
refusal
Planned
new frame

The cost frontier

The horizontal scale shows the tasks that one dollar completes. It is 1 divided by the cost of one task.

Pareto frontier Beaten Effort levels point or tab to a dot

The speed frontier

The horizontal scale shows the tasks that one second completes. It is 1 divided by the time of one task. This chart uses time in place of money. Muse Spark 1.2 and Grok 4.6 give no time data, so this chart does not show them.

Pareto frontier Beaten Effort levels point or tab to a dot

The three-way frontier

This chart shows all three measurements together. The vertical scale is the index. The two scales on the floor show the tasks for one dollar and for one second (both log). The blue surface is the frontier. A blue point at the corner of a step is equal to or better than all the configurations below the surface. The lines on the walls are the frontiers from the two charts above.

Three-way frontier Beaten Effort level Frontier surface Frontier lines on walls drag to turn · point or tab to a dot

Low effort costs very little

One dollar buys 21 to 114 tasks from GPT-5.6 Luna. The number depends on the effort level. Four of its six levels are on the cost frontier. At low effort, one task costs less than one cent. DeepSeek V4 Flash 0731 is better than Luna at xhigh effort.

More effort costs much more

The cost frontier starts at index 33.9 and ends at index 63.0. Along it, the cost of one task increases 266×. The effort levels show the same effect. Claude Opus 5 costs 5.5 times more at max effort than at low effort. It gives 10.6 more index points. The last step gives only 0.5 more points, but costs 1.3 times more.

A close result at the top

Claude Opus 5 at high effort keeps GPT-5.6 Sol at max effort off the cost frontier. The index of Opus 5 is 61.5, and the index of Sol is 60.9. The cost of one task is almost the same for both. Only the two highest levels of Claude Opus 5 are better than Claude Fable 5.

Closed models are the fastest

Every configuration on the speed frontier comes from OpenAI or Anthropic. Claude Fable 5 is on the speed frontier, but not on the cost frontier. The open-weight models are good on cost, but they are slow. Kimi K3 at max effort needs more than 9 minutes for one task.

Two good compromises

Two configurations are on this frontier only: GPT-5.6 Luna at xhigh effort and GPT-5.6 Terra at max effort. Other configurations are cheaper, and other configurations are faster. But nothing is better than these two on all three measurements. 21 of the 37 configurations with time data are on the three-way frontier.

All configurations

This table shows all the configurations. The order starts with the highest index. Point to a row to find the same configuration in the charts.

#ModelCompanyIndexCost per taskTasks per dollarTime per taskTasks per second