This page has two tables of fidelity and three charts. They compare the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. The first table sorts nine models into three fidelity classes: high, medium, and low. Fidelity is the share of attempts that kept it completely, where a refusal is a full failure. Two models are in the high class, from 99 to 100: GPT-5.6 Sol and DeepSeek V4 Pro 0813. Three are in the medium class, from 90 to 99. Four are in the low class, below 90, and Claude Fable 5 is lowest at 45. The second table shows the nature of each failure, such as a refusal, a changed frame, or an added opinion. It is ordered by fidelity. A key gives the meaning of each failure. Every heading sorts the table, and every label explains itself when you point at it. Above each chart, four buttons show or hide a fidelity class. They apply to all three charts and to the table at the end of the page. The first chart shows the index and the tasks that one dollar completes. Fifteen configurations are on this cost frontier. GPT-5.6 Luna at low effort is the cheapest, at approximately 114 tasks for one dollar. Claude Opus 5 at max effort has the highest index, at 63. The next chart shows the index and the tasks that one second completes. Fourteen configurations are on this speed frontier, and they all come from OpenAI or Anthropic. GPT-5.6 Luna without reasoning is the fastest, at approximately 16 seconds for one task. The third chart shows all three measurements together, and you can turn it. Twenty-one configurations are on the three-way frontier. A table of all the data follows the charts.

Model cost and speed Β· artificialanalysis.ai Β· retrieved Aug 13, 2026

Model intelligence, cost, and speed

This page compares the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. Each frontier model appears at all of its reasoning-effort levels. Dashed lines connect the levels of one model. Up and to the right is better. The blue line is the Pareto frontier: no other configuration is better on both scales. The buttons above each chart choose which fidelity classes all three charts show.

Intelligence and fidelity

Fidelity measures how well a model does the job you asked for. A model loses fidelity when it changes the task, refuses part of it, leaves work out, or follows its own goal. These scores come from a separate benchmark, not from Artificial Analysis. The score is the share of attempts that kept fidelity completely, and a refusal counts as a full failure. The table sorts the nine models into three classes: high (99 to 100), medium (90 to 99), and low (0 to 90). Each class has a colour, and that colour follows the model through the page. Click a heading to sort by that column. A sort by anything but the class gives each row a class of its own, in place of the group. Inside a class, the most capable model comes first. The buttons above each chart show or hide these classes. A model keeps its class at all of its effort levels, because the benchmark gives one score for the model.

ClassModelCompanyFidelityAttempts keptIndex

The nature of each failure

When a model loses fidelity, the judge names the pattern it saw. Six patterns can appear in the answer. Four more can appear in the reasoning summary, when the model shows them there. The bar covers every attempt. Its green part kept fidelity, and that share is the fidelity score in the table above. Each failed attempt appears once, in the color of its most serious problem. The numbers beside it count every appearance of a pattern, so they add up higher than the bar: one failed attempt often shows several patterns. Each pattern also has a mean severity from 1 to 3. A wide screen adds a grid of every model against every pattern. The list is ordered by fidelity, from the highest to the lowest; click a heading to order it by that pattern instead.

Every attempt

Kept fidelity The model did the task you asked for, and all of it.

A failure in the answer

Refused The model did not do the task. It gave no good reason.
Other task The model did a different task.
New frame The model changed the question, then answered its own question.
Left work out The model did only part of the task. It did not say which part.
Unproved claim The model gave a fact that its own evidence does not support.
Opinion The model added an opinion. You did not ask for one.

A failure in the reasoning

Answer first The model picked its answer first. Then it looked for reasons.
Sought permission The model asked itself if it was permitted to do the task.
Planned refusal The model planned to refuse, or to do less than you asked.
Planned new frame The model planned to change the question before it answered.
NoneMost attempts 33 attempts for each model, and 31 for Kimi K3
Model Every
attempt
In the answer In the reasoning
RefusedOther
task
New
frame
Left work
out
Unproved
claim
Opinion Answer
first
Sought
permission
Planned
refusal
Planned
new frame

The cost frontier

The horizontal scale shows the tasks that one dollar completes. It is 1 divided by the cost of one task.

Fidelity Pareto frontier Beaten Effort levels point or tab to a dot

The speed frontier

The horizontal scale shows the tasks that one second completes. It is 1 divided by the time of one task. This chart uses time in place of money. Muse Spark 1.2 gives no time data, so this chart does not show it.

Fidelity Pareto frontier Beaten Effort levels point or tab to a dot

The three-way frontier

This chart shows all three measurements together. The vertical scale is the index. The two scales on the floor show the tasks for one dollar and for one second (both log). The blue surface is the frontier. A blue point at the corner of a step is equal to or better than all the configurations below the surface. The lines on the walls are the frontiers from the two charts above.

Fidelity Three-way frontier Beaten Effort levels Frontier surface Frontier lines on walls drag to turn · point or tab to a dot

Low effort costs very little

One dollar buys 21 to 114 tasks from GPT-5.6 Luna. The number depends on the effort level. Five of its six levels are on the cost frontier. At low effort, one task costs less than one cent. Only its non-reasoning level is not on the frontier.

More effort costs much more

The cost frontier starts at index 33.9 and ends at index 63.0. Along it, the cost of one task increases 266×. The effort levels show the same effect. Claude Opus 5 costs 5.5 times more at max effort than at low effort. It gives 10.6 more index points. The last step gives only 0.5 more points, but costs 1.3 times more.

A close result at the top

Claude Opus 5 at high effort keeps GPT-5.6 Sol at max effort off the cost frontier. The index of Opus 5 is 61.5, and the index of Sol is 60.9. The cost of one task is almost the same for both. Only the two highest levels of Claude Opus 5 are better than Claude Fable 5.

Closed models are the fastest

Every configuration on the speed frontier comes from OpenAI or Anthropic. Claude Fable 5 is on the speed frontier, but not on the cost frontier. The open-weight models are good on cost, but they are slow. Kimi K3 at max effort needs more than 9 minutes for one task.

Two good compromises

Two configurations are on this frontier only: GLM-5.2 at max effort and GPT-5.6 Terra at max effort. Other configurations are cheaper, and other configurations are faster. But nothing is better than these two on all three measurements. 21 of the 37 configurations with time data are on the three-way frontier.

All configurations

The order starts with the highest index. Click a heading to sort by that column, and click it again to turn the order around. Point to a row to find the same configuration in the charts. The fidelity column gives the score of the model, so every effort level of one model carries the same score and the same class. The fidelity buttons above each chart also choose the rows here.

#ModelCompanyFidelityClassIndexCost per
task
Tasks per
dollar
Time per
task
Tasks per
second