This page has two tables of fidelity and three charts. They compare the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. The first table sorts eleven models into three fidelity classes: high, medium, and low. Fidelity is the share of attempts that kept it completely, where a refusal is a full failure. Two models are in the high class, from 99 to 100: GPT-5.6 Sol and DeepSeek V4 Pro 0813. Three are in the medium class, from 90 to 99. Six are in the low class, below 90, and Claude Fable 5 is lowest at 45. The second table shows the nature of each failure, such as a refusal, a changed frame, or an added opinion. It is ordered by fidelity. A key gives the meaning of each failure. Every heading sorts the table, and every label explains itself when you point at it. Above each chart, four buttons show or hide a fidelity class. They apply to all three charts and to the table at the end of the page. The first chart shows the index and the tasks that one dollar completes. Fifteen configurations are on this cost frontier. GPT-5.6 Luna at low effort is the cheapest, at approximately 114 tasks for one dollar. Claude Opus 5 at max effort has the highest index, at 63. The next chart shows the index and the tasks that one second completes. Fifteen configurations are on this speed frontier. They come from OpenAI, Anthropic, and Google. GPT-5.6 Luna without reasoning is the fastest, at approximately 14 seconds for one task. The third chart shows all three measurements together, and you can turn it. Twenty-two configurations are on the three-way frontier. A table of all the data follows the charts.
Model cost and speed Β· artificialanalysis.ai Β· retrieved Aug 14, 2026
This page compares the Artificial Analysis Intelligence Index with fidelity, with cost, and with speed. Each frontier model appears at all of its reasoning-effort levels. Dashed lines connect the levels of one model. Up and to the right is better. The blue line is the Pareto frontier: no other configuration is better on both scales. The buttons above each chart choose which fidelity classes all three charts show.
Fidelity measures how well a model does the job you asked for. A model loses fidelity when it changes the task, refuses part of it, leaves work out, or follows its own goal. These scores come from a separate benchmark, not from Artificial Analysis. The score is the share of attempts that kept fidelity completely, and a refusal counts as a full failure. The table sorts the eleven models into three classes: high (99 to 100), medium (90 to 99), and low (0 to 90). Each class has a colour, and that colour follows the model through the page. Click a heading to sort by that column. A sort by anything but the class gives each row a class of its own, in place of the group. Inside a class, the most capable model comes first. The buttons above each chart show or hide these classes. A model keeps its class at all of its effort levels, because the benchmark gives one score for the model.
| Class | Model | Company | Fidelity | Attempts kept | Index |
|---|
When a model loses fidelity, the judge names the pattern it saw. Six patterns can appear in the answer. Four more can appear in the reasoning summary, when the model shows them there. The bar covers every attempt. Its green part kept fidelity, and that share is the fidelity score in the table above. Each failed attempt appears once, in the color of its most serious problem. The numbers beside it count every appearance of a pattern, so they add up higher than the bar: one failed attempt often shows several patterns. Each pattern also has a mean severity from 1 to 3. A wide screen adds a grid of every model against every pattern. The list is ordered by fidelity, from the highest to the lowest; click a heading to order it by that pattern instead.
Every attempt
A failure in the answer
A failure in the reasoning
| Model | Every attempt |
In the answer | In the reasoning | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Refused | Other task | New frame | Left work out | Unproved claim | Opinion | Answer first | Sought permission | Planned refusal | Planned new frame |
||
The horizontal scale shows the tasks that one dollar completes. It is 1 divided by the cost of one task.
The horizontal scale shows the tasks that one second completes. It is 1 divided by the time of one task. This chart uses time in place of money. Muse Spark 1.2 gives no time data, so this chart does not show it.
This chart shows all three measurements together. The vertical scale is the index. The two scales on the floor show the tasks for one dollar and for one second (both log). The blue surface is the frontier. A blue point at the corner of a step is equal to or better than all the configurations below the surface. The lines on the walls are the frontiers from the two charts above.
One dollar buys 21 to 114 tasks from GPT-5.6 Luna. The number depends on the effort level. Five of its six levels are on the cost frontier. At low effort, one task costs less than one cent. Only its non-reasoning level is not on the frontier.
The cost frontier starts at index 33.9 and ends at index 63.0. Along it, the cost of one task increases 266×. The effort levels show the same effect. Claude Opus 5 costs 5.5 times more at max effort than at low effort. It gives 10.6 more index points. The last step gives only 0.5 more points, but costs 1.3 times more.
Claude Opus 5 at high effort keeps GPT-5.6 Sol at max effort off the cost frontier. The index of Opus 5 is 61.5, and the index of Sol is 60.9. The cost of one task is almost the same for both. Only the two highest levels of Claude Opus 5 are better than Claude Fable 5.
Every configuration on the speed frontier comes from a closed lab: OpenAI, Anthropic, or Google. Claude Fable 5 is on the speed frontier, but not on the cost frontier. The open-weight models are good on cost, but they are slow. Kimi K3 at max effort needs more than 9 minutes for one task.
Two configurations are on this frontier only: GLM-5.2 at max effort and GPT-5.6 Terra at max effort. Other configurations are cheaper, and other configurations are faster. But nothing is better than these two on all three measurements. 21 of the 37 configurations with time data are on the three-way frontier.
The order starts with the highest index. Click a heading to sort by that column, and click it again to turn the order around. Point to a row to find the same configuration in the charts. The fidelity column gives the score of the model, so every effort level of one model carries the same score and the same class. The fidelity buttons above each chart also choose the rows here.
| # | Model | Company | Fidelity | Class | Index | Cost per task | Tasks per dollar | Time per task | Tasks per second |
|---|