Claude Haiku 5.5 Benchmarks: Scores and What They Mean
The October 7 release reports higher Haiku 5.5 scores than Haiku 4.5 across the listed evaluations, while Sonnet 5.5 remains ahead in that table. These results help you shortlist models; they do not measure your own task accuracy, response time or HaikuChat experience.
Haiku 5.5 benchmark results at a glance
The table below transcribes the model publisher's October 7, 2026 release evaluation, reviewed October 8. The comparison is with GPT-6 Luna, not an unspecified GPT-6 model. “Not reported” preserves an empty entry in that source rather than treating it as a zero.
| Evaluation and reported condition | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 · knowledge work · score | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 · knowledge work · score | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 · offline subset | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam · no tools | 45.9% | 10.2% | Not reported | 56.9% |
| Humanity's Last Exam · with tools | 57.4% | 18.7% | Not reported | 64.5% |
| Terminal-Bench 4.0 · agentic coding | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) · agentic coding | 46.4% | Not reported | 42.4% | 52.1% · Xhigh |
| Chartography · no tools · visual reasoning | 46.4% | 6.4% | 29.1% | 61.6% |
Source and scope: all eight rows come from one publisher's release table. HaikuChat did not run these evaluations. GDPval-AA and AA-Briefcase are shown as scores, not percentages; do not compare their magnitudes with the percentage rows. The FrontierCode Sonnet entry explicitly carries an Xhigh effort label. For the full evaluation methodology, follow the linked Haiku 5.5 system card. This summary does not establish that every model used an identical reasoning budget, harness or tool setup.
What the differences do and do not show
On Terminal-Bench 4.0, the reported Haiku 5.5 score is 22.8 percentage points above GPT-6 Luna and 31.4 percentage points below Sonnet 5.5. Those are within-row arithmetic differences, not predictions that a model will solve that many more of your coding tasks.
On the OSWorld offline subset, Haiku 5.5 is 56.7 percentage points above Haiku 4.5. That is a large change in this reported evaluation, but the offline-subset qualification matters: it is not a score for every website, browser session or live service.
Humanity's Last Exam illustrates why conditions must stay attached to a result. Haiku's reported 45.9% without tools and 57.4% with tools describe different setups. A tools-enabled run cannot be silently substituted for a plain chat result. There is no Luna HLE number in this release table, so it cannot establish a Haiku-versus-Luna ranking on that test.
Do not average these rows into an overall intelligence score. They use different units, task sets and setups. A high knowledge-work score, a terminal task success rate and a chart-reading result answer different questions. For a focused choice, read Haiku 5.5 vs GPT-6 Luna or Haiku 5.5 vs Sonnet 5.5.
Match the benchmark to the work you need
Coding: Terminal-Bench concerns complex command-line tasks; FrontierCode is a separate coding evaluation. Neither number directly measures how much editing a generated code snippet needs in your repository. Use representative bugs, existing project tests and a fixed scope when evaluating a coding model. The release positions the larger Sonnet and Opus tiers for complex agentic coding, while Haiku targets narrower work.
Summaries and knowledge work: GDPval-AA and AA-Briefcase provide broad knowledge-work signals. To choose a summarizer, check whether it preserves dates, owners, unresolved decisions and source limitations in your actual notes. A polished summary that invents a deadline should fail even if a model has a strong published score.
Visual extraction: Chartography measures visual reasoning. It does not establish invoice OCR accuracy or a guarantee that fields will be extracted correctly from screenshots. Include blurry text, conflicting totals and missing values in a separate extraction check. Score field correctness before rewarding table formatting.
Computer use: OSWorld evaluates agents operating a computer. HaikuChat's chat, summary and table workflow does not operate your computer or execute terminal commands. Model-level capabilities and a product's enabled features are separate.
If you need to evaluate a more capable tier, the Haiku 5.5 vs Opus 5.5 guide covers model specifications and task selection. Opus is absent from this release's benchmark table; this page does not invent an Opus column from another evaluation.
Benchmark scores do not measure response time or task cost
This score table contains no comparable tokens-per-second, time-to-first-token or end-to-end latency series. It therefore cannot support a numerical speed ranking. Customer examples in a release are also specific to those customers' systems; they are not HaikuChat latency measurements.
For an interactive workflow, record time to the first visible answer and time to a complete usable answer separately. A quick first token can coexist with a slow final response. Note input length, output limit, effort, image use, concurrency, network and provider before comparing timings. Use several attempts and report the distribution rather than choosing the fastest run.
For cost, compare total model spend divided by accepted completions, including failed attempts and retries. A cheaper request can become a more expensive workflow when it needs repeated correction. Reasoning settings, tokenization, caching and long-input thresholds affect billed usage; a benchmark accuracy percentage alone is not a cost-per-success estimate.
The Luna comparison's cost section explains why equal base rates can diverge for long inputs. HaikuChat pricing describes this workspace's own credit packs. Published model-token prices and the price of using a finished workspace are different quantities.
A reproducible evaluation for your own tasks
Start with ten examples you have permission to use: four routine cases, three ambiguous cases and three cases with missing or conflicting information. Write the reference facts or required fields before viewing any generated result. Keep a small holdout set when revising prompts so you do not merely optimize for the examples you have already seen.
| Task | Acceptance check | Failure to retain in the results |
|---|---|---|
| Meeting summary | Correct decisions, owners and dates; uncertainty preserved | Invented commitment or omitted unresolved issue |
| Field extraction | Exact values, currency and dates; missing values marked missing | Fabricated field or incorrect total |
| Small coding change | Existing tests pass and requested behavior is demonstrated | Broken behavior, unsupported dependency or manual repair |
Keep source material and requested output consistent. Record the exact model ID, prompt version, effort, available tools, completion limits and attempt count. Blind model names while checking outputs when possible. A trial should count towards the denominator even when it times out, refuses or needs a retry; record the reason rather than silently dropping it.
Summarize usable completions, factual errors, editing time, latency and billed usage. Agree on a task-specific acceptance threshold before choosing a winner. If neither model meets it, change the workflow or compare a different tier rather than treating a small score gap as proof of suitability.
This is an evaluation protocol, not a completed HaikuChat study. To try the current workspace, open HaikuChat and check its selected model and connection status. For the broader tier decision, see Haiku vs Sonnet vs Opus.
Haiku 5.5 benchmark questions
Where can I find the Haiku 5.5 benchmark scores?
The table above lists the publisher-reported October 7 results, including Terminal-Bench, OSWorld, Humanity's Last Exam and Chartography. Benchmark and benchmarks searches are covered here with the same source and review date.
Does Haiku 5.5 beat GPT-6 Luna and Sonnet 5.5?
In this release table, Haiku 5.5 is ahead of GPT-6 Luna in rows where both have a reported score, and behind Sonnet 5.5. Missing entries, effort differences and task fit limit that conclusion; it is not a universal ranking.
Are these independent HaikuChat benchmark tests?
No. These are attributed results from the model publisher. The local evaluation checklist is a suggested method, and no HaikuChat performance or speed measurement is claimed here.
Can I use these scores to compare Haiku 5.5 with Opus 5.5?
The source table has no Opus 5.5 column. A valid numerical comparison needs the same benchmark version, task set, harness and effort or budget details for both models. Different reports should not be silently joined.
Sources and review date
Specifications checked October 8, 2026. Prices, availability and evaluation settings can change. Published benchmark results are attributed to their source; we have not presented them as HaikuChat's own tests.
Use Haiku for your next small task
Chat, summarize pasted text or extract fields into a table. Check the workspace for current availability and task limits.
Continue comparing
Claude Haiku 5.5 vs GPT-6 Luna
Compare Haiku 5.5 and GPT-6 Luna on pricing tiers, context, published benchmarks, summaries, and extraction. Choose for your actual workload.
Read comparison →When is a step up worth it?Claude Haiku 5.5 vs Sonnet 5.5
Decide between Haiku 5.5 and Sonnet 5.5 for summaries, extraction, coding, and complex work. Compare costs, context, and published benchmark evidence.
Read comparison →Everyday throughput or deeper investigation?Claude Haiku 5.5 vs Opus 5.5
Compare Haiku 5.5 and Opus 5.5 for everyday tasks and complex analysis. Understand the cost gap, reasoning differences, and when to choose either model.
Read comparison →