Claude Haiku 5.5 Benchmarks: Scores and What They Mean

The October 7 release reports higher Haiku 5.5 scores than Haiku 4.5 across the listed evaluations, while Sonnet 5.5 remains ahead in that table. These results help you shortlist models; they do not measure your own task accuracy, response time or HaikuChat experience.

Haiku 5.5 benchmark results at a glance

The table below transcribes the model publisher's October 7, 2026 release evaluation, reviewed October 8. The comparison is with GPT-6 Luna, not an unspecified GPT-6 model. “Not reported” preserves an empty entry in that source rather than treating it as a zero.

Evaluation and reported conditionHaiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 · knowledge work · score162073514371840
AA-Briefcase v1.1 · knowledge work · score157861413361824
OSWorld 2.1 · offline subset72.4%15.7%48.9%83.9%
Humanity's Last Exam · no tools45.9%10.2%Not reported56.9%
Humanity's Last Exam · with tools57.4%18.7%Not reported64.5%
Terminal-Bench 4.0 · agentic coding39.2%0.0%16.4%70.6%
FrontierCode 1.1 (Main) · agentic coding46.4%Not reported42.4%52.1% · Xhigh
Chartography · no tools · visual reasoning46.4%6.4%29.1%61.6%

Source and scope: all eight rows come from one publisher's release table. HaikuChat did not run these evaluations. GDPval-AA and AA-Briefcase are shown as scores, not percentages; do not compare their magnitudes with the percentage rows. The FrontierCode Sonnet entry explicitly carries an Xhigh effort label. For the full evaluation methodology, follow the linked Haiku 5.5 system card. This summary does not establish that every model used an identical reasoning budget, harness or tool setup.

What the differences do and do not show

On Terminal-Bench 4.0, the reported Haiku 5.5 score is 22.8 percentage points above GPT-6 Luna and 31.4 percentage points below Sonnet 5.5. Those are within-row arithmetic differences, not predictions that a model will solve that many more of your coding tasks.

On the OSWorld offline subset, Haiku 5.5 is 56.7 percentage points above Haiku 4.5. That is a large change in this reported evaluation, but the offline-subset qualification matters: it is not a score for every website, browser session or live service.

Humanity's Last Exam illustrates why conditions must stay attached to a result. Haiku's reported 45.9% without tools and 57.4% with tools describe different setups. A tools-enabled run cannot be silently substituted for a plain chat result. There is no Luna HLE number in this release table, so it cannot establish a Haiku-versus-Luna ranking on that test.

Do not average these rows into an overall intelligence score. They use different units, task sets and setups. A high knowledge-work score, a terminal task success rate and a chart-reading result answer different questions. For a focused choice, read Haiku 5.5 vs GPT-6 Luna or Haiku 5.5 vs Sonnet 5.5.

Match the benchmark to the work you need

Coding: Terminal-Bench concerns complex command-line tasks; FrontierCode is a separate coding evaluation. Neither number directly measures how much editing a generated code snippet needs in your repository. Use representative bugs, existing project tests and a fixed scope when evaluating a coding model. The release positions the larger Sonnet and Opus tiers for complex agentic coding, while Haiku targets narrower work.

Summaries and knowledge work: GDPval-AA and AA-Briefcase provide broad knowledge-work signals. To choose a summarizer, check whether it preserves dates, owners, unresolved decisions and source limitations in your actual notes. A polished summary that invents a deadline should fail even if a model has a strong published score.

Visual extraction: Chartography measures visual reasoning. It does not establish invoice OCR accuracy or a guarantee that fields will be extracted correctly from screenshots. Include blurry text, conflicting totals and missing values in a separate extraction check. Score field correctness before rewarding table formatting.

Computer use: OSWorld evaluates agents operating a computer. HaikuChat's chat, summary and table workflow does not operate your computer or execute terminal commands. Model-level capabilities and a product's enabled features are separate.

If you need to evaluate a more capable tier, the Haiku 5.5 vs Opus 5.5 guide covers model specifications and task selection. Opus is absent from this release's benchmark table; this page does not invent an Opus column from another evaluation.

Benchmark scores do not measure response time or task cost

This score table contains no comparable tokens-per-second, time-to-first-token or end-to-end latency series. It therefore cannot support a numerical speed ranking. Customer examples in a release are also specific to those customers' systems; they are not HaikuChat latency measurements.

For an interactive workflow, record time to the first visible answer and time to a complete usable answer separately. A quick first token can coexist with a slow final response. Note input length, output limit, effort, image use, concurrency, network and provider before comparing timings. Use several attempts and report the distribution rather than choosing the fastest run.

For cost, compare total model spend divided by accepted completions, including failed attempts and retries. A cheaper request can become a more expensive workflow when it needs repeated correction. Reasoning settings, tokenization, caching and long-input thresholds affect billed usage; a benchmark accuracy percentage alone is not a cost-per-success estimate.

The Luna comparison's cost section explains why equal base rates can diverge for long inputs. HaikuChat pricing describes this workspace's own credit packs. Published model-token prices and the price of using a finished workspace are different quantities.

A reproducible evaluation for your own tasks

Start with ten examples you have permission to use: four routine cases, three ambiguous cases and three cases with missing or conflicting information. Write the reference facts or required fields before viewing any generated result. Keep a small holdout set when revising prompts so you do not merely optimize for the examples you have already seen.

TaskAcceptance checkFailure to retain in the results
Meeting summaryCorrect decisions, owners and dates; uncertainty preservedInvented commitment or omitted unresolved issue
Field extractionExact values, currency and dates; missing values marked missingFabricated field or incorrect total
Small coding changeExisting tests pass and requested behavior is demonstratedBroken behavior, unsupported dependency or manual repair

Keep source material and requested output consistent. Record the exact model ID, prompt version, effort, available tools, completion limits and attempt count. Blind model names while checking outputs when possible. A trial should count towards the denominator even when it times out, refuses or needs a retry; record the reason rather than silently dropping it.

Summarize usable completions, factual errors, editing time, latency and billed usage. Agree on a task-specific acceptance threshold before choosing a winner. If neither model meets it, change the workflow or compare a different tier rather than treating a small score gap as proof of suitability.

This is an evaluation protocol, not a completed HaikuChat study. To try the current workspace, open HaikuChat and check its selected model and connection status. For the broader tier decision, see Haiku vs Sonnet vs Opus.

Haiku 5.5 benchmark questions

Where can I find the Haiku 5.5 benchmark scores?

The table above lists the publisher-reported October 7 results, including Terminal-Bench, OSWorld, Humanity's Last Exam and Chartography. Benchmark and benchmarks searches are covered here with the same source and review date.

Does Haiku 5.5 beat GPT-6 Luna and Sonnet 5.5?

In this release table, Haiku 5.5 is ahead of GPT-6 Luna in rows where both have a reported score, and behind Sonnet 5.5. Missing entries, effort differences and task fit limit that conclusion; it is not a universal ranking.

Are these independent HaikuChat benchmark tests?

No. These are attributed results from the model publisher. The local evaluation checklist is a suggested method, and no HaikuChat performance or speed measurement is claimed here.

Can I use these scores to compare Haiku 5.5 with Opus 5.5?

The source table has no Opus 5.5 column. A valid numerical comparison needs the same benchmark version, task set, harness and effort or budget details for both models. Different reports should not be silently joined.

Sources and review date

Specifications checked October 8, 2026. Prices, availability and evaluation settings can change. Published benchmark results are attributed to their source; we have not presented them as HaikuChat's own tests.

Use Haiku for your next small task

Chat, summarize pasted text or extract fields into a table. Check the workspace for current availability and task limits.

Continue comparing

All model comparisons →
HaikuChat

Sign in to HaikuChat

Continue with your Google account to use HaikuChat.

Your draft stays on this device. If you started by sending a message, its text and attached image are kept while you sign in. Existing device history is only saved to your account when you choose to import it.

By continuing, you agree to our terms and privacy policy.