Claude Haiku benchmarks

Understand what a reported score measures before choosing a model. These guides collect published results, preserve their test conditions, and explain where they do—and do not—help with everyday tasks.

Haiku 5.5 evaluation guide

Our current guide covers the release table for Haiku 5.5, Haiku 4.5, GPT-6 Luna, and Sonnet 5.5. It includes eight evaluation rows, missing results, different score units, and a way to test your own workload. Published scores are separate from HaikuChat service testing.

What to check in a benchmark

  • Match the model version, benchmark version, and test setup.
  • Check whether a number is a percentage or a score in another unit.
  • Look for tool use, reasoning effort, and results that were not reported.
  • Try the same task on your own material and count failed attempts.

Choosing between models? Read the model comparisons. To try a question, summary, or table, open HaikuChat.

HaikuChat

Sign in to HaikuChat

Continue with your Google account to use HaikuChat.

Your draft stays on this device. If you started by sending a message, its text and attached image are kept while you sign in. Existing device history is only saved to your account when you choose to import it.

By continuing, you agree to our terms and privacy policy.