Input
$0.10 / MTok for prompts up to 100,000 tokens
$0.50 / MTok for prompts over 100,000 tokens
Output
$0.50 / MTok for prompts up to 100,000 tokens
$2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.
In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])
I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
So even at the 1.5x/2x rate luna is still half the price of this. Weird pricing strategy from Anthropic. I'm sticking with Luna if I don't need a super smart model
You're judging purely by token cost I assume, not cost per completed task?
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
The cost per task from Artificial Analysis is roughly 3x higher at every reasoning effort level for Haiku than Luna. Sol 6.1 on medium has the same cost per task as Haiku with significantly higher intelligence. According to those numbers (which you should take with a grain of salt), from a pure cost vs intelligence standpoint, you're better off using Luna for economics and Sol for intelligence.
With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)
Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.
I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.
Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.
For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...
No? Don't use these lower end models to work on small well defined tasks with a frontier model orchestrating? This approach works very well for me and don't have any issue staying under 100k. I have no idea if it's cheaper but it does seem to be much faster for tasks like QA.
The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.
It's not that weird. Most companies considering paying Anthropic are probably not considering Chinese models as alternatives. Many don't even realize they exist.
The thing is: availability of near-SOTA cheap Chinese models is forcing OAI and Anthropic to bring prices down and offer more efficient models, instead of simply focusing on super expensive SOTA LLMs.
one only has to look at the providers like OpenRouter and OpenCode to see the massive swing since this summer ... or the fear mongering by Big Ai about open weights since
Not surprised by EU heading that way with how we here in 'murica are treating the rest of the world
at this point, there is no meaningful difference in day-to-day work
If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.
I said "most companies considering paying Anthropic", which is not the same as "most companies". I also don't agree that "nobody" is doing this; I have lots of anecdata suggesting otherwise. Maybe the majority of companies are using the easy option of Copilot or Gemini like you said, but it's nowhere near 99%.
Yes it is. Literally every company Ive spoken too about this only has MS teams co-pilot lol. Like all of them are using the most basic form of AI.
They dont even know about Anthropic and think its just "AI". By the way this includes one of the largest power systems design firms in the world, who helps build many datacenters... Uses only MS copilot.
Umm, just go look up the whole of software revenues globally and compare that to OpenAI. Its very small. Nobody is paying for SOTA models outside software engineering and ancillaries, except for some outliers fields like consulting. Then there's techies in every field who use it, but not companies themselves.
Companies as a whole, do not care about this technology, except tech companies and its adoption cant even be compared to CRMs. A company might buy a SaaS product with AI but most of them are not purchasing Anthropic subscriptions lol.
I have been settling on using Mimo 2.6 flash for discovery/research and glm 5.3 flash for implementation. So far I've been pretty happy with the results, but mimo can be quite slow. If speed is an issue, I just switch to glm 5.3 flash.
I haven't really tried it that much. I use GLM 5.3 a lot, especially with Coralbricks where the cache reads are free so I can let it think a lot without breaking the bank in the following turns. It's really good for bughunt, planning, and security work.
I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix.
Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.
Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...
That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.
Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.)
Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.
On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins.
They say it in the announcement: Haiku is basically useful to be called as a subagent. So you use Opus, and you want to investigate your production logs, instead of throwing them right away which is going to use a massive amount of token for mostly noise data, Opus asks Haiku to determine patterns to extract only the relevant logs and feed it back in Opus. The models are meant to be used together.
If that was the case, then Claude Code would use smaller models for subagents. It doesn't. The subagent always inherits the parent model unless you tell it specifically not to.
>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
Not really. We use Haiku 4.5 to turn users' natural language queries and requests into fairly complex structured specs for interior design and construction. It has near perfect accuracy.
There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.
Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.
You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.
Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".
If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.
If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful.
I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.
>If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue.
"less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.
it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.
it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.
It isn't an argument, and I never said it is better. It is an explanation for why the pricing break is relevant. Anyone can do the same things to take advantage of that pricing. Would you prefer I not explain a basic fact to someone who may not know what is possible?
The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful.
On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.
So what's your setup? Speaking as a "open terminal in repo, open Claude, say 'do this thing please '" kinda guy I'm interested in learning about these more advanced techniques.
I use oh-my-pi (omp.sh) - mostly with stock settings and skills but its a "fully loaded" harness that you don't really have to add anything to. I do change a couple of things: I set it to prefer subagents, and I enable rewind. It is crazy that rewind is not a default, it saves a TON of context - if the agent goes down a crazy path that burns a lot of tokens it can rewind to an earlier checkpoint with the exploration or bug hunting summary.
Presently I'm using sol-high for the default agent which does orchestration and a lot of smaller investigation and coding tasks itself. sol-max for planning and review. Luna-max for planned coding and general tasks. I also have a $10 minimax plan and use M3 for exploration and library roles, but I could probably be using Luna for that just as well and still only very rarely run into usage issues.
I don't use any plugins or skill libraries apart from Caveman and I'm not sure how useful that really is anymore so I'd start without it so you have a baseline to compare. I do think it reduces context usage a bit but I haven't measured it recently. Caveman also includes some team, agent & investigation skills - again they might be helping but I haven't re-evaluated since like 90 days ago.
Haiku doesn't seem most cost effective solution. Jev like classifier can work in a fairly smaller model which are 1/10 of the cost. Opus is SOTA so I get the use case for one being restricted to Anthropic ecosystem. I think the use case for Haiku is mainly for users using on their chat for pro subscribers and free users to maximize their quota.
9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam.
Similar ballpark to Luna in price, cost, and accuracy. These are very cheap models: $0.38 to answer 40 in-depth data analytics questions (compared to $15 for Opus 5.5 or $20 for Astra).
Overall very good at data analysis - handling all of the straightforward data analytics questions correctly. It fell short answering some of the questions that required some deeper statistical analysis like looking into other variables. In other words, it's not as persistent as other models in its analysis, which I think we'd expect from how they're positioning the model.
Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5. GPT-6.1 Sol got all answers correct, but was 10x more expensive.
Appreciate providing this benchmark. I am curious, do you think that as new models (from the same company) are released and you repeat the benchmark, they adapt/extend their training data to include your dataset, thus polluting the benchmark results? Would you be doing anything to combat this?
It's a valid concern. We're not releasing the dataset, the answers, or the full set of questions to help prevent this. At the same time, I like to share where it gets things wrong in a bit more detail, which involves sharing a bit of the exam. I expect that we'll create a new benchmark with a new dataset in 6 months.
I've considered this as well, as many others have. I tend to think that A: this is moreso a feature rather than a bug, and B: this is almost impossible to measure- it's a boogeyman imo.
In other words, combating this is impossible if it's happening. Measuring whether its happening is also not plausible, for small fish running bespoke benchmarks suites. I think we (consumers) have hit a pretty clear stride of; new model releases > some subset of evals/benchmarks are saturated > new, harder, more niche evals and benchmarks take their place. That doesnt seem unhealthy to me.
> Second, this week, we’ll roll out a new monthly API credit to all Max and Team subscribers for use on the Claude Platform. Max 5x users will get $100 in credits per month, Max 20x users will get $200, and Team subscribers will receive up to $500, pooled across their users
This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes
Not really, you have to fiddle with generating api keys and setting environment variables. Meanwhile with Anthropic it will just start charging you API prices for the tokens you are generating without even a single warning.
Of all the things I've had to deal with with migrating solutions and frameworks. Generating API keys and setting environment variables are the least time consuming of all of them.
Nah, it's pretty trivial to switch providers (especially with Claude's help, ha).
This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.
This is them sneaking in taking the Claude Agent SDK (claude -p) off of subscription plans through the back door along with a model release. They previously wanted to do this in June, but backpedaled after huge backlash:
Even after the June changes there was some allowance to use agent SDK on the pro plan. This will move me to codex tomorrow if agent SDK is really blocked on pro
Those mfers. I'm using this for work! I use my work teams plan with pi so I can do all kinds of custom workflows that I can't in Claude Code. Time to convince management I need OpenAI instead.
How credits work with `claude -p` is a common question we're seeing. we're updating the faq now to make sure it's more clear. Will share updated docs soon
I hope people notice again that this is happening this time around.
Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.
To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.
It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.
Yes. Previously the page was all about how they were going to start charging for Agent SDK use with a banner on the top saying that, actually, they weren't going to do that.
The monthly credit allocation hasn't rolled out yet either, so as of right now, we're effectively at the status quo. I'd expect the billing change to land once you can actually collect your Claude Console account.
*For now. If a company were to degrade something, it shouldn't be so obvious that the "goodwill" was just a reallocation. Just a good strategy. For example, it allows them to claim that they're "just going from 150% to 125% usage allowance, which is still more than 100%".
Just to be clear - claude -p use will now consume credits rather than count against your normal claude code usage? This seems like a significant downgrade.
The update they added here[1] yesterday makes it appear that usage is covered by your subscription by granting you credits. That's the way I'm reconciling the above comment with their update on the site.
Sorry that that help center article was misleading; we’ve updated it to clarify that `claude -p` has not been removed from subscriptions as part of this change!
That's super encouraging to hear, thank you for the clarification! I'd edit my original comment, but the edit window has unfortunately run out on it. Should be OK, though, since the help desk page is now explicit about it.
I've been using `claude -p` for personal use for a long time without issues. I do mean actual irregular, personal use though. Not using it to max out my quota, using it as main rather than sub or trying to be cheeky with it.
Are you able to speak about the degree to which non-Claude harness use is discouraged? Has Anthropic's disdain towards non-Claude Code harnesses settled down since the compute crunch + OpenClaw heydey?
More importantly, will my personal company account get banned for using a non-Claude Code harness?
Not sure if you're able to answer these questions, but I would appreciate it if you can.
I've been operating under the pretense that claude -p was removed ages ago in the OpenClaw era, it's back now? What a mess of contradictory information, hope yall can figure it out
This is massive. So on top of the regular usage, we now have USD 200,- to freely use via the API however we please, even resell? That is a statement, even knowing that inference does not cost them nearly as much as they charge, this is very developer-friendly. Does some minor de-risking for testing concepts. Terms seem to be reasonable [0].
Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.
Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.
Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
> You can paste in pngs and they'll convert them to svg with varying degrees of success.
I've found Opus to be REALLY good at converting images that are well-suited to being SVGs (decent resolution, sharp edges, clear color boundaries [though gradients do OK]) after a few rounds of back-and-forth. It'll do precise measurements to figure out curvature, the exact colors to use/gradient stepping, simplify complex paths, etc.
Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file.
Opus invented it, and wrote the player, and the songs.
My prompts were:
> I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact
> I am looking for music of the quality of the original secret of Monkey Island
And then later:
> Modify scrimshaw jukebox to add a copy-paste prompt that explains the music format, it should be shown at the bottom of the page below the readable instructions, the prompt should be designed to help any LLM tool compose music in the correct format. It should have a copy to clipboard button.
I have been using these models to make 3d models with build123d — they are great at it. I am never touching Fusion again. Astra is current a bit better than Opus 5.5.
I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all?
Then I think it is impressive! Are you able to share the prompts you used?
I mean Claude code sessions are all jsonl files that it can interrogate on its own. Get a new agent to capture how it was made and what the prompts were. No need for tedious /resume’ing and prompting.
I'm going to truncate a lot of this. But here is the root prompt that will get you a video.
Make a youtube poop video about having too many tabs that keep appearing faster than you can close them. Style it after the "too many cooks" viral youtube video that seemed to repeat over and over, getting worse and worse every time. You have ffmpeg and a plethora of programming languages at your disposal. Go completely nuts and make it as extensive and creative as you want. Render should be 1920x1080, and later 4k if you did a good job.
[discussions about acts, characters, darkening theme, etc.]
Workshopping the acts:
The play has a font issue. See screenshot.
[Image #1]
Also, can you change the thing that happens at the start of episode one the cursor clicking the one tab and it becomes two. Then it clicks another tab and it becomes four. Is that a breaking change? You can change the speed of the actions to fit it in if needed.
Here is a refinement prompt:
The last screen just before "The End", the cursor clicks on the browser window instead of the tab close button. Can you make it click on the tab itself? Here is a screenshot of where it clicks [Image #3]
Also, on the intro screen, the cursor clicks the tab and new tabs open, can you have it click the links in the page instead. See screenshot[Image #4].
You can keep the exact same timing for both of these.
Like most things, it depends. Think of the web devs of yore with MacBook Pros and Angular and React and Vim stickers stuck aside their glowing Apple logo (no longer a thing unfortunately). IMO having a pelican pin would be more of the same - advertising you are a developer of some kind. An attempt to bridge the gap to a stranger who might be into the same thing. The tech changes, the desire to signal things doesn't. The usual Exceptions(tm) apply of course.
I think models consider this composition as a meme at this point. They don’t generate a pelican on a bicycle, they generate “the pelican on the bicycle”. A well-known composition with the sun, the grass and occasionally a scarf.
I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)
I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
I thought the thinking effort was specified out of band from that, though maybe it's not. Not sure if the model was trained to listen in other areas. The biggest issue is, it's difficult to tell if it works because you can no longer see the thinking! Though I guess if you can't tell a difference in the output, was there any point to max in the first place?
I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)
Some of you might be but some of us appreciate joyous absurdity and look forward to the pelicans with great pleasure and anticipation. Chacun à son goût.
Every new model you make this post, every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not” it’s not a worthwhile conversation to have at this point, either the models can do some arbitrary thing or they can’t.
I'd turn that question around: why bother with the pelican tests at all?
Simon explained why he started them in the first place:
"I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."
That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.
I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.
I continue to do the test because I still learn something new from it every time.
This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.
Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.
I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.
I get a good initial intuition about how much they are going to cost, how much the reasoning efforts affect their output (and their duration and cost), and how much they have changed since the previous release in the same model family.
You get intuiton on capability which is contaminated by your blog post being ingested to training corpus, wihout strict comparision analysis on not pelicans, not animals, not svgs you cannot state any general capability with your prompting
The hn community is diverse. Of course some people will be tired of it but if enough people are engaging with the pelican test I would say that is likely because it still has some relevance.
My only concern is that it might actually decrease the quality of the output. So many are outputting the pelican leg on the "wrong" side of the frame (along with the left pedal if you were facing same way as the pelican) it may be reinforcing future models to draw it that way.
The benchmark is not saturated. They are in the training set in a way where the models have heard of it, but not in a way where all the models get a perfect score on it.
Yeah, I haven't been seeing the point of this exercise for a long time (apart from driving traffic to Simon's blog if I wanted to be negative) and find it boring by now.
But apparently we are in the minority since the pelicans are always upvoted to the top, so if others have fun with it then whatever, fair enough.
I too will repeat the same comment that appears on 100% of these threads, and say that actually I love seeing the Pelican and that at this point it has become a lovely tradition.
> every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not”
The answer is “they are definitely in the training set”.
They all generate the same composition and style because this has become a self-reinforcing loop. They seem to reproduce a specific solution from memory instead of designing from scratch.
Make this the new benchmark please. The prompt is funny, the result is even funnier. I think the great part is there isn’t even a right or subjective way to draw it.
Are you saying even these recent pelican pictures are not good? Vector graphics aren't meant to be photorealistic I don't think. For what they are they seem pretty good already to me. You can easily describe what is in the picture.
I meant “made up” as in truly invented - this was years ago now, they’d use a boilerplate sometimes with “everyone’s dealt with <x>”, but see what you mean.
So, just to confirm: you’re saying that it’s hallucinating “classic”, because you don’t think it’s in the training data, even though it correctly is a classic test?
Please show me your personal reasoning chain how you reached this conclusion because I think you’re hallucinating that the models are hallucinating the word “classic” here.
It's not a classic problem though. It's a very niche, novel benchmark that can't be more than 2-3 years old. That doesn't fit any definition of "classic"
I don't think any of those frames are "right". None of them has a seatpost, none of them has representative handlebars, most of them have some fatal flaws that would make the bike unrideable (like the head tube being angled the wrong way from vertical). But you are right that the first one is the only one that hallucinates entirely new tubes.
My complaint about Haiku 4.5 was that it was 10x the price of GPT-6 Luna.
> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens
Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.
So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.
Yep. I was looking at the prices of lower tier models a few weeks ago for zero/few shot tasks (pre Jev) and Haiku rates just didn't make sense at all. I ended up using 5.6-luna.
Good to know that is going back to being an actual option from perf/price perspective.
Gotta be careful about these benchmarks because they're extremely difficult questions that may be significantly more complex than questions you would ask of the model in real use.
If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this.
Agreed. I’m specifically responding to the statement that Haiku was higher on all benchmarks. Normalized by cost (and why wouldn’t you normalize by cost?), it was higher on 1/3.
At scale this make a difference. For me/us not that much.
At my work we currently don't have a subscription and use API based billing.
Some time ago the price for my usage was about 30-50€ per day when we used opus for about everything. With luna+sol its about 10€ for sol to plan or debug and 5€ for luna to implement.
At that point it doesn't quite matter to me if the cheap anthropic model is 20-50% more expensive compared to luna since luna is so unbelievable cheap.
The big gamechanger was the cheaper models being good enough to do the implementation given a good "rough" plan.
Nice that that's also possible with anthropics models.
It grinds my gears that we are still measuring $ per token, when the tokenizers for OAI and Anthropic are different. It is nonsense to compare tokens across labs, it only should be used to compare against a single lab's other models.
One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.
Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.
It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.
Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.
Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?
Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.
Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.
The monthly API credits for Max plan seems fantastic, especially considering Haiku pricing. Being able to actually use my Claude plan for other harnesses and use-cases on top of regular CC usage is everything I wanted.
Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.
yeah totally agree. esp how efficient it can be to have a subscription quota-paid orch spin up a bunch of API agents, this is kind of like free money to encourage what was already an easy way to save money (via batch pricing)
Note that this is Anthropic Trojan-Horsing the previously announced June change in with a model release, where the Claude Agent SDK can no longer be used with Claude subscriptions and is now billed with API credits only.
You could always use Claude models on other harnesses via API... just not via subscription. Now they give you $100 worth of API tokens to use on opencode or Pi. Which is better, but still not the same as OpenAI were you can use the subscription on Pi without problems.
About time Anthropic released a competitive cheap model. Haiku 4.5 has been too expensive compared to its performance for months now (in fact I don't remember being too impressed even when it was released). This one actually looks worth using in some scenarios. If it's really as much of a step up from Luna as the benchmarks they've shown indicate, it'll probably replace Luna in my workflows. 100k tokens is a pretty low threshold before the price goes up, but I tend to use these smaller models for smaller tasks anyway.
This is great! Been using GPT 6 Luna for decompiling my childhood favorite game (Age of Mythology) and this means I can throw Haiku into the mix as well. 17352/21965 functions matched so far...
And to setup a harness that will decompile the game and start doing a matching decompilation of every function. It set up a bunch of tooling and started a service in the background to do this actual decompilation campaign. I put some instructions into the main opus chat now and then to e.g. add automatic git pushing including a nice svg chart of progress and to switch model strategies here and there i.e. to do a first pass with a cheap model and then switch to opus/sol if the small model can't solve it.
I've done the same with some old games I used to play, SimTower and Oregon Trail II, both fully decompiled and now running natively on modern macOS.
I'm going to try create an interactive twitch stream where viewers can play the game through the stream and other non-player viewers can trigger events in the game via points.
Crazy time we live in.
Edit, you come to really understand the game in the process, and why things happen and how to better play the game.
And occasionally come across bugs, dev assets, assets never used, or assets all coded up, but code never triggered.
Haha yeah, I saw that post when I looked into reverse engineering SimTower, that guy saved me so much time decompiling. Didn't get so lucky with Oregon Trail II, there are some very old github repo's with attempts, but none got very far.
We're definitely in the age of ports! Interested to see how the OT2 port turns out.
Reverse engineering Redhook's Revenge binary (an old DOS game) before the advent of LLMs cost me way more hours than I'd care to admit back in the day - so I can't wait to put an LLM to work on some more obscure games like Sword Quest.
On a side note I should really give Oregon Trail II a shot. I never got into any of the successors like Yukon Trail, Amazon Trail, etc.
Surprisingly, no. I've been using Fable and Astra both to orchestrate decompilation of a relatively modern game (delivered via Steam) and they have no qualms about it.
Generally it's fair use if it's a legal copy and you have a legitimate reason such as backup/recovery or working around a bug and the models don't care.
If you stumble into vulnerability/cryptography territory sometimes they'll whine.
I want to get the original (Age of Mythology Gold Edition) running natively on macOS and then port it to WASM to run it on the web so I can easily play it with friends
Last time I gave that a try (without LLM assistance though) it was really hard as games DirectX calls cannot simply be glued to WebGL so performance was bad.
Probably. There have been dozens of examples of taking old games (Crazy Taxi, Super Monkey Ball, Quake, etc) and making WASM browser equivalents using AI to decompile them just on "Show HN" alone.
They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.
Definitely, TW can "cheat" a lot with using sprites/very low LOD at distance, only streaming the corresponding units on battle load, and it's turn-based so performance isn't critical. Although obviously you won't be doing those 50000 unit battles on the web.
Off-topic sidenote: what's with all these new projects targeting WASM instead of native, even if packaged for desktop anyway?
Probably would work although if you release/distribute/allow others to use it, you're treading into copyright infringement. Afaik personal use if you own a properly licensed copy is usually okay.
How do you validate the functions are correct? I did something similar, letting it (mostly deepseek 4.1) translate from assembly to C but it commonly made mistakes, some really hard to discover and fix.
If it's like any other matching (game) decomp, you compare the output byte-for-byte to the retail binary (e.g. any of the projects on https://decomp.dev/). Hard part to getting started is making sure you have a comparable compiler, flags, & toolchain
Their entire benchmark is deeply flawed and every model release they astroturf it (and I appear to call that out, but only after someone else independently verifies that it's bunk)
485 comments
In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])
I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.
From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.
Fixed.
The benchmark in the article showed it as lower per completed task than luna, but I guess we'll find out how representative that is. Anthropic has generally been fairly honest in their benchmarking though.
With that said, the real reason to use Haiku is that it's faster than all of these models. OpenRouter is showing an average so far of 93 tokens/sec, and AA got at least 137 in each of their benchmarks. So it might be valuable for speed at lower thinking levels. (At higher thinking levels, it's likely going to take longer to produce results than Sol on low/medium.)
https://artificialanalysis.ai/models/releases/comparisons/cl...
Neither encode nor decode are linear in compute, so providers need to price for average expected length.
This is just getting closer to the true cost of generating tokens.
For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.
These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.
In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.
Open weights models giving a distant salute from afar
the trend in industry is clear by now though
Not surprised by EU heading that way with how we here in 'murica are treating the rest of the world
at this point, there is no meaningful difference in day-to-day work
Doesn't this have more to do with LLMs getting more usage overall? I don't see how that's a sign that people are leaving OpenAI/Anthropic.
If you want to make it about 99% of real world companies, they are all on Gemini or Copilot anyway, nobody is going through legal and procurement to get models from dubious silicon valley startups when you have relations with Microsoft or Google or Amazon from ages because some benchmark is showing some minor digit benefit when vibe coding GTA 6.
They dont even know about Anthropic and think its just "AI". By the way this includes one of the largest power systems design firms in the world, who helps build many datacenters... Uses only MS copilot.
Companies as a whole, do not care about this technology, except tech companies and its adoption cant even be compared to CRMs. A company might buy a SaaS product with AI but most of them are not purchasing Anthropic subscriptions lol.
Maybe I should try the flash...
Even for simple tasks why use X if I know "Y Max" is available and on paper, better?
And why use something else when your favorite company releases something. Surely it must always be the best one to use.
And Opus 5.5 is really good.
AAI Index // Input // Output
Haiku 5.5: 43 // $0.10 // $0.50
Mimo 2.6 Pro: 46 // $0.43 // $0.87
Mimo 2.6 Flash: 38 // $0.10 // $0.28
Seems competitive to me? Plus then I don't have to manage multiple providers
If that was the case, then Claude Code would use smaller models for subagents. It doesn't. The subagent always inherits the parent model unless you tell it specifically not to.
There are plenty of workflows like translations where you'd easily be under the cap.
Your vibes don't appear to be supported by facts. From the announcement:
>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.
I'm asking to learn for a similar project, not to discount anything you're saying.
You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc
So, it is might be even worse.
For this application 100K token input is plenty.
Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.
For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.
I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.
I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.
"less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.
it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.
it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.
The services are priced this way because larger context has significantly higher costs. That is a fact about the technology and it is true for every provider. So moaning about it isn't useful.
On the other hand, there are a lot of people who don't manage context effectively - who start every session with 60K tokens - and that is significantly hurting the performance of every single thing they do with coding agents.
Presently I'm using sol-high for the default agent which does orchestration and a lot of smaller investigation and coding tasks itself. sol-max for planning and review. Luna-max for planned coding and general tasks. I also have a $10 minimax plan and use M3 for exploration and library roles, but I could probably be using Luna for that just as well and still only very rarely run into usage issues.
I don't use any plugins or skill libraries apart from Caveman and I'm not sure how useful that really is anymore so I'd start without it so you have a baseline to compare. I do think it reduces context usage a bit but I haven't measured it recently. Caveman also includes some team, agent & investigation skills - again they might be helping but I haven't re-evaluated since like 90 days ago.
If there was no Luna you would see only the >100k pricing, but because we have Luna, they had to lower price for something.
9x cheaper than Haiku 4.5 and 2 letter grades better. It's also now the fastest model (using the default speeds, not trying any of the other models "Fast" mode) to complete the exam.
Similar ballpark to Luna in price, cost, and accuracy. These are very cheap models: $0.38 to answer 40 in-depth data analytics questions (compared to $15 for Opus 5.5 or $20 for Astra).
Overall very good at data analysis - handling all of the straightforward data analytics questions correctly. It fell short answering some of the questions that required some deeper statistical analysis like looking into other variables. In other words, it's not as persistent as other models in its analysis, which I think we'd expect from how they're positioning the model.
Compared to OpenAI: GPT-6 Luna did a bit better and was about 30% the cost of Haiku 5.5. GPT-6.1 Sol got all answers correct, but was 10x more expensive.
In other words, combating this is impossible if it's happening. Measuring whether its happening is also not plausible, for small fish running bespoke benchmarks suites. I think we (consumers) have hit a pretty clear stride of; new model releases > some subset of evals/benchmarks are saturated > new, harder, more niche evals and benchmarks take their place. That doesnt seem unhealthy to me.
This is a very big benefit for me. I can now ship actual ai enhanced features behind my subscription without paying extra or fully relying on on-device models. I do worry that this is to soften the blow for user-unfriendly changes
That's 5-15 minutes of work at most. Not exactly the type of lock-in the parent is implying.
This is more to encourage people to try out adding AI into their product, which is a totally different flow and experience from using AI to build the product.
Anthropic isn't even close to being this useful.
Biggest loss is that Ant models look like they are genuinely better.
This changes on a weekly basis, I ended up with subscriptions to most of the providers (except for X.ai).
https://support.claude.com/en/articles/15036540-use-the-clau...
https://code.claude.com/docs/en/headless
So, as written, yes.
Update: We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...
Being forced through the non-OSS Claude Code with all of its quirks and issues is... such an exhausting use of force by Anthropic.
To the extent that you _can_ choose to disable telemetry and training on your traces in CC, it's not all that obvious what they gain by crippling your ability to use the subscription with other – better – tools.
It's also remarkable that it's coincident with OpenAI adding "Sign in with OpenAI", so that you can use your tokens with other tools.
You might be right and they will change this in the future, but that's speculative
This text has replaced the entirety of the page called "Use the Claude Agent SDK with your Claude plan."
What more do you need?
[1]: https://news.ycombinator.com/item?id=49999702
[1]: https://support.claude.com/en/articles/15036540-use-the-clau...
Anthropic is among the worst, their linux kernel hack is a prime example, didn't even count the "CVEs" to see Mythos can't count either
https://www.youtube.com/watch?v=NnV_cWeoo5Q
https://www.cbc.ca/news/canada/british-columbia/mother-jones...
ANT annoys me more with their ai psychosis on steroids
More importantly, will my personal company account get banned for using a non-Claude Code harness?
Not sure if you're able to answer these questions, but I would appreciate it if you can.
Of course, they don't do this out of pure kindness, but I really struggle to see a negative for subscribers already using a Claude Max subscription, especially given changing to another model is essentially frictionless via OpenRouter.
Compared with "Sign in via OpenAI" which they just announced, this is far less lock-in for anyone hosting services but less interesting for users of said services. With Anthropics approach, you can just use the allowance on your users however you see fit along with any other models and once it's used up, you can still just decide not to use their models for the remainder. With users bringing their tokens meanwhile, there is less flexibility in terms of switching for you, though might be cheaper for users.
Both interesting, each approaching this from a very different direction, each having their own trade-offs. On the OpenAI front, will be interesting whether developers can set specific temp, reasoning budgets, etc. for such "provided tokens" or whether OpenAI exposes that only via the actual API.
[0] https://www.anthropic.com/legal/credit-terms
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/I've found Opus to be REALLY good at converting images that are well-suited to being SVGs (decent resolution, sharp edges, clear color boundaries [though gradients do OK]) after a few rounds of back-and-forth. It'll do precise measurements to figure out curvature, the exact colors to use/gradient stepping, simplify complex paths, etc.
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
My prompts were:
> I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact
> I am looking for music of the quality of the original secret of Monkey Island
And then later:
> Modify scrimshaw jukebox to add a copy-paste prompt that explains the music format, it should be shown at the bottom of the page below the readable instructions, the prompt should be designed to help any LLM tool compose music in the correct format. It should have a copy to clipboard button.
https://claude.ai/share/1f721c20-2499-4d23-b368-3ab57146d956 and then https://claude.ai/code/session_01R3xuRtjVHqHo1GNbget6Tu
> Use your blender local skill to create a blender model of this faverge egg
The blender local skill is this one: https://github.com/simonw/gpt-6-astra-blender-pelican-bicycl... - which I described here: https://til.simonwillison.net/llms/blender-coding-agents-mac...
https://imgur.com/a/i7KAgB6
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
https://kq-styles.specr.net
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
https://www.youtube.com/watch?v=2EqMplbt0gU
Workshopping the acts:
Here is a refinement prompt:> "What's up with the pelican?"
Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...
The pelicans all start to look the same after a while.
But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.
Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
I'm definitely not the average person though.
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
Opus 5.5 had similar response on max: This is a classic test request
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)
That said... here's "Generate an SVG of an armadillo in fishnet tights jaywalking on Mars" on xhigh for comparison: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (and here's the same thing from other models: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...)
Every new model you make this post, every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not” it’s not a worthwhile conversation to have at this point, either the models can do some arbitrary thing or they can’t.
"I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."
That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.
I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.
https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle
This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.
Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.
I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.
Can you share in what ways these learnings affect your decisions or behaviors?
The hn community is diverse. Of course some people will be tired of it but if enough people are engaging with the pelican test I would say that is likely because it still has some relevance.
But apparently we are in the minority since the pelicans are always upvoted to the top, so if others have fun with it then whatever, fair enough.
The answer is “they are definitely in the training set”.
They all generate the same composition and style because this has become a self-reinforcing loop. They seem to reproduce a specific solution from memory instead of designing from scratch.
Please show me your personal reasoning chain how you reached this conclusion because I think you’re hallucinating that the models are hallucinating the word “classic” here.
That’s quite a big assumption from your end as well.
I think it's credible to call it a "classic" though, in a field that moves this fast. Name another comedy benchmark for LLMs with the same traction.
Two pelicans on a tandem?
They are training on your prompts.
> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens
Haiku and Luna now have the exact same price up to 100,000 tokens. Luna is now cheaper for anything after 100,000 tokens, even after Luna's own price increases at 270,000 it's still less than Haiku.
So it sounds like they've directly addressed that problem. Their self-reported benchmarks are all higher than Luna too.
Good to know that is going back to being an actual option from perf/price perspective.
In OSWorld 2.1 Haiku is better.
On GDPval-AA v2.1 Haiku is equal or worse than Luna.
On Humanity’s Last Exam they don’t seem even be comparing Haiku with Luna.
For these baby distillations of flagships, I expect their users to be very price sensitive.
If Haiku is "noticing" this and working harder to improve quality, you could still see similar or better cost-per-task in easier domains. ObviousBench is a good test of this.
At my work we currently don't have a subscription and use API based billing.
Some time ago the price for my usage was about 30-50€ per day when we used opus for about everything. With luna+sol its about 10€ for sol to plan or debug and 5€ for luna to implement.
At that point it doesn't quite matter to me if the cheap anthropic model is 20-50% more expensive compared to luna since luna is so unbelievable cheap. The big gamechanger was the cheaper models being good enough to do the implementation given a good "rough" plan.
Nice that that's also possible with anthropics models.
Haiku 5.5: https://html.non.io/lcars-haiku-5.5/
Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5
Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...
One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.
It's meant to be a good test, not a good design.
Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.
TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5
All results: https://jonclegg.github.io/pacman-bakeoff/
Hasn't this always been the case with Haiku?
Anthropic has really been doing all the right things in the past few weeks, while OpenAI continues to fumble the bag.
https://support.claude.com/en/articles/15036540-use-the-clau...
"You can still use the Claude Agent SDK, claude -p, and third-party apps with your subscription limits."
That's not a good sign for Conductor...
Update: We just updated the docs to clarify how API credits can be used w/ claude -p: https://support.claude.com/en/articles/15036540-use-the-clau...
Wait what? This has gotten their blessing?
I could now one-shot a new game, yeah.
I'm going to try create an interactive twitch stream where viewers can play the game through the stream and other non-player viewers can trigger events in the game via points.
Crazy time we live in.
Edit, you come to really understand the game in the process, and why things happen and how to better play the game. And occasionally come across bugs, dev assets, assets never used, or assets all coded up, but code never triggered.
https://news.ycombinator.com/item?id=49676394
Reverse engineering Redhook's Revenge binary (an old DOS game) before the advent of LLMs cost me way more hours than I'd care to admit back in the day - so I can't wait to put an LLM to work on some more obscure games like Sword Quest.
On a side note I should really give Oregon Trail II a shot. I never got into any of the successors like Yukon Trail, Amazon Trail, etc.
If you stumble into vulnerability/cryptography territory sometimes they'll whine.
Like could total war become a browser game?
They often ship the original assets in a somewhat brazen disregard for basic copyright law even when the games are still for sale on places like GOG though.
Off-topic sidenote: what's with all these new projects targeting WASM instead of native, even if packaged for desktop anyway?
My tests for Haiku 5.5: https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
Twice as expensive as Luna, but also considerably smarter too:
https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
https://aibenchy.com/compare/anthropic-claude-haiku-5-5-xhig...
haiku x-high: Cost $0.028, Tokens 54,579
luna high: Cost $0.004 Tokens 6,679
Luna was faster and cheaper, but the output didn't work.