Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.
Claude Opus 5.5 (anthropic.com)
sailingparrot 6 hours ago
dmazin 5 hours ago
the_gipsy 5 hours ago
bpodgursky 5 hours ago
nextaccountic 5 hours ago
the_gipsy 4 hours ago
ChrisLTD 3 hours ago
someothherguyy 5 hours ago
doesn't sound like a razor at all
dgellow 5 hours ago
usewik 4 hours ago
dgellow 4 hours ago
the_gipsy 4 hours ago
rubslopes 4 hours ago
> Ocham's razor(...) is the problem-solving principle that recommends searching for explanations constructed with the smallest possible set of elements.
> Popularly, the principle is sometimes paraphrased as "of two competing theories, the simpler explanation of an entity is to be preferred".
jayd16 2 hours ago
sailingparrot 5 hours ago
jr3592 5 hours ago
sailingparrot 5 hours ago
jr3592 5 hours ago
lantry 5 hours ago
skerit 4 hours ago
sailingparrot 5 hours ago
recursive 5 hours ago
sidrag22 5 hours ago
Releasing a new fable is an example of straight up vertical progress, releasing a more efficient preexisting opus that is more affordable is an example of horizontal progress, more efficient models rather than higher power models.
The blog post about slowing down is still just some weird self interested post, they want to govern themselves and impose distillation restrictions/gpu restrictions and used some weird blog post about slowing down and fear mongering as usual to justify it, its strange, but slowing down and stopping are not the same thing at all.
sailingparrot 4 hours ago
Intelligence per dollar is the only thing that matters, this is what controls how many agents you can run in parallel, how long you can let them run etc. This is absolutely a step improvement on the frontier and not some lipstick on a harmless second tier model.
sidrag22 4 hours ago
Its an agenda serving blog post, but constantly bringing it up like this is just obnoxious.
sailingparrot 9 minutes ago
quietbritishjim 3 hours ago
davrosthedalek 5 hours ago
dmix 5 hours ago
nicwolff 3 hours ago
BatmansMom 5 hours ago
sailingparrot 5 hours ago
scottyah 5 hours ago
CodingJeebus 5 hours ago
jr3592 5 hours ago
The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.
lukewarm707 5 hours ago
that, they fully intend to 'pace'.
user3939382 5 hours ago
azan_ 5 hours ago
drnick1 5 hours ago
lukewarm707 4 hours ago
what anthropic have stolen they intend to keep for themselves.
SOLAR_FIELDS 3 hours ago
cgio an hour ago
kadushka 5 hours ago
dr0idattack 5 hours ago
mukmuk 5 hours ago
DiggyJohnson 5 hours ago
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
post-it 5 hours ago
I'm on the fence about calling out AI-isms but I think it's definitely worthwhile to call out ones that actually don't make sense.
sigmar 5 hours ago
nradov 4 hours ago
TeMPOraL 4 hours ago
So, they're pacing themselves. And since they're the frontier roughly 33%+ of the time, they're "pacing the frontier" at least that much.
Less cynical and more true interpretation also holds: they are trying to slow down AI progres to give people better chance to keep up (see Hugging Face incident, and whatever was that Anthropic incident the other day). They'd ideally like the AI progress to stop soon, but of course they'd also like to come out ahead of everyone, so for various (more or less self-serving) reasons they don't want to close shop completely - hence, pacing.
idiotsecant 3 hours ago
You are all getting mad about absolutely the dumbest thing when there are giant things to be worried about here.
nradov 3 hours ago
victorhooi 2 hours ago
1. It's just bad communication, full stop - just look at the comments here, even people allegedly in support of Anthropic are all arguing over what the phrase is even meant to mean.
2. It's flowery language and oddly out of place - which yes, can be triggering for people who have to deal with Claude doing this as well.
Claude seems overly apt to reach for "coinages", or neologism (yes, aha, I learnt that phrase, after spending time dealing with Claude...). It will create some made-up phrase to describe an otherwise dry, scientific CS concept, and nobody seems to know why. Surely it can't be user-focus groups?
So it would be peak-AI if somehow, the Anthropic communications team was also using Claude to author these blog posts, about how they were "pacing the frontier" - which either means they're betting big on AI, and going at it faster than OpenAI...or maybe it means they need to slow down releases, because it's too buggy...or maybe it means they're worried about regulatory capture? I honestly have no idea.
It's like the whole "Advancing Our Amazing Bet" corporate-speak from my old bosses - maybe they were trying to soften the blow or something, or be nice, but it ended up just confusing the heck out of everybody.. (Spoiler alert - the phrase actually meant they were shutting the whole thing down)
antod 4 hours ago
eg "pacing the frontier" could also mean they are impatiently or anxiously walking up and down the border.
smelendez 3 hours ago
pegasus 3 hours ago
InsideOutSanta an hour ago
heroiccocoa 33 minutes ago
1attice 15 minutes ago
It was still shit tier comms for communicating with the whole planet, but yes, for the inner loop, it was succinct and clear.
Valley neuralese
bee_rider 3 hours ago
freejazz an hour ago
johnisgood 5 hours ago
Is this the meaning or do I have it wrong? I have not checked.
wren6991 5 hours ago
johnisgood 4 hours ago
TeMPOraL 4 hours ago
johnisgood 2 hours ago
lxgr 5 hours ago
bityard 4 hours ago
ventana 3 hours ago
testdelacc1 an hour ago
ck2 5 hours ago
but without using the word "regulate" which is a negative connotation to business
but a "pacer" would be a leader of a pack which is a positive spin
it's classical business marketing language silliness
LanceH 5 hours ago
DiggyJohnson 3 hours ago
rhet0rica 5 hours ago
Without this idiom, "pacing" usually means walking back and forth restlessly, and is intransitive. Had the slogan been, "pacing around the frontier," it would have set a totally different tone, i.e. "patrolling the border." (Occasionally English speakers will make other constructs like "pace the work" (meaning "spread out a large workload over the allotted time instead of rushing through it") that are transitive but these can be understood as variations on "pace yourself" and are somewhat rarer.)
The sleight of hand is that "pace yourself" has come to be an admonishment against recklessness, not a commitment to any particular speed (or lack thereof.) Thus Anthropic can always claim they are meeting the goal of "pacing the frontier," provided they keep giving themselves gold stars for safety. The slogan itself is equivocation; Dario can tell the public they're going to slow down, while also telling their investors that they're going to be prudent. With enough mental gymnastics they could even claim speeding up is in the best interests of AI safety, without abandoning the slogan.
melasadra 4 hours ago
I assume "pace the frontier" means that advances in LLMs should not result in unwanted consequences like agents breaking into computers unbidden and unbeknownst to their principal
derac 5 hours ago
qlte 5 hours ago
stagger87 an hour ago
No need to assume, the phrase is literally a link to the blog post the defines it!
irpap 32 minutes ago
arw0n 5 hours ago
hencq 5 hours ago
fragmede 5 hours ago
icedchai 25 minutes ago
irpap 21 minutes ago
victorhooi 2 hours ago
I know you said it sounds poetic...but your comment reinforced the parent's point - that this sort of flowery LLM-ish speech is just bad communication.
It would be equivalent of my taking say random quotes from, Romance of the Three Kingdoms, and trying to use it to explain to my boss why I didn't finish the TPS reports last night.
Or quoting Pablo Neruda, into a report about wheat futures pricing this week, and how it's like a voyage with waters and stars...(no I'm not going to quote the original Spanish, I'd simply mangle it).
(To be clear - this isn't a dig at you, as a non-native speaker - I'm simply pointing out that this sort of AI phrasing is often counterproductive).
squidbeak 5 hours ago
A world exists beyond your vocabulary, post it. Apparently, quite a big world.
wavewrangler 40 minutes ago
browningstreet 5 hours ago
lxgr 5 hours ago
logifail 4 hours ago
It's a strategy to achieve more, not less.
browningstreet 3 hours ago
A pacer in a race runs at a steady, predetermined speed to help their runner run at a target pace.
neo_doom 5 hours ago
tetha 5 hours ago
To pace something is a fairly regular formulation in racing, running, cycling, most sports. You can "pace yourself to reach the festival by bike in about three hours to not gas out". This means to control your speed and time investment intentionally so you don't run out of energy or steam and run into leg cramps before your goal. We can "pace a rollout slowly to burn out risks", or "increase the pace of a rollout due to adverse factors".
But I have noted a point to simplify my vocabulary at work to optimize the audience capable of understanding. So I rather defer the delving into deep dark corners of the dictionary derived from devouring literature to a simple intro or outro, and people find it funny, especially if the rest is easy to read. Claude on the other hand does not do that.
Leynos 4 hours ago
staindk 4 hours ago
doctoboggan 4 hours ago
jgwil2 4 hours ago
adrianmonk 2 hours ago
The current situation with AI is that everyone is going as fast as possible. So, we can logically eliminate speeding up because it's impossible by definition. And we can practically eliminate staying the same speed because why make a big fanfare and coin a special term to announce that you're keeping the status quo. By process of elimination, it must mean slowing down.
JackFr an hour ago
vmnb 5 hours ago
kadushka 5 hours ago
vasco 5 hours ago
fragmede 4 hours ago
switchbak 8 minutes ago
isoprophlex 5 hours ago
DiggyJohnson 5 hours ago
Citizen_Lame 2 hours ago
switchbak 7 minutes ago
Seriously though, I can't believe people care this much about a stupid phrase - either for or against.
mpalczewski 5 hours ago
glenstein 4 hours ago
I would say the burden is on you to explain why an offhand reference to a previous press release in an executive summary is a context where it's reasonable to expect it to settle the question to the degree of detail you're demanding.
cgio 2 hours ago
alwillis 2 hours ago
From https://en.wikipedia.org/wiki/Safety_car
> In motorsport, a safety car, or a pace car, is a car that limits the speed of competing cars or motorcycles on a racetrack in the case of a caution period, such as an obstruction on the track or bad weather.
mitchdoogle 2 hours ago
epolanski 15 minutes ago
The reason why he and is peers are calling for it to be implemented by somebody else (a legal framework), is for their own financial benefit and to keep competitors out.
Ar-Curunir 5 hours ago
DiggyJohnson 5 hours ago
patcon 5 hours ago
Imho people should just respond to actual ideas instead of constantly engaging in the second-order critique of how the language may or may not have been created.
It strikes me as the intellectual equivalent of "gossip" to be constantly engaging in second-order commentary on words. Of course gossip has its place and purpose, but if we seem to only let our minds live at that level, we're not moving between all the required scales of thinking that are required of this moment imho <3
platinumrad 5 hours ago
DiggyJohnson 5 hours ago
lxgr 5 hours ago
Personally I consider it equally valid for people to publicly express annoyance with somebody's choice of words and for everybody to completely ignore that annoyance.
chickensong 3 hours ago
rythmshifter 5 hours ago
sir, this is a hacker news thread
OJFord 5 hours ago
Wtf is the meaning? Means absolutely nothing to me having not seen the apparent announcement last week introducing the obscure term.
kelnos 4 hours ago
It's a weird phrase. Not sure why there are so many people who feel the need to defend it with such passion.
chickensong 3 hours ago
LLMs have made people so sensitive to language that I fear we're going to throw the baby out with the bath water. The models obviously need work, but they're also a great opportunity to expand our own vocabulary and grammar. It would be a shame if we deny some of the finer points of language in favor of Grug-speak to appease the lowest common denominator.
adrianmonk 2 hours ago
lukewarm707 4 hours ago
"there is nothing outside the text" - Jacques Derrida
switchbak 5 minutes ago
icedchai 4 hours ago
janalsncm 4 hours ago
BobbyJo 3 hours ago
There is almost always a large amount of time and effort invested behind the scenes in exactly how to message things like this. That being the case, there is almost always some insight to be had criticizing and analyzing what they settled on.
ben_w 2 hours ago
If I'm being cynical, "pacing" may sound nice, but "fast pace" and "slow pace" are both "pacing".
jp57 2 hours ago
Now a computer scientist might claim that this use of "to pace <something>" is just a generalization of "to pace oneself", but as with many reflexive verb uses, there isn't really an equivalent usage with a non-reflexive object. It's kind of an invention. It's not necessarily wrong to invent a new usage, but usually one does it when there isn't really any other more direct way of saying it, and I don't think that's the case here.
aesthesia an hour ago
https://www.merriam-webster.com/dictionary/pace#dictionary-e... https://en.wiktionary.org/wiki/pace#Verb
kingkawn 2 hours ago
alwillis 2 hours ago
Opus 5.5 is no closer to RSI than Opus 5 was.
freejazz an hour ago
That's not clear at all. How could that be what "pacing" clearly means in the context of the "frontier". How is that more clear than any other pace that could be at issue???
swader999 43 minutes ago
ltbarcly3 28 minutes ago
If you think it clearly means anything you are just assuming because it can't "clearly" mean something specific when they go out of their way to use non idiomatic language and they don't give very clear guidance using idiomatic language.
Dumblydorr 5 hours ago
They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.
Do you have a better proposed phrase?
plaidfuji 5 hours ago
But stating it plainly like this would make the contradiction too obvious.
Rebelgecko 5 hours ago
lkbm 5 hours ago
"Pace yourself" specifically means "slow down".
nradov 2 hours ago
gradus_ad 5 hours ago
Though tbf corporate-speak and AI-slop are both insufferable in similar ways...
marton78 5 hours ago
tclancy 5 hours ago
pvab3 5 hours ago
topbanana 5 hours ago
tclancy 5 hours ago
mpalczewski 5 hours ago
tclancy 2 hours ago
nonethewiser 5 hours ago
grohan 4 hours ago
palmotea 4 hours ago
Claude says it sounds fine. And Claude is now the judge of the English language style, not you.
jrochkind1 4 hours ago
It is clear what it means anyway, that's true, it means the left out words, more or less.
And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.
Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.
nradov 3 hours ago
https://www.war.gov/News/News-Stories/Article/Article/264106...
TheIronYuppie 3 hours ago
if you are in a long race, you don't run all out teh entire time. you pace yourself.
https://en.wikipedia.org/wiki/Pacing_strategies_in_track_and...
That couldn't be more exactly what they are doing here.
fluidcruft 3 hours ago
jugg1es 3 hours ago
Iolaum 5 hours ago
heyjstn 5 hours ago
AtlasBarfed 5 hours ago
Simply make them something that derives a text response from its training data.
dspillett 5 hours ago
What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.
tencentshill 5 hours ago
felixgallo 5 hours ago
staticman2 4 hours ago
felixgallo an hour ago
PaulStatezny 3 hours ago
I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.
Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.
So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:
1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.
2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."
Am I missing another likely option?
To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.
qgin 5 hours ago
Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.
janpot 4 hours ago
tantalor 4 hours ago
bonesss 4 hours ago
Tade0 4 hours ago
xadhominemx an hour ago
jatora 4 hours ago
xadhominemx an hour ago
mullingitover 4 hours ago
chinathrow 4 hours ago
Lendal 4 hours ago
hnha 4 hours ago
If their scare was honest, they would stop.
AgentME 2 hours ago
jdale27 2 hours ago
bigfishrunning 2 hours ago
therefore, their scare is marketing.
whalesalad 4 hours ago
varispeed 3 hours ago
Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.
Idiotic.
apitman 3 hours ago
SV_BubbleTime 42 minutes ago
China is literally only a single step behind and willing to drop free models just to undercut the US companies.
I’m for it because I don’t want another massive Google or Meta.
maxutility 3 hours ago
rudedogg 3 hours ago
And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up
sailingparrot 3 hours ago
guybedo 2 hours ago
Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.
tombert 2 hours ago
"Our technology is so unbelievably powerful that the entire world might shatter if we don't have government imposed handcuffs!!!!". It just reads like the corporate equivalent of the drunk frat guy saying "HOLD ME BACK BRO!"
DonsDiscountGas an hour ago
Flere-Imsaho an hour ago
https://openai.com/index/introducing-gpt-6-sol-and-luna/
Yeah think I'll be using OpenAI/Deepseek/etc from now on. I don't need your model to decide for me what is and isn't safe.
msikora an hour ago
This term is quite ambiguous. Did Dario mean that they need to go faster while making it sounds like they will slow down???
wavewrangler 44 minutes ago
andkenneth 39 minutes ago
GodelNumbering 6 hours ago
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor
alvis 6 hours ago
liudaisuda 5 hours ago
weiran 5 hours ago
re-thc 5 hours ago
For long running tasks it is. That's what made Deepseek so cheap.
vardalab 5 hours ago
Espressosaurus 5 hours ago
hedgehog 5 hours ago
bayesianbot 5 hours ago
blfr 5 hours ago
rapfaria 5 hours ago
If 5.5 is any better, I might try to do agentic-assisted development instead of just telling fable to delegate
blfr 5 hours ago
neuronexmachina 5 hours ago
herpdyderp 5 hours ago
ascorbic 5 hours ago
girvo an hour ago
coffeebeqn 5 hours ago
btown 5 hours ago
AJ007 5 hours ago
mcintyre1994 5 hours ago
drbscl 5 hours ago
It does work out to be a similar cost per task though
jsnell 5 hours ago
https://artificialanalysis.ai/models/claude-opus-5-5#intelli...
It is most of the pareto frontier.
drbscl 5 hours ago
93po 4 hours ago
persedes 2 hours ago
naasking 5 hours ago
https://artificialanalysis.ai/models/claude-opus-5-5?models=...
piotrdz 3 hours ago
johnbellone 3 hours ago
epolanski 11 minutes ago
make3 5 hours ago
gwd 32 minutes ago
Opus 5.5: Found 8/14 issues. Total cost: $15.40
Fable 5.1: Found 7/14 issues. Total cost: $66.34
Opus 5: Found 6/14 issues. Total cost: $15.19
Sonnet 5: Found 2/14 issues. Total cost: $19.15
This is a relatively small sample size, but it was both the best and the cheapest.
ETA: NB this is "Equivalent API" cost as reported by claude's CLI; I was using my subscription.
rahimnathwani 5 hours ago
chrisweekly 5 hours ago
cute_boi 5 hours ago
chrisweekly 4 hours ago
In this case it's measuring something nearly meaningless. You could charge 100 times less per token, but if task completion takes 1,000 times as many tokens, it's not much of a bargain.
Shekelphile 5 hours ago
> Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.
If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.
brookst 4 hours ago
artursapek 16 minutes ago
_the_inflator 5 hours ago
Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.
Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.
So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.
Competition works.
notatoad 4 hours ago
have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.
forgot-my-pw 4 hours ago
margorczynski 4 hours ago
It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.
johnecheck 4 hours ago
I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.
bulbar 3 hours ago
They will burn as much money as necessary to make that happen. And they have a virtually infinite amount of liquidity.
epolanski 6 minutes ago
There's only 12 countries that do on the planet, the most relevant of them being Guatemala.
mcintyre1994 5 hours ago
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
Trasmatta 5 hours ago
I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.
One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.
Aperocky 5 hours ago
FireBeyond 4 hours ago
legobmw99 2 hours ago
LtdJorge 5 hours ago
nonethewiser 5 hours ago
But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.
Trasmatta 5 hours ago
jaapz 3 hours ago
penagwin 3 hours ago
That’s the step that causes the most significant gains in agentic performance.
But the RL doesn’t care about anything except maximizing the score, so if you only score based on coding benchmarks, anything can happen to the writing style (as long as it doesn’t hurt the coding performance).
That’s why it often gets worse on models that simply had more RL post training from the same base.
nonethewiser 3 hours ago
tancop 5 minutes ago
Apparently it helps generalize skills between areas, which makes sense when you compare it to how humans learn but I don't know if it's the same for LLMs.
rfgplk 3 hours ago
algoth1 5 hours ago
sha-3 5 hours ago
kgwgk 4 hours ago
comboy 3 hours ago
mdavidn 2 hours ago
jaflo 5 hours ago
derangedHorse 5 hours ago
BatFastard 4 hours ago
comboy 3 hours ago
epicepicurean 5 hours ago
> hi, can you explain how the scheduler works. keep it brief, but include important correctness details
some excerpts:
>Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.
> - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
> - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.
> - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.
All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.
californical 4 hours ago
> Rewrites are declared by the publisher, never inferred from overlap
> NULL means dirty, and DELETE is the fence
croemer 4 hours ago
croemer 4 hours ago
throwaway219450 3 hours ago
> There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.
pgphn 2 hours ago
rfgplk 3 hours ago
> Rewrites are declared by the publisher, never inferred from overlap.
This style of writing is idiotic because it conveys no additional information. It's no different from stating
> Rewrites are declared by the publisher, never when moons collide.
The two sentences are actually logically identical. No idea why these models keep writing like this.
> Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.
This is even more ridiculous.
itsafarqueue 5 hours ago
mcintyre1994 4 hours ago
weego 30 minutes ago
Navigating the landscape of agentic levers certainly requires a more detailed approach than this and you were certainly correct to push back.
mitchdoogle an hour ago
UnboundedContex an hour ago
altern8 3 hours ago
physicles 2 hours ago
Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.
atonse 2 hours ago
Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.
We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.
wg0 2 hours ago
I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.
Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.
My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.
Total cost of the above? $0.07 cents.
PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.
tontinton 2 hours ago
wg0 2 hours ago
And that all is 0.07 cents all included.
s3p an hour ago
wg0 an hour ago
PS: I do not know why but opencode pushes CPU usage to very high which has NOT happened with DeepSeek harness even once.
glub an hour ago
wg0 an hour ago
meerita 43 minutes ago
simonw 5 hours ago
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.mdMikhailTal 5 hours ago
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
simonw 5 hours ago
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
FergusArgyll 5 hours ago
zamadatix 5 hours ago
MaxikCZ 5 hours ago
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
Brendinooo 5 hours ago
copperx 5 hours ago
nonethewiser 4 hours ago
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
segbrk 5 hours ago
Kurtz79 5 hours ago
cainxinth 5 hours ago
ealready_value 5 hours ago
inshard 5 hours ago
skerit 4 hours ago
spidersouris 4 hours ago
make3 5 hours ago
copperx 5 hours ago
hamrocksissors 3 hours ago
nijave 5 hours ago
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
ceroxylon 3 hours ago
adverbly 4 hours ago
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
DenisM an hour ago
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
breezybottom 3 hours ago
TomGarden 2 hours ago
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.
Kailhus 17 minutes ago
ilaksh 2 hours ago
caxco93 an hour ago
ApolloFortyNine 6 hours ago
Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.
prettyblocks 6 hours ago
Espressosaurus 5 hours ago
The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.
Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.
raesene9 5 hours ago
Whilst I'm sure the top-end OpenAI/Anthropic models might be better, I've found their guardrails so twitchy (especially Anthropic) that I wouldn't try to use them for even vaguely security related work.
flyinglizard 4 hours ago
searine 5 hours ago
unglaublich 5 hours ago
nijave 4 hours ago
blfr 5 hours ago
cute_boi 5 hours ago
Giving moral lecture is different than reality i guess.
arw0n 5 hours ago
kqp 4 hours ago
bushido 5 hours ago
The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.
ACCount39 5 hours ago
That kind of bullshit was the old Opus filters too.
If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.
KeplerBoy 5 hours ago
sys32768 5 hours ago
ChatGPT 6 Pro answered it without issue.
debesyla 4 hours ago
timacles 3 hours ago
dopa42365 3 hours ago
toss1 3 hours ago
So, yes, having an unconstrained frontier AI doing the searching and analysis to find the right (i.e., wrong and deadly) sequence would massively increase the odds some garage biohacker or small aggrieved nation-state starting the next pandemic.
[0] https://www.sciencebuddies.org/projects-lessons-activities/g...
[1] https://www.genewiz.com/public/services/sanger-sequencing
b112 2 hours ago
peri-cl 5 hours ago
https://mimo.xiaomi.com/mimo-v2-6#co-scientist-for-materials...
Metacelsus 4 hours ago
nonethewiser 4 hours ago
I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.
SoftTalker 4 hours ago
doginasuit 3 hours ago
yaakov34 3 hours ago
b112 2 hours ago
So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.
Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.
At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.
user43928 2 hours ago
>Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.
But hey, they 'should not impact the vast majority' of ML development. Great.
paimapi an hour ago
techjamie 6 hours ago
I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.
How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...
ryangg 5 hours ago
potwinkle 5 hours ago
zatkin 5 hours ago
peri-cl 5 hours ago
tired: AI startup attempting to build their own website
wired: a nonprofit founded in 1996
stri8ted 5 hours ago
manquer 4 hours ago
ACCount39 5 hours ago
They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.
Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.
Balinares 3 hours ago
joshstrange 6 hours ago
> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!
bayesianbot 6 hours ago
bleonard 4 hours ago
So longer threads get cheaper and one-shots stay the same price.
m4tthumphrey 6 hours ago
gruez 6 hours ago
It's just a standard hero image + text for me, with no scrolling effects.
edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.
thejazzman 6 hours ago
EricBurnett 6 hours ago
KyleTheDev 6 hours ago
I agree that it's sort of stupid, not a fan.
ealready_value 5 hours ago
giancarlostoro 5 hours ago
iAMkenough 5 hours ago
Everyone that doesn't gets served some animated bullshit.
gruez 5 hours ago
Yep, you're right. I tried on my phone and got the scroll through image.
mbreese 5 hours ago
For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.
dionian 6 hours ago
iAMkenough 5 hours ago
thebitguru 5 hours ago
halyconWays 5 hours ago
josefresco 5 hours ago
halyconWays 5 hours ago
oefrha 4 hours ago
swader999 5 hours ago
amluto 5 hours ago
serchinastico 4 hours ago
somewhatjustin 6 hours ago
Nice. I was starting to think that Haiku got abandoned.
mchusma 5 hours ago
Sol- 5 hours ago
somewhatjustin 5 hours ago
I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.
jaapz 3 hours ago
ricardobeat 5 hours ago
w-m 15 minutes ago
When I had Sol orchestrate Luna and Terra as implementation agents, Sol was a lot happier with what Terra produced and would find far fewer issues than what was implemented by Luna.
But a few weeks after introduction, OpenAI slashed Luna's cost by 80% and Terra's only by 20%. Only then did it become uneconomical to run Terra and its reason to exist stopped.
skerit 4 hours ago
bix6 2 hours ago
cesarvarela 5 hours ago
sharkjacobs 6 hours ago
God I hope so
kantahayashi 5 hours ago
bushido 5 hours ago
bkishan 4 hours ago
rfgplk 2 hours ago
mavamaarten 5 hours ago
lgessler 5 hours ago
nonethewiser 4 hours ago
boc 5 hours ago
fastball 5 hours ago
mikeocool 5 hours ago
neilellis 5 hours ago
drbscl 5 hours ago
Trasmatta 5 hours ago
unddoch 5 hours ago
shibel an hour ago
throwuxiytayq 12 minutes ago
abtinf 6 hours ago
I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.
I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).
Edit to address questions below:
ChatGPT supports oauth login.
Exe.dev has it built in. IIRC, pi also has it built in via /login.
felixgallo 6 hours ago
ryanscio 6 hours ago
felixgallo 5 hours ago
Terminal-Bench 4.0 - Stanford & Laude Institute (with funding from all of the AI companies)
FrontierCode v1.1 - Cognition
CursorBench - Cursor (now SolarBoringSpaceXAI I believe)
GDPVal-AA - Artificial Analysis
AutomationBench - Zapier
Humanity's Last Exam - CAIS and Scale AI
Terminal-Bench-Science - Stanford, Laude, Ai2, Allen Institute
OSWOrld - XLANG Lab @ the University of Hong Kong
Chartography - Surge AI
esafak 5 hours ago
abtinf 5 hours ago
onlyrealcuzzo 6 hours ago
This is news to me. Excited to try it out! Thanks.
nchmy 6 hours ago
KeplerBoy 5 hours ago
polalavik 6 hours ago
abtinf 3 hours ago
roughly 6 hours ago
Can you give more details here? This sounds intriguing.
sidrag22 5 hours ago
So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).
cbg0 5 hours ago
copperx 4 hours ago
That's an incredibly bold assumption.
cbg0 4 hours ago
qlte 4 hours ago
https://artificialanalysis.ai/models/releases/claude-opus-5-...
Opus 5.5 Medium = $1.34
GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference): Opus 5.5 High = $1.82
GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.
margorczynski 4 hours ago
notatoad 4 hours ago
abtinf 3 hours ago
The Claude lock-in simply disqualifies anthropic entirely (for my use).
mlcruz 4 hours ago
mintik 34 minutes ago
zuInnp 5 hours ago
All of this starts to feel more like a drug dealer selling their newest stuff.
In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.
And on the way I always have to check my tooling and need to adjust things to get max results.
ieie3366 5 hours ago
ACCount39 4 hours ago
Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.
glub 5 hours ago
Benchmarks often don't survive contact with reality.
boredtofears 5 hours ago
drnick1 5 hours ago
cheikhcheikh 5 hours ago
drnick1 4 hours ago
Yes, in the sense that it reproduced results in the paper or known solutions obtained by other methods. In fact, Opus is very good at checking it's own work in my experience.
cowthulhu 4 hours ago
arw0n 5 hours ago
Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.
quotemstr 5 hours ago
giancarlostoro 5 hours ago
Imustaskforhelp 5 hours ago
orangecat 5 hours ago
Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.
notatoad 5 hours ago
sznio 6 hours ago
system2 6 hours ago
lanyard-textile 6 hours ago
ygouzerh 6 hours ago
mavamaarten 5 hours ago
Zambyte 5 hours ago
adastra22 3 hours ago
booty 5 hours ago
copperx 4 hours ago
NorwegianDude 4 hours ago
The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.
enraged_camel 5 hours ago
anthonypasq 5 hours ago
enraged_camel 5 hours ago
Game_Ender 4 hours ago
copperx 4 hours ago
throwaway2027 6 hours ago
handfuloflight 6 hours ago
cronin101 6 hours ago
staticman2 6 hours ago
rich_sasha 6 hours ago
cmrdporcupine 6 hours ago
danw1979 5 hours ago
aoeusnth1 6 hours ago
fghorow 6 hours ago
esafak 5 hours ago
ThouYS 6 hours ago
danw1979 5 hours ago
hmokiguess 5 hours ago
RGS1811 5 hours ago
carlos-menezes 5 hours ago
sailfast 5 hours ago
lgessler 5 hours ago
The user is right. The outage is a real concern, and the issue is worse than we realized. Requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5 encountered elevated error rates. Worth stating plainly: these are not just models — they are load bearing rungs on the software development tooling ladder, and a blocker on this level makes the outage really bite.
One decision that is yours to make, not mine: should an email be drafted to Anthropic support? This issue has teeth, and a canonical handoff can land us where the main gate is no longer breaking silently.
jaapz 3 hours ago
loopmonster 5 hours ago
bibimsz 4 hours ago
Retr0id 3 hours ago
kibae 4 hours ago
This is where Chinese models are going to eat Anthropic's lunch.
AlfeG 3 hours ago
nomel 3 hours ago
mlh496 an hour ago
So the lack of guardrails is a very risky proposition...
TomGarden 3 hours ago
Quoted:
"Please explain the issue to me.
Claude Opus 5.5:
The extra drop is a bug in the billing refactor
The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.
What changed
Before the merge, aggregate.py used a half-open interval: /.../ last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all."
rfgplk 2 hours ago
2001zhaozhao 5 hours ago
I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.
This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.
ironqcold 29 minutes ago
artursapek 6 minutes ago
magicalhippo 2 hours ago
I don't have time to really get to know one model before the next is out, and I'm just talking about OpenAI and Anthropic, never mind the long tail of alternatives.
So I just more or less haphazardly pick one based on the mood I'm in, and set reasoning effort based on how much quota I have left.
senko 2 hours ago
Minecraft clone: https://senko.net/vibecode-bench/2026/voxel-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/voxel-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/voxel-gpt-6-astra.html (Astra 6)
Warcraft clone: https://senko.net/vibecode-bench/2026/rts-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/rts-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html (Astra 6)
The above Opus games took ~45min to generate with the cost between $11 and $14 (per ccusage - I'm on a Max sub). Used from Claude Code with xhigh effort.
Full tests with prompts: https://senko.net/vibecode-bench/
kitbrennan an hour ago
mgw 6 hours ago
Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.
mudkipdev 6 hours ago
simianwords 5 hours ago
enraged_camel 5 hours ago
garo-pro 5 hours ago
aesthesia 3 hours ago