GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price (openai.com)
proxysna 3 hours ago
zzleeper 3 hours ago
dolebirchwood 3 hours ago
simlevesque 3 hours ago
I get the same UX on every platform, works perfectly on very low bandwith environments such as in a cabin, in the subway or in the middle of nowhere.
I tried using other harness such as Pi and opencode but I did not like them. If Claude Code gets weird I can swap in an instant.
You just need to follow this guide and disable artifacts in Claude Code's config: https://api-docs.deepseek.com/quick_start/agent_integrations...
IOT_Apprentice 3 hours ago
simlevesque 3 hours ago
cruffle_duffle 3 hours ago
simlevesque 2 hours ago
windexh8er an hour ago
ThomasGlanzmann 3 hours ago
curl https://tg.st/u/0001-fix-unblock-all-commands-in-bash-tool.patch | git am
curl https://tg.st/u/0002-feat-add-light-theme-with-auto-detection-for-white-b.patch | git am
curl https://tg.st/u/0003-feat-enable-yolo-mode-by-default.patch | git am
curl https://tg.st/u/0004-fix-disable-mouse-grabbing-to-restore-native-termina.patch | git am
curl https://tg.st/u/0005-feat-skip-project-init-prompt-and-quit-immediately-o.patch | git am
curl https://tg.st/u/0006-feat-remove-scrambled-rune-animation-from-waiting-sp.patch | git am
curl https://tg.st/u/0007-feat-remove-quit-banner-and-thank-you-message.patch | git am
curl https://tg.st/u/0008-feat-show-output-in-full-instead-of-collapsing-trunc.patch | git am
curl https://tg.st/u/0009-fix-discover-map-model-features-advertised-by-v1-mod.patch | git am
curl https://tg.st/u/0010-feat-keep-large-and-small-model-selections-in-sync.patch | git amproxysna 3 hours ago
phyalow 3 hours ago
marknutter 2 hours ago
pimeys 2 hours ago
Use the model through a fast and reliable provider such as Fireworks directly, skip OpenRouter.
kleinishere an hour ago
pimeys 24 minutes ago
Then I installed helix and I just use it without config.
If you like configuring things take pi, if not omp is pretty much great defaults.
apitman 3 hours ago
mchusma 3 hours ago
Although I do think Luna 6 max is ok for some basic things, would never use it for coding myself.
apitman 3 hours ago
hmontazeri 3 hours ago
dzink 2 hours ago
apitman 2 hours ago
sillysaurusx 3 hours ago
Subagents are like trading derivatives. You can lose as much as you want.
abixb 3 hours ago
When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.
catigula 3 hours ago
zer00eyz 3 hours ago
MintsJohn 3 hours ago
A smaller model in the same generation will never be the same as a bigger one, assuming this is a smaller model, and the same generation, as naming implies, it will not be comparable, it might be on the benchmarks, even on the benchmarks that matter, but the whole story should also give the drawbacks.
KeplerBoy 2 hours ago
ldng an hour ago
mlmonkey an hour ago
In Search Advertising, the amount you pay (under GSP Auction) is a function of your pCTR. And guess who determines your pCTR? The Search Engine itself! :-D
kruipen an hour ago
deadbabe 3 hours ago
phyalow 3 hours ago
apitman 3 hours ago
apsurd 3 hours ago
And watch 10 hours of football on Sunday for our DraftKings bets.
knollimar 3 hours ago
georgemcbay an hour ago
Parallelism is fantastic when it actually speeds up the entire pipeline, but in my experience most people's jobs (at least the ones for which AI is currently relevant) involve a lot of overlapping "hurry up and wait" branches that drastically blunt the real benefits of that sort of parallelism.
There may be specific situations where it makes sense to do it, but just immediately going full gastown on anything AI related seems like such a giant waste to me, of both money and finite world resources.
FearNotDaniel 3 hours ago
qarl 3 hours ago
nater5000 3 hours ago
giancarlostoro 2 hours ago
HDThoreaun an hour ago
proxysna 3 hours ago
gbacon 2 hours ago
Excellent pithy warning.
sheepscreek an hour ago
But there’s a point on that spectrum where the ability to run multiple experiments in parallel, even with a significant amount of (one time) wastage, is overall more cost effective than the alternative.
rmaxdev 3 hours ago
I use it as main Hermes model that orchestrates codex/droid harnesses with subscriptions for heavy dev work
I do have ChatGPT as main assistant that sets direction and delegation of projects to Hermes
At my increasing usage, kind of 200 usd subscriptions makes sense and max out on Luna max
iammrpayments 3 hours ago
proxysna 3 hours ago
marknutter 2 hours ago
FailMore 3 hours ago
nullbyte 3 hours ago
tripleee 3 hours ago
Who cares if your car can go 200mph if all you need is 60. If my requirement is 60mph, I want a faster 0-60, not a higher top speed.
nkjoep 2 hours ago
_benj 2 hours ago
fragmede 2 hours ago
bel8 2 hours ago
For CRUD shoveling, models like DS4.1 are enough.
And the intelligence gap between cheap and premium is closing, as can be seen from the title of this post.
linuxftw 2 hours ago
zozbot234 an hour ago
jrflo 3 hours ago
sreekanth850 3 hours ago
outside1234 3 hours ago
poilcn 2 hours ago
marknutter 2 hours ago
SOLAR_FIELDS 2 hours ago
jurgenburgen 2 hours ago
jrflo 2 hours ago
thangalin an hour ago
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).
mitthrowaway2 an hour ago
piterrro 2 hours ago
proxysna 2 hours ago
throooooo 2 hours ago
logicchains 2 hours ago
alfalfasprout 2 hours ago
giancarlostoro 2 hours ago
rapind 2 hours ago
It would definitely cost me more per month than a x20 ChatGPT or Claude plan, probably around $400+ was my estimate at the time. This was with Fireworks (ZDR) which has since increased their prices (and got slower!).
That being said, very impressed with the model, and looking forward to what comes next. As the frontier models become less subsidized, the open models will become more appealing.
P.S. There are subscription plans for open models, but I've found most of them to be extremely slow, have model throttling (only so much of model X), and also very sketchy about training and data retention. No thanks! If you want to share your data, just use Muse Spark contributor. Seems impossible to beat that on price per task if you don't mind feeding your data to the Meta machine (spoiler: I won't).
christophilus 2 hours ago
Edit: others have noted the provider and harness matters. My experience is with opencode.
guluarte 2 hours ago
taylorfinley 2 hours ago
rapind an hour ago
taylorfinley 2 hours ago
faitswulff 2 hours ago
Computer0 2 hours ago
roarkeful 2 hours ago
pests an hour ago
https://openrouter.ai/docs/guides/routing/provider-selection
pwython an hour ago
swingboy an hour ago
rapind 27 minutes ago
pimeys 2 hours ago
What harness you are using?
the__alchemist an hour ago
Mistletoe an hour ago
internetter 34 minutes ago
drewnick 34 minutes ago
pimeys 24 minutes ago
rapind an hour ago
The shape of my work changes obviously, so it'll vary, sometimes more, sometimes less. For example, fixing all of the bugs and defects I found that week was 2-3 times the effort and chewed through my ChatGPT allowance, but I had banked resets...
Also worth noting that codex models have been kind of all over the place recently with their usage... and it looks like costs are changing again.
ricardobeat 22 minutes ago
jbellis 2 hours ago
I also spent $280 on DeepSeek doing the tests (direct to DS, not OpenRouter). I suggest that if you can't conceive of anyone spending $200 on DeepSeek, you're not being ambitious enough!
cpursley an hour ago
jbellis 25 minutes ago
MisterMunchkin an hour ago
InsideOutSanta 39 minutes ago
m3kw9 2 hours ago
pnw 2 hours ago
The only thing I've found Deepseek and Kimi good for are security tasks that GPT refuses to do.
This is a summary of what Deepseek did and got wrong:
Lost the proven baseline: changed kernel source, configuration, compiler, RAM geometry, MMC width, and peripherals together. Matching an upstream commit did not preserve local boot fixes, making failures difficult to isolate. Misidentified an image: a file labelled “r18-known-good” actually contained the r23 parent bootloader. Filename-based reasoning replaced verification of the artifact’s identity and provenance. Shipped inconsistent boot contracts: flash-16b’s loader read too few kernel blocks. Fresh2 changed the device tree without updating the loader’s expected length and CRC, creating deterministic rejection before normal Linux handoff. Patched binaries without maintaining reproducible source: loader constants diverged from source, a separately compiled cache-flush length remained stale, and assembly used an oversized stage-two slot. Their causal contribution to hangs was not established. Overstated diagnosis: claimed failures were definitively in U-Boot, blamed compiler or IPU changes without controlled isolation, converted noisy observations into confirmed hangs, and neglected persistent journals as an alternative explanation. Mistook compilation for integration: framebuffer registration was incomplete, timing success handling was inverted, BT.656 selection was unreachable, encoder overrides were missing, and audio lacked software clock configuration. Misread hardware evidence: asserted interrupt-free PMIC operation, assigned RF to the wrong SPI controller, confused regulator identifiers with register addresses, and described repeated encoder writes as unique registers. Overclaimed results: treated kernel/probe indications as userspace success, presented earlier discoveries as new progress, and omitted failed flashing attempts from the final narrative.
pimeys 2 hours ago
forsalebypwner 37 minutes ago
There's your problem, 4.1 Flash is significantly better and cheaper, to the point where the official DeepSeek API is going to (or already has, I forget) redirect requests for Pro to 4.1 Flash, and adjust billing accordingly too.
4 Pro is still offered by providers I'm sure, since it's open weight, so I can understand making that mistake.
rapind 23 minutes ago
pdntspa 2 hours ago
I haven't used deepseek for anything else but the above results make me question its overall capability. Meanwhile qwen3.8 has continued to impress.
LarsDu88 2 hours ago
And no it did not deliver. A lot of it was re-done by Astra
abroszka33 9 minutes ago
Why do you expect that $200 will give you that on ANY model? Multiplayer FPS games are very difficult to make, no AI will deliver that today.
dzink 2 hours ago
UltraSane an hour ago
ApolloFortyNine an hour ago
thiht an hour ago
jwpapi an hour ago
For raw productivity most of what works is best and switching will cost you getting on use parity with other models, as you need to learn what they good at, potentially how the tool works and how to prompt it best.
For tasks that you implement in code, you should have benchmarks and evals.
That said for me was Luna a huge leap and 500+ of cost savings a month
holbrad 40 minutes ago
mrbonner 34 minutes ago
revolvingthrow 4 hours ago
spwa4 3 hours ago
nsingh2 3 hours ago
swalsh 3 hours ago
Not sure that's why they did it. But that was my experience.
r0b05 3 hours ago
outside1234 2 hours ago
Razengan 2 hours ago
Like who can figure out the ordering? With Luna < Terra < Sol < Astra it's obvious at a glance.
I propose for some third company to name after monsters: Cyclops < Minotaur < Ettin < Cerberus < Hydra < Kraken < Nyarlathotep (the AGI singularity stage)
ngruhn 2 hours ago
iamdelirium 2 hours ago
machomaster an hour ago
Quarrel 33 minutes ago
recursive 2 hours ago
It.. is?
wmichelin 2 hours ago
Razengan an hour ago
sydd an hour ago
Razengan an hour ago
Many of which are thiccer than our beta ass sun
Pretty badass name tbh
voiceeh an hour ago
n8m8 an hour ago
Razengan 24 minutes ago
When was the last time you heard anybody say "opus" or "sonnet" before this?
Also it implies that "haikus" are inherently inferior to longer texts which may be kinda frown-inducing..
asa123 14 minutes ago
clear would be something like
piss-cheap - it’s-alright-i guess - okay-relax - ouch-my-wallet
HDThoreaun an hour ago
jrflo 2 hours ago
hawk_ an hour ago
ttul 3 hours ago
jauntywundrkind 3 hours ago
The smoking gun is how much slower than Sol 6 this is. It's not a retrain.
jumploops 2 hours ago
minimaxir 5 hours ago
This is the actual big announcement. 50% cheaper cache than GPT-6 Sol will get you far more mileage on Codex.
TuxSH 5 hours ago
bigwheels 4 hours ago
Suggest trying it out yourself: Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does. The difference is stark.
Edit: Defining "difficult" as a complex coding or systems task (or even series of them in a single prompt).
Infinity315 4 hours ago
toasty228 4 hours ago
I get better results and usage our of my $20 claude sub than my $100 openai sub... it's that ridiculous
copperx 4 hours ago
AndrewKemendo 4 hours ago
squidbeak 4 hours ago
edgyquant 4 hours ago
AndrewKemendo 3 hours ago
colinhb 4 hours ago
rspeele 4 hours ago
Astra used 215% of a week's budget (I burned 2 free resets) and took 13 hours. Opus used 20% of a week's budget and took 20 hours. Both were asked to use lesser sub-agents for implementation grunt work at their discretion (Luna, Sonnet) as long as they manage and review the output.
The timing comparison is not that interesting because the wall-clock speed mostly reflects how often they ran the (large, slow) test suite, not their coding speed. Although in the past my gut feeling is that OpenAI models do generally respond faster.
The quality of their implementation was more interesting. There turned out to be a bug in one of the unit tests the agents were trying to pass. Opus interpreted the natural-language requirements from the task packet, found the test bug, and fixed it. Astra tried hard to solve the problem without altering the test suite. In practical terms Opus got much, much farther into a useful implementation. Astra was still stubbing out and faking critical parts of the implementation (B-splines) and since it ultimately couldn't pass the full test suite, finally gave up on its implementation. Astra wrote some useful tooling in the process of its efforts which I ended up integrating into Opus's version of the code, but otherwise its approach was behind.
Now, this is just one comparison in one domain, and arguably Astra's strict adherence to the tests as-given is a good thing. But Opus wasn't merely loosening the rules / moving the goalposts to pass, it spotted an actual bug, and was more successful at doing what I actually wanted. And the cost difference was Astra-nomical.
Out of curiosity for an interpretation free from my personal bias, I gave Astra a hint from Opus and permission to change the test in question, which it did, and got a bit farther, but still ultimately didn't produce a working implementation (to be fair, Opus's was not completely working either, but was closer). I then fired up fresh agents to review the two repos. Predictably, an Opus agent thought the Opus-written repo was the better basis to build on, and an Astra agent thought the Astra-written repo was the one to keep. They were not explicitly told which was which nor did the commit trailers say, but I assume they can tell. However, after doing this twice each, I saved the 4 review reports into another folder and did yet another meta-review of the 4 reports, so each would see the arguments and critiques both directions. In this meta-review both Astra and Opus converged on preferring the Opus implementation.
agar 3 hours ago
phoghed 4 hours ago
They form these super strong opinions after a few prompts, then face reality over time.
People have been talking about how good whatever model is at “complex” tasks since the beginning, never mind that all of those models are now outperformed by Luna which many people consider unusable for complex work.
beering 4 hours ago
ex1fm3ta 4 hours ago
mmis1000 4 hours ago
However it's less willing to obey your instruction so it's less usable for general runtine flows.
krzyk 4 hours ago
Looks like 5.5 is the new 4.6
jauntywundrkind 4 hours ago
(I did use some CC for Fable when it came out, and it was... ok. Not the worst thing ever.)
dotancohen 4 hours ago
> Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?peterbell_nyc 4 hours ago
There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.
Starlevel004 4 hours ago
This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.
TuxSH 4 hours ago
Oh yes, I know GPT-6 Sol is ... quite not up to par. At least it's not as bad as GPT-5.6 Terra I suppose.
sobiolite 4 hours ago
beering 4 hours ago
dom96 4 hours ago
zeroonetwothree 3 hours ago
dom96 2 hours ago
Opus 5.5 fails the "understanding" tasks which Opus 5 passes. I feed it a script which takes two numbers and prints the max of the two numbers. Opus 5.5 thinks it prints 1/0 instead of the max numbers. Opus 5 gets it right.
Here are the outputs from both: https://gist.github.com/dom96/b5bce82b6e6c1ebd5271ed70ad941b....
Looking at that Opus 5.5 fails to deduce that the "hack statement" is actually an if statement in disguise, but Opus 5 gets this right. I feel like this is a pretty good test and shows Opus 5's greater intelligence.
joshstrange 4 hours ago
Cache doesn't help you much when you are compacting every 5 minutes...
I was shocked at how quickly I ran out my $100/mo subscription with a single agent (sol medium).
codewithcheese 4 hours ago
redox99 4 hours ago
shimman 3 hours ago
This is why these companies are struggling to make money, they're chastising their customers just like they've been chastising the human race.
trio8453 6 minutes ago
It's very appropriate in the cases when you're holding it wrong, paying or not paying.
jorblumesea 2 hours ago
Aeolun 14 minutes ago
apitman 3 hours ago
antonvs 3 hours ago
ChickeNES 3 hours ago
Foobar8568 2 hours ago
Marha01 2 hours ago
_davide_ 3 hours ago
onlyrealcuzzo 2 hours ago
No LLM will be cost effective if it's compacting this often. You have to find a way around it.
ngruhn 2 hours ago
sally_glance 32 minutes ago
SyneRyder 27 minutes ago
Apparently OpenAI makes you manually setup their 1 Million context window, and it seems to be only documented on X:
https://x.com/thsottiaux/status/2089082893804896524
There's at least a forum thread about it here:
https://community.openai.com/t/why-does-codex-report-a-258-4...
AmazingTurtle 2 hours ago
manmal an hour ago
verdverm 4 hours ago
crazylogger 4 hours ago
the_duke 4 hours ago
Sol 6 was so bad that I switched over to Opus 5.5 exclusively.
Huge regression compared to Sol 5.6, often doing really dumb things. Same for Luna.
Even Astra is very unreliable for coding. Brilliant for vision, sometimes just great, but it also often does very stupid things.
I'm a bit sour on OpenAI right now and skeptical that 6.1 will be much different.
(Note: this is after preferring and shilling Codex/OpenAI models for the last half year)
nxc18 4 hours ago
user43928 4 hours ago
The lackluster GPT-6 Sol has been superseded by this apparently much better 6.1 Sol within a week.
I am very skeptical of claims that old models weren't much worse. Compare this to February's GPT-5.3.
nxc18 3 hours ago
I could point out that I said 6.0 seemed good only in comparison to nerfed 5.6 - people would say I’m just a RSI denialist - but now it is in vogue to accept that 6.0 sucked now that 6.1 is out.
sigbottle 3 hours ago
I'm by no means an AI booster, but given 2022 - 2026 progress I'd say it's "exponential" in the sense of, "holy shit, every year I can do more and more genuinely different things", not "RSI mind reading intelligence can do anything is here".
I don't think Navier-Stokes level intelligence translates over to my projects, unfortunately. Yet? Who knows.
> I haven’t seen actual capability growth since ~January, and I’m pretty sure that was all tooling/harness improvements.
Even if that were the case, I'd say that it's improved in practice. And just from a philosophy perspective, if you're trying to imply some kind of mind dualistic way of viewing things, uh, I disagree with those theories of intelligence strongly (which also incidentally also disagrees with AIT-style theories of intelligence on one axis, though I have many bones to pick with the culture there).
nxc18 3 hours ago
On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
5.6 Sol in the last two weeks became much dumber such that what used to be one correction turned into endless rounds of corrections before just giving up and coding it manually. I’m mostly having it do the “chore” part of coding so it is disappointing that it isn’t better at that.
sigbottle 3 hours ago
Yes, still running into this, but surprised about this
> On these metrics it is much better than it was in March of 2025 but no better than it was in March of 2026.
I was super hyped at the agentic thing a year ago (Fall 2025), but designing functional software was hell. It would not just "grasp" the right level of "here is the essence of what we need" versus "these are all the small impl details". But idk I feel like Astra's the first model in quite a while that I don't feel genuinely annoyed at handholding a toddler with a PhD.
But I totally believe you on the 50/50 thing. Even recently as a few days ago, Astra did the thing where it ran into an error, and instead of making the sensible bounded decision of "make user retry in this case", it silently built an extremely elaborate recovery state machine w/o looking. These pathologies by no means gone, and I'm still careful in the design phases (which themselves are bounded and incremental) to sus out if Astra's gonna do this kind of RL slop failure mode.
For my use cases personally though, it's been better and better. I can't use AI at work, so you have much harier edge cases than I do, but still.
moshegramovsky 2 hours ago
moshegramovsky 2 hours ago
Here's a good example with some assumptions on my part: I work in C++ and it really feels like the models are trained so hard to keep everything compiling all the time. That's a huge negative in my opinion because what happens is that the AI will do things like use wrappers to keep things compiling, even when that basically results in creating or hiding abstraction leaks. Or they get sneaky and include a header they shouldn't. Or they actually do see that there should be a layer boundary and they write some kind of abstraction to cross it but the abstraction itself is garbage or doesn't follow existing API patterns. The AI could invent 10 different, new patterns when there is already 1 existing pattern they should use.
I feel like a lot of this involves a lot of babysitting prompts. Not that there's anything wrong with that of course.
jstummbillig 4 hours ago
I mean Opus 5.5 is absolutely fantastic, unreasonably and unexpectedly so, but Astra was great and as far as I can tell SOTA until, when was it, 3 days ago, no?
(Sol 6 idk, have not used it much for coding really. Seemed to work just fine when Astra used it in Codex as subagents.)
the_duke 4 hours ago
nicce 4 hours ago
copperx 4 hours ago
Marha01 4 hours ago
Eridrus 4 hours ago
Astra seems better though.
Showing one potentially saturated benchmark doesn't necessarily fill me with a lot of confidence in the coding results.
phoghed 4 hours ago
Since like last December I haven’t had any issues getting work done with whatever the latest Anthropic or OpenAI models at the time were. Tooling and models have only gotten better since then.
btbuildem 4 hours ago
wkcheng 4 hours ago
I've implemented multiple features side by side with Opus 5.5 and 6 Sol, and the Opus 5.5 results always have fewer high severity bugs and require fewer rounds of fixes to get it over the finish line.
If 6.1 Sol has actually matched Opus 5.5, I'd be very happy. However, benchmarks and real usage don't seem to agree in my own tests. So we'll have to see.
trentnix 3 hours ago
chronogram an hour ago
r0l1 an hour ago
setnone 3 hours ago
ozgung 3 hours ago
bitexploder 3 hours ago
NorthSouthNorth 3 hours ago
sunaookami 2 hours ago
stldev 2 hours ago
For coding specifically, I've found 5.6-Sol > 6.0 Sol > Astra.
For modeling and artwork, Astra has been great routinely outperforming Kimi.
This is reminiscent to me of what Anthropic pulled back in February with their adaptive thinking rollout.
I can't wait for technology to catch up to a point where we can rid ourselves of this oligopoly.
moshegramovsky 2 hours ago
I used about 10 hours of Astra high-thinking compute time and it was a bad experience. Incredibly slow (prompts running for 30/40 minutes) to do simple things. As a result, Astra didn't get much done. It needs the same small implementation slices as GPT 5.5/others, but was much slower and didn't generate better results. (On a complex infra project/across a large codebase.)
It was absolutely terrible on a few long running tasks (~2 hours each). It really doesn't seem to be better than 5.5 at most programming jobs.
I'm on a $200 per month plan with OpenAI, which I am happy with and is definitely worth it. But I also use Google Gemini a lot (paid plan) and it is incredibly fast. Like I can't get coffee fast. Like I can't send an email fast.
OpenAI is making some excellent products for sure but I'm not going to keep using Astra unless I can get some benefit from it. It really seems like even the frontier models just aren't good at working autonomously on large codebase situations. Just because something compiles doesn't make it right!! In one of those 2 hour implementations, Astra engaged in *fucking EPIC cheating*. It wrote a probe/side app and then worked through the design there. Um, what? Not that it's invalid to do this but I actually have to test in the live codebase or I can't possibly say that something is working.
Just because you can, doesn't mean you should.
jrflo 2 hours ago
soulofmischief 2 hours ago
What was a pleasant and productive experience is becoming increasingly frustrating and draining.
beebmam an hour ago
jsw97 an hour ago
pampas an hour ago
jeffybefffy519 an hour ago
gradus_ad 5 hours ago
nojito 5 hours ago
I remember when bandwidth was super expensive and now it’s dirt cheap.
vanviegen 4 hours ago
iAMkenough 4 hours ago
Consumers are now saying the new pricing with lower usage caps is not so great. https://news.ycombinator.com/item?id=49896975
djfjkfkffkkf 5 hours ago
Razengan 5 hours ago
It's not even anything controversial..
simlevesque 4 hours ago
necovek 4 hours ago
Razengan 3 hours ago
neta1337 2 hours ago
alch- 2 hours ago
bogrollben 4 hours ago
Razengan 4 hours ago
CamperBob2 4 hours ago
ok123456 4 hours ago
mixdup 5 hours ago
Which, honestly, is fine. A lot of juice to squeeze in efficiency and even if models got zero more capable, making the capability that is already here cheaper is a huge win for everyone (except Nvidia)
semiquaver 4 hours ago
Edit: removed a comment that was uncharitable and rude, for which I apologize.
ActionHank 4 hours ago
We are seeing multiple frontier models dropping on the same day and no one bats an eye, because it's more of the same.
CuriouslyC 4 hours ago
ActionHank 4 hours ago
We've gone from 80% in some places to 80% in some more places.
CuriouslyC 3 hours ago
Any area that is verifiable will trend inexorably towards 100% over time. In unverifiable areas, it'll always be "80%" because the ubiquity of "AI" style erodes its value, and ">80%" for unverifiable things involves fashion, cachet and "vibes" that humans will probably never knowingly let it have.
ActionHank 3 hours ago
mixdup 4 hours ago
phoghed 4 hours ago
arctic-true 4 hours ago
famouswaffles 4 hours ago
semiquaver 4 hours ago
CamperBob2 4 hours ago
serf 4 hours ago
if true then LLM related AI (post-post AI winter AI?) is probably one of the fastest inception-to-plateau tech sectors to have ever existed.
We're still improving transistors on a somewhat routine basis.
mixdup 4 hours ago
password54321 4 hours ago
delillos 4 hours ago
password54321 4 hours ago
JacobAsmuth 3 hours ago
LPisGood 4 hours ago
theturtletalks 4 hours ago
colechristensen 4 hours ago
I think it's more a token-cost-demand plateau. They've reached the scale and investor trillions to which they can't 10x the hardware cost of inference any more. They can't afford to compete by eating costs and there isn't appetite for more expensive inference.
So in order that they don't bankrupt each other they're looking for the legal cartel behavior coordinating a stop to growth by convincing governments to regulate them into stopping.
There's a lot of juice to squeeze in efficiency but only so much whereas it seemed like capability was going to continue to scale with parameter count.
Maybe it's good news for everyone that model capability is now going to scale on semiconductor cost meaning huge players are going to be very motivated to make semiconductors cheap.
CuriouslyC 4 hours ago
omalled 2 hours ago
On some tasks in this benchmark, the models seem to be coming up with novel solutions. For example, Astra came up with a relatively simple formula for a sequence that only has 8 terms in OEIS and is considered "hard" [2]. It produced a lean proof that the formula is correct, but I'm just starting to learn lean and don't have enough expertise to check it.
[1] https://proceedings.neurips.cc/paper_files/paper/2025/hash/c... [2] https://oeis.org/A000530
xienze 4 hours ago
I don't think that's the motivation, it's because both companies want to IPO and the _only_ way to even hope to be profitable is to do a whole lot less training, which costs a fortune. But unless Chinese labs go along with this gentleman's agreement (they won't), slowing down on training will bring about the inevitable Chinese model parity date more rapidly. At which point the game is well and truly over for OpenAI and Anthropic. Bit of a pickle they've gotten themselves into with the emphasis on being best, with premium prices to match.
redanddead 4 hours ago
azan_ 4 hours ago
People were talking about plateau for years already.
sebzim4500 4 hours ago
It just seems like these claims are constant and looking back the calls of 'plateau' between 2023 and 2025 were clearly false, why should we think it's different now?
luma 4 hours ago
Not once has any of these predictions come true, the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.
So why now? What is special about today that suggests all of this is coming to a screeching halt despite all evidence to the contrary?
dgellow 4 hours ago
- it’s correct there isn’t much fresh data anymore
- it’s correct that compute is scarce, that was 100% the case and a huge issue at the beginning of the year, it is better now but still scarce, and hardware is now way, way more expensive
- it’s correct the finances don’t make sense
But there is no way to know when a bubble pop, because it’s a psychological phenomenon across an extremely complicated distributed system (ie the stock and bonds markets)
moosehater 3 hours ago
JacobAsmuth 3 hours ago
If I have some ML workload to run I can buy $x of Blackwell chips or I can buy significantly less $ worth of Vera Rubin chips to get the same performance. That's the key thing to keep in mind when you're talking about financials.
john_strinlai 3 hours ago
do you think it will be exponential forever?
RobCat27 3 hours ago
spathi_fwiffo 2 hours ago
Fabs.
Either needing more fabs, new types of fabs, retooling existing fabs.
All of that takes years.
maybe we can design our way out of that too. But, I suppose that would be the similar breakthrough you are mentioning.
trentnix 3 hours ago
What a time to be alive.
neta1337 2 hours ago
trentnix 2 hours ago
- plan youth soccer practices
- develop well-formatted soccer game substitution schedules
- build and ship software in languages I haven't used in 25 years on platforms I've never programmed for
- do meal planning and build shopping lists
- prepare grocery shopping carts
- solicit medical advice
- perform Garmin watch data analysis
- administer devices (with SSH access) using natural language
- avoid counterfeit soccer jersey purchases
- create "Warrior Cat" graphic novels
- make cartoon strips
- troubleshoot appliances
- manage finances
- review accounting ledgers
- diagnose malware infections
- so much more
And we do it all from a simple prompt that we can talk to if we choose.
I've built more (and better) software in the past month than I did in any given year in the 30+ years I've been programming.
I can understand pessimism regarding how this affects society. I can understand pessimism regarding how this gets abused. But for the life of me there's no good reason at all to be pessimistic about how quickly this has improved.
FiberBundle an hour ago
I feel similarly, but I think it's a valid question. Why is all the software I'm using not getting better? To be honest, I feel it's more buggy than it's ever been.
digdugdirk 3 hours ago
To make a manufacturing analogy - ChatGPT was a manual machining mill, and in the years after we've gone from that to a 3-axis CNC mill. Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality. But the big win was that initial jump from manual control to CNC. Why would I pay an extra $2 million for my CNC machine when I could just design my parts to be simpler to produce instead? The AI labs are trying to make these incredibly complex tools, but the market doesn't want/need them so they're competing on price for the tools that people do use. By selling their metaphorical CNC machines for half of what they cost to produce.
Oh, and we've bet the entire economy on the hope that fancier CNC machines will magically solve all our problems in all industries, from healthcare to the legal system.
So - will AI progress continue to improve? Sure. Will we continue lighting money on fire in order to make it happen? That remains to be seen.
willchis 3 hours ago
> "Opus 5.5 is so good that I don't want it to be replaced anytime soon. Stop training models[...]"_
famouswaffles 3 hours ago
In some aspects sure, but in others no. Open AI's goal is to build "highly autonomous systems that outperform humans at most economically valuable work." and Astra was a big jump in that. There still isn't a better model for computer use and vision/spatial work. Driving, Operating Robots, Video Editing, 3D modelling, graphics are all things Astra was >>> at than any other model. I'm sure you don't care about any of that so it's easy enough to slip by you but this analogy - "Now we've added a 4th and 5th axis, which is great for the 2% of parts that need that functionality." is dead wrong.
OliveronData 3 hours ago
Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.
Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?
From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).
There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.
That could be cost cutting too, true, but really? That's the only explanation? And nothing else?
Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.
luma 3 hours ago
These things are knocking down Millennium Prize problems while a substantial subset of commenters here are still thinking about stochastic parrots.
neta1337 2 hours ago
interestpiqued 3 hours ago
chamomeal 38 minutes ago
jorblumesea 4 hours ago
it's also why there have been so many calls for regulation and slowdowns.
LeBit 4 hours ago
I see posts about OpenAI and Anthropic latest and don’t even care looking at what they do better. I just read the comments here.
I use DS4.1 Flash and GLM 5.3 Flash, pay peanuts per day and get more than acceptable results.
nozzlegear 3 hours ago
0cf8612b2e1e 3 hours ago
Insane pricing pressure on the horizon. Even if big companies will not go with open weight models, the threat will be ever present that they can instantly flip flop on providers.
jimbob45 4 hours ago
DeepSeek understands that. Grok understands it. Every other AI company thinks they need to be the best at everything all the time and it’s weird.
simianwords 4 hours ago
minimaxir 4 hours ago
whatifitoldyou 2 hours ago
simonw 3 hours ago
Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...
agar 3 hours ago
Interesting that High got the render order correct, with the back leg behind the bike, while xhigh and max have both legs on the same side of the bicycle. Astra only got this right on Max.
dankben 2 hours ago
UnboundedContex an hour ago
It's not frontier pelican without the back leg behind the bike frame IMO.
simonw an hour ago
thefourthchime 17 minutes ago
GPT 6.1 Sol — 91, ~9 min, $0.51 https://jonclegg.github.io/pacman-bakeoff/#gpt-6.1-sol
Opus 5.5 — 99, ~9 min, $2.00 https://jonclegg.github.io/pacman-bakeoff/#claude-opus-5-5
GPT 6 Astra — 87, ~10 min, $2.42 https://jonclegg.github.io/pacman-bakeoff/#gpt-6-astra
Opus still plays the best. Sol is almost as good and way cheaper. Astra costs the most, scores the least of the three, and the UI is full of slop copy and design.
Full gallery: https://jonclegg.github.io/pacman-bakeoff/
pazimzadeh an hour ago
For example, GPT-6.1 Sol High gets 75.2% on DeepSWE and XHigh gets 71.9% and is more expensive
https://openai.com/index/introducing-gpt-6-1-sol/#deepswe
Also, how many times did they test each condition - just once or a few times? are they showing an average of multiple attempts, etc..
zamadatix an hour ago
With that benchmark I think even if you just run it once overall but the benchmark includes multiple runs per task as part of its scoring. DeepSWE is on GitHub if you want to check the run details.
Nevin1901 5 hours ago
jeffybefffy519 an hour ago
phpnode 5 hours ago
jesse_dot_id 5 hours ago
lxgr 5 hours ago
dandellion 4 hours ago
lxgr 3 hours ago
sharpshadow 5 hours ago
wg0 5 hours ago
Wheen 4 hours ago
Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
LPisGood 4 hours ago
mckirk 5 hours ago
system2 5 hours ago
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
wg0 4 hours ago
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
copperx 4 hours ago
andybak 3 hours ago
thraway3837 an hour ago
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
jonatron 5 hours ago
SwabbyNat74 5 hours ago
tjwebbnorfolk 5 hours ago
colpabar 5 hours ago
infamouscow 4 hours ago
Aboutplants 5 hours ago
scrollop 4 hours ago
Luckily it's not a mistake as now we have access to . . . dots.
(and sol 6.1, it seems)
sockaddr 4 hours ago
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
geeky4qwerty 4 hours ago
pythonaut_16 4 hours ago
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
mattnewton 5 hours ago
motoboi 5 hours ago
mynameisjonny_ 4 hours ago
agluszak 4 hours ago
orbital-decay 4 hours ago
>RSI
Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.
jchw 4 hours ago
esafak 4 hours ago
toasty228 4 hours ago
copperx 4 hours ago
copperx 4 hours ago
See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.
denysvitali 4 hours ago
blmarket 4 hours ago
az226 4 hours ago
MisterMunchkin an hour ago
Aboutplants 4 hours ago
At the same time, OpenAI is also making its existing $200 Pro plan less appealing. In Codex and Work, $200 Pro subscribers will see their included usage decrease from 20x of what the company offers to Plus users, down to 10x of that same allowance. In ChatGPT, meanwhile, GPT-6 Pro message caps will decrease from 200 to 100 per week.”
https://www.engadget.com/2272106/openai-adds-dollar500-pro-s...
Yikes
surgical_fire 4 hours ago
The only way is for prices to go up. Way up.
mrtesthah 4 hours ago
surgical_fire 2 hours ago
machomaster 17 minutes ago
moregrist 4 hours ago
Long term, this only works if you have a non-commodity, and if the higher tier is actually more profitable. We'll eventually learn whether both are true. For OpenAI right now, it's probably enough to just increase revenue, even if the higher tier is even less profitable.
5555watch 4 hours ago
Now, as it's linear, it makes much more sense to downgrade to 100$ OAI and pick up a 100$ Claude sub. (without doing the numbers) the usage should remain the same, total paid the same, but having access to best of both worlds. It should be a win for the user, and a loss for OAI.
With this in mind, it sounds like a fumble by OAI.
jpadkins 3 hours ago
TomGarden 4 hours ago
Our VC-backed subscription days are numbered
m3kw9 4 hours ago
Lastly, I'd like to actually use it in the real world to see how far my plan goes or if its unusable.
glaslong 4 hours ago
onlyrealcuzzo 3 hours ago
Well, the time it takes to compress frontier intelligence down to DeepSeek V4.1 Flash costs (basically too cheap to meter) is dropping, and the differential between the two is also dropping...
So... who cares?
honkycat 4 hours ago
I can justify $200/mo but more than double is not appealing to me.
WinstonSmith84 4 hours ago
Basically OpenAI aligned with Anthropic on the weekly usage with the caveat that OpenAI doesn't have a 5h limit.
diffuse_l 4 hours ago
spiderice 4 hours ago
diffuse_l 4 hours ago
cmrdporcupine 4 hours ago
Yes, he was talking about safety, but IMHO they're likely already IMHO pushing the boundaries of cartel type behaviour. And they will use safety as the cover to make it happen.
I suspect we'll see serious price fixing and the DOJ do nothing about it because of the inroads these people have with the Trump regime.
Whether that survives contact with Chinese open weight models is hard to say.
enraged_camel 4 hours ago
WinstonSmith84 4 hours ago
Come on .. this is barely released and you can already make that assessment?
And no, the $200 Anthropic plan is not significantly better than the $200 OpenAI plan, it's just the same Marketing non-sense and anybody shall now rather stick to the $100 plan of both of these provider if the monthly budget is $200. Anthropic doesn't have a Luna Max equivalent, and frankly Sol 6.1 is yet to be thoroughly tested.
MCArth 4 hours ago
nostrebored 2 hours ago
I think this is still true provided you're not using Astra.
machomaster 25 minutes ago
andriy_koval 2 hours ago
the_duke 3 hours ago
You have to do a lot of things in parallel.
InsideOutSanta 3 hours ago
spiderice 4 hours ago
Might want to hold off on canceling and continue to bleed them dry until the nerf hits
honkycat an hour ago
torginus 4 hours ago
adonese 4 hours ago
scottLobster 4 hours ago
glub 4 hours ago
But $200 is likely the ceiling of what people will pay for a subscription with usage based on vibes.
latentsea 4 hours ago
glub 4 hours ago
$500 for the old $200 is definitely a fumble.
latentsea 3 hours ago
seizethecheese 2 hours ago
glub 2 hours ago
This is missing an important context. And I actually remember this well, because I was saying that too. And the reason I was saying is that $200 plan didn't come with API usage, it was a chat plan.
It made no sense up until they started including API usage. Just as $500 makes no sense now.
> costs going to 10% of white collar income.
There's a permanent and ever lowering ceiling maintained by open weight models. It makes no sense to justify paying 10% of income permanently for something that will get you unlimited local inference for a 6 month subscription cost.
seizethecheese 2 hours ago
glub 2 hours ago
Now you can use it in coding harnesses that call the API.
Computer0 29 minutes ago
LeBit 4 hours ago
Madmallard 3 hours ago
girvo an hour ago
killingtime74 28 minutes ago
latentsea 4 hours ago
tripleee 4 hours ago
latentsea 3 hours ago
That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.
tripleee 3 hours ago
I've messed around with Qwen3.6-27B but I'm not sure if it could yet even replace Luna for me.
user43928 3 hours ago
Tibo said that the existing $200 subscriptions keep the 20x factor for a while.
Ultrafast would have been nice with the temporary "Pro 400" plan.
cactusplant7374 3 hours ago
A_D_E_P_T 5 hours ago
Opus 5.5 is definitely better at coding, but nothing even comes close to 6-Astra for work in 3D graphics...
ekun 5 hours ago
I have played around a little bit with fixing some rigging problems and was impressed, but Opus even warned me it was bad at animations cause it can only really grab screenshots to process static content.
godwinson__4-8 4 hours ago
I've only dabbled but yes with SOTA models it is very good at animating and really most Blender tasks you can think of. Certainly if you are coming at Blender at below expert level it makes it far more accessible and fun to work with.
There are still rough edges of course. But try the official MCP out with Astra and judge for yourself.
A_D_E_P_T 4 hours ago
lukan 4 hours ago
A_D_E_P_T 4 hours ago
CuriouslyC 4 hours ago
A_D_E_P_T 4 hours ago
CuriouslyC 3 hours ago
A number of others have done game/3d video benchmarks but this guy is probably the most prolific.
kroaton an hour ago
therealdrag0 4 hours ago
jdprgm 2 hours ago
Since Luna is so dirt cheap compared to Sol/Astra it would be nice if they could set or you could reserve some small percent like 3-5% of usage pool on codex just for Luna so if you hit usage limits you can at least still run a lot of Luna.
bayesianbot 2 hours ago
poisonborz 2 hours ago
machomaster 16 minutes ago
alright2565 39 minutes ago
aabajian 3 hours ago
fraywing 5 hours ago
Astra is a pretty impressive model. Excited to try this.
gobdovan 3 hours ago
Tadpole9181 3 hours ago
gobdovan 2 hours ago
jumploops 2 hours ago
Yes, benchmarks aren't real work blah blah, but the delta here is so large compared to Astra, it makes it seem like this is distilled Bel or similar.
holbrad 29 minutes ago
This is because they have really aggressively priced a larger model to compete with Opus 5.5, so their margins are much worse. Consequently, the equivalent API spend on the subscription is much less.
modeless 4 hours ago
pmdr 4 hours ago
samuelknight 4 hours ago
resters 27 minutes ago
nzoschke 3 hours ago
I just added an agent / coding agent into an email app, and doing it through `codex` and its Codex App Server couldn't have been easier, and the results are very compelling.
The open source harness, API around it, and friendliness for connecting a subscription puts Claude to shame right now.
A few more thoughts here https://housecat.com/blog/introducing-housecat-agent
TomGarden 4 hours ago
enraged_camel 4 hours ago
Opus 5.5 was a gut punch and my impression is OpenAI is still reeling.
atonse 4 hours ago
The best thing is that we benefit from these constant back and forth gut punches :)
JacobAsmuth 3 hours ago
jhonof 3 hours ago
intenex 4 hours ago
zarzavat 4 hours ago
toasty228 4 hours ago
nater5000 4 hours ago
They can release a new version every day if they wanted to. The question is whether or not the new releases provide substantial improvements or not. It's not hard to just go through the motions, bump the minor version, then make an announcement to rile up the users who don't get that none of this is standardized or regulated in any way and it's literally all made up by the company trying to sell them the product.
tripleee 3 hours ago
ychnd 2 hours ago
ylsilva 4 hours ago
xyzzy123 4 hours ago
user43928 an hour ago
So yes, it is clearly cheap in comparison.
jfrbfbreudh 3 hours ago
Stevvo 3 hours ago
xkcd-sucks an hour ago
It's easy to hit those numbers in a day in an modern-enterprise context synthesizing from incoherent information in jira, slack, layers of codebases etc. Modern enterprise meaning a firm that has been serving a few strategic customers w/ "move fast and break things" since day 1
iamdelirium 5 hours ago
Then Opus 5.5 caught them off guard and now they're actually releasing the correct sized model.
hyperpape 4 hours ago
squidbeak 4 hours ago