Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra (cognition.com)
postalcoder 17 hours ago
enraged_camel 17 hours ago
nullbio 17 hours ago
thereitgoes456 17 hours ago
While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.
Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.
selectodude 16 hours ago
throwaway240403 15 hours ago
throwup238 16 hours ago
Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.
mediaman 17 hours ago
Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
felixgallo 17 hours ago
letmevoteplease 16 hours ago
Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.
> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)
And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.
vlovich123 16 hours ago
kzrdude 15 hours ago
That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.
Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.
xnickb 4 hours ago
kzrdude 3 hours ago
general_reveal 17 hours ago
You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.
Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.
iLoveOncall 16 hours ago
Yes? Just like every single model from every single AI lab.
postalcoder 16 hours ago
A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?
Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.
Lucasoato 11 hours ago
Wait a second, are we taking into account the massive difference in terms of resources of these two companies?
solenoid0937 11 hours ago
You're either competitive or not.
asdfsa32 3 hours ago
nrmitchi 16 hours ago
Yes.
dpweb 16 hours ago
Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.
gunalx 10 hours ago
willcmcc 15 hours ago
Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.
ben_w 13 hours ago
It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.
Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.
conmod278 13 hours ago
TedDoesntTalk 12 hours ago
Muromec 12 hours ago
NewJazz 12 hours ago
podocarp 5 hours ago
Also no idea what viruses and cancers have to do with this. Cancer is surely very poor reproductively because they never spread to other hosts.
injidup 4 hours ago
Cells have a certain optimised DNA mutation rate kept in check by various machinery. Multi cellular life expects each of these little replication machine to co-operate in the grand scheme of running a body. But it's also required in the grand scheme for DNA to mutate a little bit to ensure population variance. So you could say that cancer is the tax paid for having cooperative yet flexible and adaptive nano machinery.
So yes the propensity for cancer developed under evolutioniary pressure towards a non zero level.
The population could have optimised for zero cancer but it would not have paid for itself in terms of overall population adaptability and survival.
yieldcrv 10 hours ago
These firms are literally hiring professionals from all fields to teach procedure
To teach processes that can subsequently be done agentically or in automated chains
Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth
So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?
eranation 16 hours ago
tonychang430 15 hours ago
fallingbananna 16 hours ago
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
Readerium 15 hours ago
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
nicce 14 hours ago
p1esk 14 hours ago
thereitgoes456 14 hours ago
nijave 13 hours ago
mokre 12 hours ago
Also a lot of questions to benchmark because opus 5 is completely useless model right now.
I think that the main problem with opus that they try to solve context size optimization problem, and that is the main reason why it speaks like alien with only one technical dictionary at hand. So why it is so good?
gpt5 6 hours ago
didibus 5 hours ago
airstrike 12 hours ago
throwatdem12311 16 hours ago
First, almost all models are within spitting distances of eachother.
Second, it never translates to being better for my own workloads.
You just need to make your own benchmarks.
walrus01 15 hours ago
Readerium 15 hours ago
thefourthchime 15 hours ago
I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!
Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
dudeinhawaii 13 hours ago
That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.
gruez 16 hours ago
https://www.youtube.com/watch?v=tNmgmwEtoWE
As others have mentioned this is post trained from Kimi k3, which is already quite capable, so it can't be that bad, but any claimed improvements in performance should be taken with a grain of salt.
notfromhere 16 hours ago
fishtoaster 15 hours ago
htrp 15 hours ago
unshavedyak 15 hours ago
Or at the very least, make more mistakes.
darkwizard42 14 hours ago
sorry so many buzzwords to say, the capabilities to do this kind of work are more accessible and easier to manage, so now it works!
Good to see, and agree they were severely overhyping their product back then.
tyre 12 hours ago
They seem to love a good overpromise.
wetpaste 12 hours ago
arjie 12 hours ago
klardotsh 12 hours ago
paimapi 11 hours ago
I see issues with other harnesses too but not with the regularity I was getting from this. And the one moat they had with the better UI for per-project multi-agent orch disappeared and now is standardized
fishtoaster 5 hours ago
I will say that the lack of parity between Devin cloud and Devin desktop is downright embarrassing. It's very clear that the latter is a thinly-reskinned Windsurf. A visually similar UI with vastly different capabilities. Definitely a black mark on the whole thing.
Saline9515 11 hours ago
princevegeta89 7 hours ago
Everything feels dull and they're always several features behind while Cursor is just killing it every other week.
I'm back to VsCode plus Copilot Pro.
airstrike 7 hours ago
esafak 15 hours ago
sterlind 14 hours ago
justincormack 13 hours ago
Ohentis 13 hours ago
deet 10 hours ago
It's coming from a different starting place than Claude Code or Codex are as individually controlled single-developer tools. Devin has been more persistent in pursuing the direction of something that operates more autonomously at the team level, as a peer. And while it might be slightly behind in raw harness ability (maybe?) it's probably ahead on the team-focus.
thereitgoes456 9 hours ago
deet 7 hours ago
Our experience might also not be typical because we have built infrastructure around making Devin and similar agents work better. And for the record no ties to Devin/Cognition. Just pay them too much as a customer.
nullbio 17 hours ago
I think these competing labs need to realize that no one wants another closed-weight model provider... We aren't even happy with the two we have right now, and their days are entirely numbered. If DeepSeek 4.1 flash is really as good as it's benching, we're probably a month away from 1/3rd of users moving off the closed-weight models in favor of something they have more control over (or is cheaper).
The big labs love to release their new model and quantize after the first week. You don't have that problem using dirt cheap API rates on OpenRouter. DS 4.1 flash is also faster than fast mode Astra. OAI's subscription rates are good value, but now these new open-weight models are nearly as cheap on API usage rates. I honestly can't wait for the day we're not beholden to the two big labs anymore. No wonder there's so much fear pumping happening at the moment from Anthropic and their funded NGOs.
eru 17 hours ago
Only in the world where the incumbents don't react. Eg if they saw lots of users moving away, they'd drop prices or do something else.
bayesianbot 16 hours ago
btw I've had a ton of fun with the new deepseek today, I was waiting for my OpenAI 5h limit reset and decided to give it some problems for fun, got pretty great results. Tried some harder problems and still got great results. I don't expect it to be Sol class or anything but I really didn't expect it to be anywhere near this good so we'll see where it ends up. And it's really fun throwing crazy amount of tokens at the wall for ~free instead of watching the subscription limits tick closer while your agents churn away.
cmrdporcupine 13 hours ago
I'm willing to tolerate babysitting things a lot more if I know I'll get almost instant results.
eru 8 hours ago
user43928 13 hours ago
But great that we have a new leader in performance/price in that segment.
eru 8 hours ago
For example, it was quite good to get a decent Sashiko review. Sashiko is a Linux kernel review agent with interchangeable LLM driver. It's very good, but it eats tokens like crazy.
eru 7 hours ago
Losing most of your customers tends to sharpen the mind a bit. They could eg stop pushing out the absolute frontier for a while and focus on making what they have run cheaper. Or they go and do more lobbying against China. Or a million other little things that take more than 30 seconds to come up with when writing a HN comment, but less than a week for someone who's smart and paid to do this for a living.
colingauvin 14 hours ago
On the one hand you, if you bought a lot of compute a couple years ago (perceived demand, perceived shortage) you are in a good spot temporarily. But the counter to that is that everyone else is becoming more compute efficient so maybe that advantage isn't what people thought it would be. I can almost, almost run DS4.1 Flash at home. 4 sparks can do it at 200+ tokens per second. I have two Sparks, so I am not in the club. Neither is your average laptop owner or gamer either. But your average HN software engineer can probably easily swing 2 sparks.
user43928 12 hours ago
That's like four years of ChatGPT + Claude subscription.
Eight years if only ChatGPT, or sixteen years of the Pro 5x subscription.
colingauvin 9 hours ago
eru 8 hours ago
I'm not sure? If we have techniques to use the hardware even better, that will make the hardware even more valuable, won't it?
colingauvin 5 hours ago
notfromhere 16 hours ago
Basically any successful AI based service will do this because at scale the frontier models are expensive and you’ll have enough data to fine tune your own.
Same reason Harvey is doing models now and basically every other provider
didibus 5 hours ago
onel 3 hours ago
pizza234 16 hours ago
DS 4 Flash requires large amounts of memory to run at reasonable quants (I think a system with 160 GB or so). DS 4.1 Flash is even larger, I think around 250 GB.
Any DS version is dumb when compared (in realworld tasks) to Astra/Opus 5, which means, one would spend thousands of dollars, and still need to rely on cloud services to do jobs that are non trivial.
pkilgore 17 hours ago
AznHisoka 16 hours ago
PolCPP 15 hours ago
I used to use windsurf as my main editor until they changed their pricing model. Now i use it just to burn my weekly tokens on fable/astra if i remember to that on a task and that's it.
chris_st 15 hours ago
dominotw 15 hours ago
0l 13 hours ago
fschuett 14 hours ago
hightrix 14 hours ago
It's a great product compared to Copilot. It is also the first AI tool I used heavily outside of creating random images or one off questions.
I'm now using all three, Devin, Claude, Codex. I'm finding Claude and Codex to be much better. One of my biggest gripes is that the web client and desktop client for Devin are two completely different harnesses, so the quality of responses varies greatly.
klardotsh 12 hours ago
> Devin CLI is easily the worst-in-class coding harness I've ever been subjected to using, rife with bugs (up to and including dropping answers to the question tool), so I can't say I'm exactly inspired to try anything else the company produces.
SaltyBackendGuy 11 hours ago
This made me laugh a bit. I was forced to do an evaluation of their shit product twice due to being backed by the same PE firm; "take a look at it again, it's much better now". It sucked the second time also...
dvfjsdhgfv 2 hours ago
TheJCDenton 17 hours ago
On the one hand I would have expected a completely new model, on the other hand it's an RL-ed K3 go Fable 5 capabilities, which demonstrate that this is probably possible, which is nice.
htrp 15 hours ago
ianm218 15 hours ago
US has OpenAI/ Anthropic/ Google/ meta/ SpaceX atleast trying to make frontier foundation models 2 are using VC money + cash flow, last 3 are mainly cash flow + equity and debt.
China has state banks and similar willing to fund lower margin open source labs.
nijave 13 hours ago
mydreamof 17 hours ago
harmonic18374 17 hours ago
Also the submitter's account is very new which makes me suspicious of self-promotion.
bobtheborg 17 hours ago
Looking forward to 2 -- maybe it'll be usable
captainregex 12 hours ago
CyLith 15 hours ago
skrhee 15 hours ago
CyLith 12 hours ago
mohamedkoubaa 11 hours ago
Take8435 15 hours ago
wy35 7 hours ago
bluelightning2k 17 hours ago
I used to really like Windsurf. (Now Devin. Kind of? But also now Antigravity.) I still use it as my editor but haven't touched the agent for a while simply due to the rise of Codex.
andai 14 hours ago
https://cognition.com/frontiercode
Which is too bad, since all of the gains here appear to be from massively reduced output tokens?
The model SWE-2 is based on, Kimi K3, is cheaper per token than Sol, but costs more per task (ArtificialAnalysis) due to using way more tokens.
Whereas, based on the graphs, SWE-2 appears even more token-efficient than Sol! That might have been worth showing off, if true.
eyeris 17 hours ago
The write-up from yesterday was by somebody from cognition using Devin to translate existing cpu sieving methods to gpu and to optimize the gpu sieve.
pelorat 14 hours ago
sbseitz 15 hours ago
handoflixue 12 hours ago
Dang is pretty good at enforcing stuff. There's a flag button and you can email reports if you're really bothered.
gexla 4 hours ago
MaxikCZ 4 hours ago
gexla 3 hours ago
alansaber 15 hours ago
scronkfinkle 17 hours ago
samyok 17 hours ago
:)
Disclaimer: I work at Cognition, although was not involved in SWE-2
scronkfinkle 17 hours ago
jkelleyrtp 13 hours ago
randomblock1 16 hours ago
samyok 12 hours ago
wren6991 16 hours ago
anthonypasq 15 hours ago
vopi 15 hours ago
breznev 8 hours ago
CamperBob2 16 hours ago
yipinwong 11 hours ago
Same for AI models trained on Kimi-3 or other models like Chinese models do. They suffer from the same issue.
monkeydust 17 hours ago
arrowleaf 16 hours ago
ccapitalK 16 hours ago
IIRC cognition boasted about hiring a lot of competitive programmers and algorithms experts back when they released Devin, so it tracks that they'd use the term.
walrus01 15 hours ago
Tsarp 17 hours ago
airstrafer 17 hours ago
Maybe still worth it if their "64% cheaper" figure holds.
Tsarp 17 hours ago
samyok 17 hours ago
teddyX 16 hours ago
Bolwin 17 hours ago
airstrafer 16 hours ago
FergusArgyll 15 hours ago
Fine tuning you have the actual model weights of the original model, you then train that model to answer in a different (or better) way.
Bolwin 7 hours ago
What you're describing is just synthetic data.
Note Anthropic misused the term in their post about Chinese model distillation, deliberately I assume.
FergusArgyll 6 hours ago
xlbuttplug2 17 hours ago
I wonder if, similar to the American labs, they'll become stingy with their weights once they start getting immediately undercut by a wave of slightly better derived models.
llmslave 17 hours ago
I now just work from my phone, and speak into the agents as they run. I dont write code and I dont write documents. I work on very complicated distributed systems. I dont open my laptop most days. Its a legacy brick I carry around.
Some of my coworkers are still doing things by hand, and are working long hours to produce 25% of the output (when considering hours worked). I stay quiet with my setup. We are in the end times for this job for the people that can see clearly how to automate their own job
hnedeotes 17 hours ago
llmslave 16 hours ago
hnedeotes 16 hours ago
mogwire 12 hours ago
Watching the AI slop my sales reps put in their emails is disgusting but the reply telling them how great of a job they are doing and how insightful their email was says differently.
Many people are laughing to the bank while you are still running `--help` to figure out how to run a complex command.
hnedeotes 3 hours ago
The problem is retards that can only function on a cocktail of drugs, and as they were never good at anything other than anal retentive stuff built and continue to build these retarded systems. Those peddling RoR apps even when they couldn't serve more than 3 or 4 concurrent requests, JS backends to handle complex workflows that even after 2 years of dev. still have bugs and accrued a sprawl of crap to hide the issues of their own making, etc, and yet charge thousands of dollars, those that write shit software that's not even worth to clean your ass with, even though they have 20 years of experience, but then go give conferences and write books about their amazing architectural skills, those that write utils behind the "oh, it's open source, if you don't like it just fork it" and due to marketing get their crap everywhere, while making holes everywhere for their paycheques. Or the nepo babies that need their mexico border run to get their fix so they can have these "humanity changing" ideas? I bet they're the same that before would weasel a 2 week sprint to change the borders of a button. Or burn through 10k in meetings for irrelevant crap. Or get VC funding for a CSS styling company or a two prompt company. Or go on about the value of ideas, but then can't even get that going without outsourcing or an AI to help them have those same "ideas".
Ultimately, you just need to turn into a little pig and party in the pigsty, it's not that difficult either, they say pigs are very close anatomically to humans.
At least AI can help untangle the crap the anal retentive retards have built, and thank god, the pig-mor, this society can't even fuck to replacement levels (perhaps they'll manage now with AI).
gigatexal 6 hours ago
microdrum 10 hours ago
ltsSmitty 17 hours ago