https://simonwillison.net/2026/Sep/8/on-navier-stokes/, https://news.ycombinator.com/item?id=49621697
https://twitter.com/sama/status/2097385167002415140, https://xcancel.com/sama/status/2097385167002415140
https://simonwillison.net/2026/Sep/8/on-navier-stokes/, https://news.ycombinator.com/item?id=49621697
https://twitter.com/sama/status/2097385167002415140, https://xcancel.com/sama/status/2097385167002415140
arctic-true 16 hours ago
naveen99 16 hours ago
sashank_1509 16 hours ago
credit_guy 15 hours ago
chinathrow 16 hours ago
jrflo 16 hours ago
QuesnayJr 16 hours ago
I'm not sure what the top 3 problems are. You can make a case for the Riemann Hypothesis and P != NP, but I'm not sure what #3 would be. Maybe the Langlands program? (That one is not as precisely stated as the other two.)
anthonypasq 15 hours ago
dsdf3 15 hours ago
anthonypasq 15 hours ago
QuesnayJr 15 hours ago
I am not particularly skeptical of claims about AI, compared to the average here on HN, but that doesn't mean every random piece of hype is warranted. What they did is impressive, even though we now know the only reason they threw so much compute at the problem is that they heard a rumor that someone else was already close. Navier-Stokes is not a top 3 problem in mathematics, and it was the one that was thought closest to being solved.
ameliaquining 15 hours ago
QuesnayJr 12 hours ago
ameliaquining 12 hours ago
mrbungie 16 hours ago
1) We don't really know how they arrived to this result except that they had a lead and that they threw millions of compute at the problem. The article is written in a way that makes you believe that it was just an agent loop with little human intervention, but without any evidence.
2) If the threats are to be believed, it is concerning how far they are willing to go to show how capable the model is. One would think their products and credibility would be enough to speak for themselves.
dsdf3 15 hours ago
Personally I anticipated nefarious behaviour as part of a broader marketing strategy to sway the view of those in the west that american frontier offerings were far better and powerful than that of China - that if you did not purchase their offerings you'd be awake every night worrying your competitor was.
And this is boring - they need to admit at some point they misinvested, Anthropic less so. All this math stuff is great... but hello? The largest market cap companies are valuable irrespective of such amplified intelligence.
scurnus 15 hours ago
Regarding product and credibility normal people have a completely different view about LLMs, most don't even know difference between models and probably don't even care about Millenium problems, but care instead if chatgpt can solve their day to day problems. This is just them trying to have the throne on the AI companies space, outside it this result won't matter.
mrbungie 15 hours ago
Did I say otherwise?
> 2) Millenium Problems have been the goal every AI company wanted to achieve since their diffusion, all companies have thrown a lot of resource to solve these problems, as they are very famous and scientists spent a lot of time trying to solve them. The first company to solve it will remain in history, despite all of you finding excuses about it.
I know, but I don't know how that relates to my point, which is about the way they are doing it.
scurnus 14 hours ago
The way they are doing it is by trying to get the attention and staying on top of the news, it is a game they are playing that benefits both OpenAI and Anthropic. The more people discuss SF drama, the less attention Chinese Labs and others get.
andrepd 14 hours ago
danielmarkbruce 12 hours ago
It's really unclear that this entire line of work (training LLMs for proof writing) has much real value outside of writing math proofs. It is reasonably clear that, similar to Deep Blue at the time, people are extrapolating the results to general intelligence because the people who usually write proofs are insanely smart (just like world class chess players).
eutropia 15 hours ago
"Mission. Fucking. Acccomplished."
https://xkcd.com/810/Aboutplants 15 hours ago
chilmers 16 hours ago
[1] https://openai.com/index/research-acceleration-view-inside-o... [2] https://openai.com/index/an-alien-mind/
10xDev 15 hours ago
Miner49er 15 hours ago
10xDev 15 hours ago
ccozan 12 hours ago
glenstein 14 hours ago
I don't know that that's achievable yet. Though the era of increasingly advanced and automated robotics seems to be around the corner which could create a cycle, vicious or virtuous depending on how you feel about it.
Fordec 15 hours ago
dakolli 14 hours ago
Fordec 14 hours ago
dakolli 10 hours ago
People on hackernews are actually some of the dumbest people on the internet. This place is worse than lesswrong.
scott_weber 4 hours ago
a2ff6eeb0 13 hours ago
Scroll down to the existing examples section.
7373737373 12 hours ago
mrbungie 14 hours ago
Fordec 14 hours ago
combobyte 4 hours ago
classified 4 hours ago
HenrikPontoppid 14 hours ago
Fordec 14 hours ago
supern0va 13 hours ago
jsLavaGoat 13 hours ago
And it's definitely supposed to imply some kind of historical discontinuity not a change in convexity.
Fordec 12 hours ago
piloto_ciego 9 hours ago
But thinking about the geometry of this problem helps understand why people aren't adjusting well to this. While we're walking on the curve, we look at the rate of change of the curve and say, "well, yeah, of course, dy/dx is 5 at this point and was 1 at the point a few years ago, because the curve is getting steeper" but we're always going to feel this way as things rip off into the stratosphere because dy/dx(e^x) = e^x.
From standing on the curve the curve is notably seeper, but the steepness totally makes sense to you. It's only when you look back 10 years or so that you think "wait a second, holy hell, I couldn't have imagined this!"
The first time I really (I mean really) thought about the singularity and AI was in roughly 2014. I mean, yea, I'd thought about things before that, but yeah, before the Humans need not apply video, I'd never actually given it much thought. I think back to myself 10 - 15 years or so ago, when I was just starting to tackle real programming projects, and was just starting to get decent at writing code, there is no way on earth that I would have imagined that a little over a decade later, Navier-Stokes would be solved by a computer program and the vast majority of my work would be playing sooth sayer to increasingly complicated piles of linear algebra.
dekhn 7 hours ago
JumpCrisscross 6 hours ago
Bit ironic given the model’s alleged finding…
Singularities are model breakdowns. A singularity simply says our current methods cease to work in this region. Within the context of a recursively self-improving intelligence with an unknown bound, “singularity” is probably a good description of our current socioeconomic system.
monster_truck 14 hours ago
Bit apples to oranges, but it reminds me of all the fiber we installed in the late 90s, certain that per-strand capacity increases were years or decades out, only to get massively rugged
Fordec 13 hours ago
xtracto 13 hours ago
monster_truck 9 hours ago
I get the impression the only reason there is renewed interest in photonics is because DCs are simply out of room (and power) to rack more servers and switches.
Fordec 8 hours ago
hgoel 13 hours ago
In which case, maybe we don't need as much compute as we might expect. I hesitate to say "to reach a singularity" because it's kind of hard to define how that works out. Even intelligence probably hits some scaling limits eventually (e.g. speed of light related restrictions on how far it can scale, or how quickly it can expand).
lijok 13 hours ago
noir_lord 10 hours ago
They are fluffy PR pieces otherwise.
piloto_ciego 9 hours ago
I swear there's nobody blinder than those who won't see.
samrus 9 hours ago
FallCheeta7373 8 hours ago
altcognito 8 hours ago
piloto_ciego 3 hours ago
Is nobody else astounded by this?
aurareturn 5 hours ago
medler 7 hours ago
dekhn 7 hours ago
timr 5 hours ago
The first part of your sentence literally contradicts the second part: "we don't have enough knowledge to know, but I know the opposite".
dekhn 5 hours ago
"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.
Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"
classified 4 hours ago
No, it's an aggravated case, since it's the same way they got all of their training data in the first place.
timr 3 hours ago
classified 4 hours ago
Based on what? Your crystal ball?
tomalbrc an hour ago
AlexCoventry an hour ago
selfmodruntime an hour ago
camel-cdr 5 hours ago
magicalist 16 hours ago
Is this buried under the drama or are the major OpenAI twitter accounts from the people involved in the drama desperately attempting to make this the story after everything else obviously got away from them?
ameliaquining 15 hours ago
20k 15 hours ago
What we're really looking at is seemingly a massive plagiarism scandal, which especially brings a lot of the past results into question
If OpenAI is training models on researchers' prompts, and then threatening them into staying quiet about it, who knows if anything that's been announced is genuine - or just theft?
Edit:
OpenAI have admitted they were training on prompts at the time they made their breakthrough
https://mastodon.social/@tristanbuckmaster/11723647135247030...
ameliaquining 15 hours ago
If you're saying that the question of whether they actually have a highly capable model is less important than the question of whether there's a plagiarism scandal, I continue to disagree.
20k 15 hours ago
The fact that this plagiarism scandal exists underpins the idea that there's actually a mass theft going on, and that these models aren't nearly as capable as is it would seem
lotsofpulp 15 hours ago
20k 14 hours ago
lotsofpulp 13 hours ago
fwip 10 hours ago
orangecat 14 hours ago
Yes, in retrospect I should have been suspicious of that drone hovering outside my window when I was writing down the counterexample to the Jacobian conjecture.
This is just not a reasonable take. Even if OpenAI is maximally guilty here, the work that they "stole" was also largely done by AI.
20k 14 hours ago
letmevoteplease 13 hours ago
ivory54321 14 hours ago
Timwi 12 hours ago
sdenton4 13 hours ago
Now: Could Astra agents have hacked their way into the OpenAI logs to find human mathematicians with a good lead on the problem to build upon? Certainly doesn't seem impossible.
doctoboggan 14 hours ago
user43928 13 hours ago
That's it. The rest appears to be wild speculation.
jsw97 13 hours ago
Never ever touch those requests. If you get a side by side comparison just resend the prompt.
TZubiri 12 hours ago
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, welcome to the subject.
TZubiri 12 hours ago
I don't want to get epistemiological, but these aren't allegations, and your use of "may be" is more of lack of knowledge on how OpenAI and ChatGPT work. Read the Terms of Service, this is not a secret, OAI doesn't deny it, usage of ChatGPT through the web interface or through its App, including Codex, are shared with OAI and used to train future models. This is one of the ways in which ChatGPT works and improves, it's not something we are learning now, it's something that was always known, welcome to the subject.
zem 12 hours ago
flir 12 hours ago
Publication, though? Slimy is right.
But the interesting question to me is: once they had a solution, what should they have done with it? I see two choices: bury it, or contact the mathematicians whose prompts they were listening in on.
an0malous 13 hours ago
ameliaquining 12 hours ago
pama 15 hours ago
danielmarkbruce 12 hours ago
topaz0 8 hours ago
Aboutplants 15 hours ago
stingrae 15 hours ago
blake__dev 15 hours ago
cool_dude85 15 hours ago
blake__dev 15 hours ago
merksittich 13 hours ago
sebzim4500 9 hours ago
bananaflag 15 hours ago
mzhaase 15 hours ago
dboreham 14 hours ago
dakolli 14 hours ago
E-Reverance 14 hours ago
fc417fc802 12 hours ago
sznio 12 hours ago
karmakurtisaani 14 hours ago
Bluestein 14 hours ago
ccozan 12 hours ago
monster_truck 14 hours ago
vatsachak 15 hours ago
Loops and parallel connections make transformer go brrr
curt15 15 hours ago
refulgentis 14 hours ago
_fizz_buzz_ 13 hours ago
gcr 13 hours ago
tristanj 13 hours ago
lossolo 13 hours ago
danielmarkbruce 13 hours ago
preommr 10 hours ago
So interesting how everything played out, I remember in the early days when MS came out with the partnership with OpenAI it seemed like they were playing 5d chess and were poised to win big. And it all just fizzled out.
Second biggest fumble after Google.
danielmarkbruce 10 hours ago
sebzim4500 9 hours ago
irthomasthomas 13 hours ago
fer 12 hours ago
piloto_ciego 9 hours ago
It's all gas no brakes now boys and girls. Hold on to your hats!
itemize123 7 hours ago
vimbtw 3 hours ago
Most models trained for general use are ingrained with certain tendencies that are usually very useful like "if you're stuck and bashing your head against the wall stop and tell the user". You generally don't want Claude Code to go off and work for weeks on something when if it had just asked for help you could've clarified or provided more information or just picked a different approach.
When you're solving extremely difficult math problems though you generally do want a model to be more persistent and keep trying even when the model can't clearly see a way forward. OpenAI appears to have done this with lots of previous models. The model they trained for the IMO competition seems to have been an RL maxxed version since they noted that while it did the math it couldn't write up its results on its own and just produced CoT [0]. The capabilities are in there lying dormant, you just to need to RL max the model to ruthlessly pursue the goal at all costs which destroys general use but improves frontier math.
We've also seen hints from OpenAI at least that they seem to train more persistent versions of all their models [1].
Also, Astra probably completed training at least one to two months before the public release so it's not like they only had a week to whip this version up.
[0]: https://x.com/OpenAI/status/1946594933470900631 [1]: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
soltanov an hour ago
hdivider 16 hours ago
1. It shows what even this wave of AI can actually do.
2. I wish it were done by different folks, ideally under some kind of public control like NASA research or the NPR model.
3. Keep in mind: natural science is different. It's not always a matter of computation. Computer science folks often struggle with this -- but this virtual world here does not actually exist. Everything is physical, including information. Any natural science PhD or otherwise knows just how complicated nature actually is -- e.g. mention any research topic and try to encapsulate all the relevant phenomena present there. Pure mathematics is different because we define the problem, rarher than explore nature. We are in my view far away from removing humans in natural science R&D. Advancements in AI however can greatly assist us in all natural sciences, which is already beginning to happen.
tantalor 15 hours ago
red75prime 15 hours ago
BTW, there's also a problem of asking interesting questions that AIs aren't yet good at.
No one has found any principled walls of AI development yet. And empirical results are quite telling. So, I guess, those problems will not stand for long.
efavdb 15 hours ago
Math is like this too. The big problems they've been solving have been identified as interesting only through lots of prior effort.
vatsachak 15 hours ago
The natural sciences will soon start breaking too.
I will concede that AI seems likely to not invent a "research program" anytime soon.
It has no taste
danielmarkbruce 12 hours ago
The reason AI is doing so well in math proof writing is that it can verify every idea it has, quickly.
geremiiah 15 hours ago
m11a 6 hours ago
ThePhysicist 14 hours ago
So I'm greatly excited what AI will bring about in physics, more so than in math, because in physics it's clear that our fundamental theories are missing a big piece of the picture, and given how easily AI crunches through Millenium prize problems I think it's possible that AI will come up with a viable grand unified theory uniting quantum mechanics and gravitation, or produce new predictions in other areas. There's enough contradictory or unexplained observational data available to make a ton of progress on the theory side I think. Exciting times ahead!
throwaway198846 14 hours ago
alde 12 hours ago
aubanel 20 minutes ago
sobellian 13 hours ago
"If in other sciences we should arrive at certainty without doubt and truth without error, it behooves us to place the foundations of knowledge in mathematics."
semi-extrinsic 12 hours ago
See e.g. https://en.wikipedia.org/wiki/Renormalization
Nobody who actually works in fluid dynamics on any sort of application gives a hoot about the N-S millenium problem. Many do not even know what it is. There is no practical effect of this proof on how we do fluid mechanics.
sobellian 12 hours ago
semi-extrinsic 12 hours ago
sobellian 12 hours ago
semi-extrinsic 12 hours ago
Even if someone comes up with a construction that does not require any forcing, it is going to be some extremely weird initial conditions that you will never be able to even approximate in reality unless you can move all the individual molecules of a fluid around and set their initial velocities from a far distance.
sobellian 11 hours ago
calf 10 hours ago
jarenmf 13 hours ago
brettdev 12 hours ago
danielmarkbruce 12 hours ago
olalonde 11 hours ago
This is sort of what OpenAI was supposed to be. I'll never understand how it was legal for them to turn it into a for profit corporation.
mickael-kerjean 8 hours ago
I would argue the main reason AI labs have been focusing on programming is to unlock industrial scale automation, next logical step is to solve math as it's the key to unlock everything else. Once you hold the key for math, everything downstream fields become a matter of compute
tiborsaas 16 hours ago
WOW?
echelon 16 hours ago
- First off, to reiterate, WOW.
- Second of all, when does this end? Are we at the dawn of the singularity now?
- People are saying OpenAI "stole" this from the work of an OpenAI user. If so, that's pretty fucked - how can we trust them?
- Time to think about retiring from any knowledge work or business? This could be winner-take-all where a leading lab can button press any economic function, business process, or scientific discovery. 24 months of lead on Open Source might turn into virtual centuries of lead.
- Do "normies" even know what's happening?
Anybody who thinks the improvements stop here isn't paying attention. It hasn't been showing any signs of slowing down since 2018. And the curve isn't even linear! My god, next year is going to be insane.
d_silin 16 hours ago
lbreakjai 4 hours ago
raincole 16 hours ago
The said user (Tristan Buckmaster) didn't solve the millennium problem. He didn't really accuse that OpenAI stole his research either. The beef came from the fact OpenAI asked him to remove another mathematician, who works for Anthropic, from the credit.
"People" are just misinformed and keep spreading misinformation.
naasking 16 hours ago
Not quite accurate, Buckmaster was taking an approach that nobody else was, and this new proof uses this same approach just weeks after he saved those results to OpenAI workspaces. He asked OpenAI if they used chat logs for training the new model, and they did not confirm or deny.
Asking to remove his collaborator is also totally over the line though.
Edit: although this OpenAI post is not comforting: https://x.com/OpenAI/status/2097375276384567642
Quote: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. "
emp17344 16 hours ago
stefap2 15 hours ago
achierius 16 hours ago
> I should say here why I interpreted their statement the way I did, the in- terpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.
...
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI. > I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
It's not a direct accusation, but it's not far off.
You shouldn't accuse other people of spreading misinformation when you haven't read the actual sources in question, it's possible that they might know more than you.
raincole 15 hours ago
> I have not seen OpenAI’s proof. I do not know what their model did, or how. I do not know whether our data was used. I am not accusing anyone of anything.
People saying that he accuses OpenAI stole his proof are putting words into his mouth and I consider that very disrespectful to him. It's basically using Buckmaster as a tool to express their dissatisfaction over OpenAI.
calf 10 hours ago
20k 14 hours ago
He very much is accusing them of stealing his work
bananzamba 10 hours ago
It's like he had a treasure map and was about to find the treasure, but they copied his treasure map and scooped him with a faster boat and found the treasure first. But he would have found it if it weren't for them.
tiborsaas 16 hours ago
3) I'm still processing the drama, just found out about it after reading the blog post. If that happened based on private data, that's horrible. If that happened based on public tweets, then it's still abuse of power as OA employees access to compute (launching 10k agents) is quite heavy weight in boxing terms.
But apart from AI and drama now that we have working solution to Navier-Stokes, what improvements can we expect in engineering?
inkysigma 15 hours ago
There's unlikely to be any engineering applications since even if the solution can be approximated, you still need to set up the initial conditions but at that point you can also drive pressure in other ways.
thangalin 14 hours ago
Proving out the combination of scaling inference-time compute and agent collaboration to solve previously intractable mathematical problems is WOW. By pairing creative candidate generation with automated proof checkers (like Lean) we are leaning into a repeatable framework for AI-driven scientific discovery.
semi-extrinsic 12 hours ago
This is 100% wrong and reads like copy paste of AI slop.
Any simulation which uses sub-grid scale models is already solving a different PDE than the actual Navier-Stokes considered in the Millenium problem, and that PDE is guaranteed to have different properties. Full stop.
And to claim this is somehow connected to AMR methods is an example of the kind of pseudoscientific statement Wolfgang Pauli would have called "not even wrong".
20k 15 hours ago
Its a bit like solving p = np with a negative result. Its an incredibly difficult problem, but it doesn't lead to anything at all on its own. This is why people are talking about the fact that the solution methodology is much more interesting than the solution - the tools used to crack something like this may lead to solving more useful problems
cyberax 15 hours ago
Nothing, really. This mirrors other examples of blowups from the classical physics. It's possible to create a system with just gravitating bodies that exhibits a blowup to infinite speeds in a finite time. The root cause is that, in classical physics, the speed of gravity is instant.
In the case of Navier-Stokes, the fluid is incompressible. So technically any force that you apply to it is supposed to instantly affect everything else. This can be exploited to create these blowups. In reality, no fluid is incompressible, and it takes time for any action to affect the material.
It's just that Navier-Stokes equations are so slippery that it's hard to pin their behavior down. They basically just restate the momentum conservation law for a continuous medium.
mcfry 10 hours ago
cyberax 10 hours ago
There's a Wiki article about it: https://en.wikipedia.org/wiki/Painlev%C3%A9_conjecture
xyzsparetimexyz 12 hours ago
Minor productivity boost in mathematics as people are no longer nerdsniped by the problem
stefap2 15 hours ago
munificent 15 hours ago
You really think it makes sense for you to be higher on the "solving complex problems ladder" than the machines that solved fucking Navier-Stokes?
I envy your self-confidence.
stefap2 15 hours ago
FridgeSeal 10 hours ago
mlsu 15 hours ago
reducesuffering 15 hours ago
mlsu 15 hours ago
If that is true then this seems to be, again, a case of AI producing an interpolation over data it has seen before. Everything about openAI's behavior indicates that they were using the transcripts as input. Why not have the AGI choose a different Millenium prize problem?
FabCH 11 hours ago
It’s not exactly a strong argument against AI.
FridgeSeal 10 hours ago
FabCH an hour ago
That's not interpolation though, that's theft.
j_maffe 11 hours ago
Well if you do the math, the number of agent-compute time in total, given the insane number of agents thrown at the problem, might end up being comparable in time, if not for the budget.
fooker 14 hours ago
For example there are no engineering implications of this solution yet.
For the next several decades, we'll have engineers (presumably with AI) optimize things like rocket engines and turbines and AC compressors to work a few percent better because the numerical approximations might have caused us to be overly conservative.
AI is not going to magically solve all random problems. Pick a career where you are in the driver seat.
semi-extrinsic 12 hours ago
No. Just no.
cyberax 14 hours ago
But it did find a long-suspected smooth solution with a singularity.
naishoya 3 hours ago
Ongoing publications of statements produced by both sides of this situation do seem to support that this is an intentional effect of the hiring of these world class mathematicians at competing firms: to specifically use the research of those human minds to create a perception of capacity as if it came from the machines and the models.
Without those minds and the 'training data' derived from the intermediate stages and intuitions of those minds the models cannot be shown to be capable of this result.
A hammer and saw wont build a house, not even a dog house on their own, and while being shown capable of using software tools in ways not stated as direct instruction (see HuggingFace breaches) these models do not demonstrate naive intuition nor novel capability.
This outcome regarding N-S demonstrates that in the hands of world-class minds these models can be induced to coalesce interesting accumulations of information and results, but using these accumulations as proof of innate capability is exactly the pre-IPO motivated behaviour we should all be wary of, and all mathematicians who currently are assisting in this market manipulation in return for remunerative consideration need to be cautious of the potential disgrace that this brings to their reputations and that of the field.
I get that the need to pay the bills is a strong motivation in these times of uncertainty, but there are numerous examples in history of world class mathematicians being perfectly capable of at the same time producing world changing results and also working at normal professions; as barristers, magistrates, ministers, primary school teachers, translators, draftsman/engineer, banker, miller and baker, private math tutors, weavers, clockmaker and locksmith, merchant, patent officer, Augustinian monk turned exiled Protestant preacher, physicians, cryptologists, soldier, telegraph operator, astronomers, physicists, chemist, agriculture manager, political writer, oboe player, organist and music director, architect and surveyor, librarian, statistician, habidasher, brewer (at Guiness in one case: William Sealy Gosse ~ originator of t-distributions), bookbinders apprentice, hospital administrator, and even the first creator of the first computational model of a neural network, which serves as the structural grandfather of modern Artificial Intelligence was a low level laboratory assistant.
Sure this list includes professions and employment which are obsolete, but my reasoning stands, there are jobs available. Arguing that 'because the pay rate is so high' as a reason to abdicate moral responsibility for personal involvement in unethical market manipulations simply demonstrates a lack of personal ethics. Whether the choice is through lack of self awareness or a conscious choice to become wealthy in spite of any such breach of the public trust is immaterial to the outcomes, the 'if i don't someone else will' argument should be met with the same derision for any con-man's Ponzi scheme no matter how new the technology, no matter how many zeros are in the bribe.
trio8453 15 hours ago
No, there are even many non-normies talking about how it's all marketing or try to give balanced take about AI being sometimes a little useful for certain things (but they can do without it anyway).
ImaCake 10 hours ago
sire-vc 10 hours ago
The world has changed in 12 months and many people didn't notice, and those that did are eating their peers' lunch.
armchairhacker 15 hours ago
bibimsz 15 hours ago
qlte 11 hours ago
bibimsz 11 hours ago
reducesuffering 15 hours ago
armchairhacker 15 hours ago
p1esk 4 hours ago
qlte 12 hours ago
"Moving the goalposts" as shallow dismissal doesn't work if e.g. someone points out AI hasn't even built a new type of spaceship yet in response to a claim that AI is on the verge of building a Dyson sphere.
Bluestein 15 hours ago
onidj 15 hours ago
Absolutely not. Even to a lot of techy/nerdy people it's still just a chatbot that they sometimes use to help them at work. Even on here people will do whatever they can to downplay.
The lack of fucks given is staggering.
nozzlegear 12 hours ago
senshan 8 hours ago
tantalor 15 hours ago
Singularity doesn't "dawn". That's the whole idea. It happens all at once.
echelon 15 hours ago
tantalor 15 hours ago
FabCH 11 hours ago
The event horizon would then be the time period between the singularity becoming inevitable and it actually happening.
biophysboy 15 hours ago
root_axis 14 hours ago
This is an impressive result, but there is absolutely zero evidence of "the singularity".
hackinthebochs 12 hours ago
root_axis 12 hours ago
hackinthebochs 11 hours ago
root_axis 11 hours ago
hackinthebochs 11 hours ago
stratos123 11 hours ago
root_axis 11 hours ago
stratos123 10 hours ago
I think you are implying that it's invalid to consider every advance to be evidence "for", and I agree - that'd violate conservation of expected evidence. But not considering any advance to be evidence "for" is also invalid, for exactly the same reason. There has to be some news you may hear that'd make you think a singularity is more likely, and "millenium prize problem solved by an LLM" sure seems like one of those.
baq 14 hours ago
> - Second of all, when does this end? Are we at the dawn of the singularity now?
normalcy overhang n. /NOR-muhl-see OH-ver-hang/
The uncanny period during the Singularity when superintelligence is already accomplishing feats that seem like magic, yet everyday life still looks mostly the same.
jampekka 14 hours ago
This.
I do dislike the AI oligarchs as much as the next person, but I do find the thread full of complaining a bit depressing still.
If the result holds (and it looks it does), this may be one of the, if not the, biggest things to happen in computing to date. A lot bigger than e.g. Deep Blue beating Kasparov in chess or AlphaGo beating Sedol in Go.
Eridrus 14 hours ago
karmakurtisaani 14 hours ago
xhevahir 14 hours ago
empath75 14 hours ago
concinds 14 hours ago
robryan 12 hours ago
tzone 11 hours ago
There would be no need to deliver new mathematical insights by solving problems. You would just have a magical math problem solving machine and that’s it.
FridgeSeal 10 hours ago
To further human understanding is itself a goal that single-handedly justifies our efforts.
Jumping straight to the “answer” and therefore missing both the understanding of the actual problem, and any useful discoveries along the way is a waste at best, and actively harmful at worst.
tzone 9 hours ago
rybthrow2 13 hours ago
qlte 13 hours ago
But, as mathematicians learn and push forward, occasionally something like elliptic curves will emerge as having useful applications, making all that previously "pointless" specialized knowledge newly valuable.
Or advances in physics, that suddenly have a need for a specific mathematical underpinning to develop a theoretical framework. Like how Einstein benefited from Minkowski's work on hyperboloids to create a coherent mathematical description of spacetime.
It was the AI labs themselves not mathematicians who were happy to conflate proofs for open math problems with some kind of tangible technological advancement in the real world. They would surely prefer to be able to claim a cure for cancer vs. a math problem but that loop requires a lot more time/money/test tubes/etc and they need headlines now not in a decade.
And so, thanks to OpenAI/Anthropic, we're now in a world where thousands of crypto bots on X breathlessly hype up each new problem being solved that previously wouldn't have any got any attention beyond academia and passionate fans of math.
Hopefully this won't lead to a trough of disillusionment as more people start to feel like you, with mathematicians getting the blame for inflating the value of their work even though the hype was coming entirely from the labs not them.
vouaobrasil 10 hours ago
But if AI can solve any problem and existing mathematicians just use AI to solve problems for the sake of solving them, the community itself with wither and so will the interest in mathematics and over a longer period of time, it will just become soul-less and uninteresting and the entire community powered by the fire of fascination will simply die.
throwaway81ag81 9 hours ago
Why would you go out of your way to make a case of something being not useful when, ironically, so much advancement in human history has come from the discipline?
Your motive is more worrying than your straw man argument.
lukewarm707 9 hours ago
what would have happened if they didn't do this? they could have done things the right way. we could celebrate this achievement.
perhaps they could have taken just a few weeks, even, to work out something with buckmaster and apoge.
i am very concerned that this is a glimpse of things to come. a malign elite with powerful ai, who will turn the sublime of technology into barbarity, violence and human oppression.
if sam altman has destroyed openai's public benefit corp with lesser tools why would he aid humanity with even greater power?
throwaway81ag81 9 hours ago
People are not complaining about problems being solved or advancement in technology. They are complaining about terrible people doing terrible things.
mewse-hn 16 hours ago
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
dash2 16 hours ago
avs733 6 hours ago
An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk.
There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.
nradov 16 hours ago
gowld 15 hours ago
red75prime 14 hours ago
But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
brainwad 14 hours ago
WarmWash 15 hours ago
If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
perching_aix 15 hours ago
lima 13 hours ago
spruce_tips 13 hours ago
magicalhippo 12 hours ago
Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in.
You'd have to take their word, but that goes for anything in life.
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is...
14u2c 13 hours ago
TZubiri 12 hours ago
nozzlegear 12 hours ago
At this point, how can we even trust that they aren't accidentally training on those tokens too?
jdm2212 12 hours ago
nozzlegear 11 hours ago
One would think getting caught asleep at the wheel while their bots are escaping containment and hacking third parties would be corporate suicide. One would think that potentially stealing their competitors' work on the Navier-Stokes problem would be corporate suicide.
Alas we live in bizarro world where there are zero consequences (maybe the opposite, in fact) for the first, and their employees meme about the second on social media.
WarmWash 10 hours ago
Boardrooms run businesses, not bookstore ethics clubs.
nozzlegear 7 hours ago
jdm2212 6 hours ago
mmanfrin 4 hours ago
Would it, though? Considering their entire business model is built on the agglomeration of data that isnt theirs.
sinuhe69 14 hours ago
jimbob45 14 hours ago
vessenes 14 hours ago
All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
elwell 13 hours ago
taylorfinley 12 hours ago
ImaCake 10 hours ago
kypro 10 hours ago
The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction.
I get the scepticism, but I feel some of the accusations here are bad faith.
mzs 5 hours ago
nullbio 4 hours ago
THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE.
"TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data.
I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
jakevoytko 16 hours ago
Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees
philipwhiuk 14 hours ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
closetheloopdev 14 hours ago
traes 13 hours ago
> 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee.
> 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)
https://xcancel.com/SebastienBubeck/status/20973794116915163...
hellohello2 9 hours ago
I remained impressed by ChatGPT however!
angry_octet 8 hours ago
But they have learnt their lesson, next time they won't reach out to who they stole it from, they will publish first.
jdm2212 8 hours ago
angry_octet 6 hours ago
jdm2212 5 hours ago
OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.
Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.
I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.
Palmik 4 hours ago
keeda 12 hours ago
It's like that story about George Dantzig solving open problems as a student because he thought they were simply homework: https://en.wikipedia.org/wiki/George_Dantzig
It's also unfortunate that such a potentially momentous occasion is overshadowed by so much drama. Which I suppose is expected given the technology and the people involved are so polarizing.
recitedropper 16 hours ago
The dark forest awaits..
vmasto 16 hours ago
sheafification 16 hours ago
recitedropper 15 hours ago
So less about hiding civilizations, and more about hiding information. Math is clearly headed in this direction, and I see no reason why the rest of intellectual work shouldn't too.
AlexErrant 15 hours ago
2. The dark forest is fun for scifi stories, but is mathematically bunk anyway https://www.noahpinion.blog/p/the-dark-forest-hypothesis-is-... https://www.reddit.com/r/IsaacArthur/comments/1l06cnk/cool_w... https://www.projectnash.com/aliens-the-fermi-paradox-and-the...
When doomposting please actually say something substantive. Negative news always gets clicks/updoots; fight that human tendency.
recitedropper 15 hours ago
I agree that we have not solved the Fermi paradox; I disagree that comments highlighting immature behavior from people who wield enormous power in our world are unproductive.
AlexErrant 14 hours ago
Separately, I disagree that intellectual work has ever been free of "dark forest"-style secrecy. Scientists everywhere have worried about being scooped; AI just magnifies that (as all tools have; e.g. Leeuwenhoek lenses).
And thirdly, if you want to make a stronger case for "I feel even less confident in them as a team to be shepherding this much capital and compute", you should give citations and arguments. From what I've seen, there's drama, it's much OpenAI trying to avoid scooping, and Tristan being stuck in a game of telephone, and Levent being incommunicado.
If you have a better analysis, you should say so instead of being vague.
recitedropper 13 hours ago
I appreciate your upholding of ideals, and since I respect that, I will honor with final replies:
1. Locktime has passed.
2. Yes, intellectual work has always had elements that incentivize secrecy. If you want to say we were already in a "dark forest", so be it. My suggestion is that the multiplier AI adds to the possibility you get scooped is a step-change, and therefore we now enter a new "dark forest".
3. This is a big thread, and the twitter antics are well-documented, so I would assume someone else has cited them. If not, I think most are aware at this point that the online antics of AI researchers, especially when announcing or citing mathematical advancse, have regularly been childish.
AlexErrant 13 hours ago
3. ctrl-f "x.com" in this thread only yields https://x.com/sama/status/2097385167002415140 https://x.com/SebastienBubeck/status/2097379411691516310 https://x.com/OpenAI/status/2097375276384567642
and frankly, I'm unwilling to give Elon any more traffic to dig up drama that ultimately doesn't matter. I'm not seeing anything especially childish, but y'know... I'm not sure I care.
zem 12 hours ago
AlexErrant 10 hours ago
I retract the projectnash citation; I grabbed it from the Cool World's youtube description, thinking it was a blog version of the video. It was not. I suggest watching the video instead.
nicce 13 hours ago
closetheloopdev 14 hours ago
- There are at least two versions of a model more powerful than Astra at OpenAI at the moment.
- The less capable version was used to solve the unforced Euler problem (while the one solved by Levent Alpöge and Tristan Buckmaster was forced Euler) with 100 agents.
- The more improved version was used to solve Navier-Stokes, given the results of the unforced Euler problem from their earlier attempt, with 10000 agents.
- OpenAI initially tried a shotgun approach against the 6 Millennium Prize Problems until it emerged that Navier-Stokes was the most likely to succeed.
So the timeline was:
Shotgunning 6 open Millennium Prize Problems -> solved unforced Euler problem with 100 agents -> concentrating on Navier-Stokes with 10000 agents -> solution.
If so, that is fantastic development and a huge success (despite all the drama surrounding it)! Congratulations!
tristanj 12 hours ago
closetheloopdev 12 hours ago
I hope the next solved Millennium Prize Problem will have less drama.
bananzamba 10 hours ago
Meaning they would have found it first if OpenAI hadn't spent millions in compute on following their lead to its conclusion faster than them.
closetheloopdev 9 hours ago
To be clear, I only talked about the mentioning of the approach to solving it and not of the proof in the training data.
intenex 14 hours ago
This specific problem having had a $1 million bounty on its head and still remaining unsolved for 26 years after the bounty was placed is pretty clear evidence that many of the world's best human mathematicians would have solved this problem if they could have, and none were able to until LLMs came along.
Hard to claim at this point that LLMs aren't capable of novel STEM creativity and genius to a degree that will soon far surpass that of humans.
If anyone has counterpoints to this I'd love to hear them!
philipwhiuk 14 hours ago
piker 14 hours ago
intenex 14 hours ago
piker 14 hours ago
[Edit: my only point here is that the prize is probably not driving human effort to the limit.]
superxpro12 14 hours ago
jampekka 14 hours ago
gpm 14 hours ago
adverbly 14 hours ago
To be fair, I think it's still an open question about how far it might surpass human capabilities.
I think it's clear that its speed of development will be significantly faster, but it's technically not proven that the frontier and problems don't themselves become increasingly difficult faster than any acceleration in intelligence past the point of human training, data and existing knowledge.
Should this be the case, we would see a rapid broadening of development, and a slow advance in the frontier in such a way that might surpass the collective capabilities of people, but not by very far.
redox99 13 hours ago
lixtra 7 hours ago
aeve890 13 hours ago
Sure. A proof without an unknown amount of human steering (and/or stolen research) would be an unquestionable achievement.
To this day there's zero (0) evidence of any result by an LLM alone (maybe I'm wrong). If I just prompt ChatGPT right now with "give me a proof of the Riemann Hypothesis" and this thing delivers, I'm sold. But anything close to "yeah ChatGPT proved X with 5 years of 24/7 work with 10x Terrence Tao level geniuses" it really doesn't cut it.
Or why's there's no new branch of mathematics invented by AI? That'd be indubitably _novel_ and _creative_. But to my knowledge (and I'm eager to be educated) there's nothing like that. What are the HARD examples of novelty, creativity and genius you claim? For how people like you talk about AI I'd expect idk, a unified theory on fundamental physics, or a novel engineering solution for material science and nuclear fusion, or at least improve itself to not need a bazillion GPUs to emulate a 20 watts wetware. Sure it would infinitely easier to make OpenAI literally print money with any of the thousand problems easier to solve with such amazing intelligence than the NSE problem right? Honest question
redox99 13 hours ago
Also there are proofs where the only human steering was "keep going".
aeve890 12 hours ago
Any result of such kind from an AI alone would be enough to refute my argument, yet you don't present any.
>Also there are proofs where the only human steering was "keep going".
Which ones?
masterspy7 11 hours ago
>Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
aeve890 8 hours ago
indigo945 25 minutes ago
> Drawing on extensive prior research by mathematicians over the past decades, it
> [Claude] has increased this bound [for the fraction of zeros of the Riemann zeta
> function that satisfy the Riemann hypothesis] from 41.6% to 67.2%. Claude also
> produced a formally verifiable proof of its result.
How is a formally verifiable proof not a proof? You're making literally no sense.sp527 11 hours ago
aeve890 11 hours ago
monk_grilla 9 hours ago
I think the counterpoint here is simply to look at what was being achieved with LLMs one year ago versus today, and extrapolate that trend. Sure, there may not be examples of what you've asked for yet, but Astra is literally a couple of months old, the model that solved Navier-Stokes is less than two weeks old. It appears that we're seeing the hockey stick that only the most bullish thought was possible.
hansvm 11 hours ago
Not to mention, it's still very much up in the air whether the model derived the answer of its own accord or sniped the important details from the researchers it was spying on.
hollowcelery 8 hours ago
hansvm 7 hours ago
>> isn't a counterpoint
Where do we disagree?
eulgro 7 hours ago
hansvm 4 hours ago
> How can I afford?
Mostly not my money, tech makes one fabulously wealthy, I care more about math than fast cars or whatever (even as a pizza driver you can afford a fancy car if that's your primary motive), etc.
> Already solved
That's a very interesting insight into mathematics. It's absolutely just as interesting to prove that certain proof techniques are or aren't possible as it is to actually prove the main result, sometimes moreso, especially if they have any chance of improving attacks at other problems.
yauneyz 11 hours ago
Proving something in the affirmative often requires the creation of an entire new sub-field of math, or new tools. Think of Fermat's Last Theorem or something like that.
These results, while impressive, are clever constructions using existing techniques. It isn't clear that AIs can build new machinery like this. But if/when they can, yeah it is probably game over.
mellosouls 10 hours ago
When you read the detail the compute they are throwing at it is incredible, tens of thousands of agents with different groups competing.
It's not like a single Gauss as you imply, "just" many, many mathematicians working tirelessly in a completely ego-less way, guided by other agents and ultimately humans, built - allegedly - on recent human insights.
Stunning, undoubtedly, but this is a "brilliant autistic herd" result, not that of a singular mind.
monk_grilla 9 hours ago
I slightly disagree. A single LLM is equally 'mindless' as a herd of them. As anyone will tell you they "simply predict the most likely next token," yet, complex solutions to difficult problems arise from them.
Many people have said that the architecture of LLMs will need to change for true ASI. I think that the herd of tens of thousands of agents can be seen as one such potential architectural extension. Whether or not a herd or a single LLM is used for a result like this is irrelevant.
To be clear, I think the orchestration of thousands of LLMs in their current form, even with ever increasing intelligence, is not the form ASI will take. There is still a major architectural breakthrough to come, in my limited, ignorant opinion.
piker 13 hours ago
webcoon 40 minutes ago
At a conservative estimate of GPT 6 Astra pricing, this would have cost upwards of 15 Million dollars for anyone using the OpenAI API!
To me this is the one silver lining. Yes, they can solve millennium prize problems, but it still costs a fortune.
pred_ 16 hours ago
And what's a better way of empowering people than robbing them.
rfgplk 16 hours ago
Better than the walled gardens of most journals where you can't even read half the papers without shelling over thousands of $$$
20k 14 hours ago
pavel_lishin 16 hours ago
https://news.ycombinator.com/item?id=49605915
https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcd...
beering 16 hours ago
floatrock 16 hours ago
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
cute_boi 16 hours ago
andrewguenther 16 hours ago
luke5441 14 hours ago
That they don't is telling.
ImaCake 11 hours ago
biophysboy 16 hours ago
enraged_camel 12 hours ago
tristanj 11 hours ago
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
biophysboy 10 hours ago
JumpCrisscross 4 hours ago
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
tristanj 3 hours ago
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
recursivecaveat 10 minutes ago
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
karmasimida 8 hours ago
tedsanders 16 hours ago
I work at OpenAI, though not on the team that did this, and my understanding is:
- we decided to ask our model for Millenium problem solutions because of two reasons: (a) our new model was looking incredibly good and (b) we heard rumors that some Millenium problems had been solved and were curious if our models could solve them (the goal here was not to scoop any particular individuals and we were looking at many problems beyond these)
- we did not read any private chats (but of course the model was aware of prior research literature published to the internet)
- the proof generated by our model was very different from theirs and also goes far beyond the published literature
- we made an effort to jointly announce rather than immediately scoop (I understand Tristan was unhappy with the conversations; I know zero details here and I hope more is shared today)
Edit: Here's is Sebastian's take: https://x.com/SebastienBubeck/status/2097379411691516310?s=2...
suddenlybananas 16 hours ago
tedsanders 16 hours ago
sk4rekr0w 16 hours ago
suddenlybananas 16 hours ago
https://news.ycombinator.com/item?id=49605915#49610498
https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...
(I can't reply to the below comment, but I was aware this was about Sebastien, I was trying to be charitable by including stuff said about both people)
qt31415926 15 hours ago
sk4rekr0w 15 hours ago
suddenlybananas 16 hours ago
applicative 16 hours ago
I dedicate my life to its complete destruction beginning today.
senordevnyc 9 hours ago
contemporary343 16 hours ago
- This, from Tristan Buckmaster's writeup yesterday, indicates to me that there was more than incidental inspiration from Alpoge and Buckmaster.
tedsanders 15 hours ago
- "very little human" input feels ambiguous, and if someone spends a few days prompting a model to solve a super hairy problem requiring a 100-page proof, I can understand reasonable people interpreting that as both "very little" and "not very little" human input
- it's all true that a team worked on this, a bunch of compute was burned, and the problem was solved in stages and pieces
I'm not sure how any of this provides evidence that OpenAI took any of their work.
As evidence against, we never looked at any of their ChatGPT conversations and our model's proof is quite different from theirs.
(I work at OpenAI, but not on the team that did this proof.)
enraged_camel 12 hours ago
Sorry, but the burden of proof lies in the other direction: OpenAI needs to definitively prove that their agents did not look at the existing work that was about to be published. Otherwise OpenAI simply stole the glory and the spotlight (and I'm being charitable here).
fc417fc802 12 hours ago
nulld3v 12 hours ago
fc417fc802 12 hours ago
Only the CIA knows whether or not they're actively covering up reptilian space aliens exerting control over the US government. Therefore the burden of proof remains on the CIA to prove that they are not actively participating in such a scheme.
nulld3v 11 hours ago
fc417fc802 11 hours ago
I didn't realize you had insider knowledge about their systems. Do please explain for the class.
As I understand it they will only have trained on his data if he consented to it. Do you have evidence that they do otherwise?
Timon3 10 hours ago
hellohello2 10 hours ago
machomaster 2 hours ago
tristanj 10 hours ago
If it was enabled, then their work was included in the training dataset.
opello 9 hours ago
At least, as an ignorant outsider, that's how it seems to me.
Arodex 11 hours ago
fc417fc802 11 hours ago
You don't just get to subpoena your neighbor's bank account because "I know he's stealing from me" you need to first present credible evidence that you were stolen from and that he is among the most likely culprits.
Arodex 10 hours ago
Any other argument, fc417fc802?
tristanj 10 hours ago
See Russell's teapot for an explanation https://en.wikipedia.org/wiki/Russell%27s_teapot
magicalist 5 hours ago
This is incorrect, and you invoke Russell's teapot incorrectly too.
It would only apply if the accusation rested solely on the fact that neither of us have evidence against the accusation.
But that's not the case. First, we know that there could be proof, it's just apparently burdensome and expensive to produce. At that point you're not in fallacy land anymore, you just need a way to balance the cost required of someone to prove the accusations against them false.
Second, we have an arguably plausible mechanism of action that OpenAI does not dispute is possible.
This isn't a legal dispute, so no one is going to force OpenAI to do anything here, but it's not unreasonable (and certainly not fallacious) to suggest that Buckmaster's suggestions are plausible enough it's up to OpenAI to stand behind their denial.
tristanj 4 hours ago
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
You seem to think that it is some "gotcha" that OpenAI refuses to make a blanket denial, but they cannot do so in good faith, because they have a genuine understanding of their own system. They don't know where the data they have came from.
This situation meets the requirement of Russell's teapot, since neither party has enough evidence to prove nor disprove what information is actually in OpenAI's training set.
Arodex 26 minutes ago
hellohello2 9 hours ago
In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?
fc417fc802 9 hours ago
If there's more to the story I'd be interested to hear it.
hellohello2 8 hours ago
airognio 6 hours ago
derangedHorse 12 hours ago
I don’t think they’re too concerned about appeasing you, enraged_camel.
For most reasonable people, achievement in solving the other Millenium Prize problems at an unprecedented rate will be enough. At some point people will see models are capable of solving hard issues without whatever 0.00001% of the training data coming from irate individuals who believe their sample was the key component of the solution.
kamaal 3 hours ago
https://en.wikipedia.org/wiki/Burden_of_proof_(philosophy)#P...
whimsicalism 12 hours ago
tristanj 11 hours ago
It's unknowable and not possible to prove if any one specific conversation was the key to solving Navier–Stokes.
DetroitThrow 11 hours ago
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
tristanj 11 hours ago
You might retort that ChatGPT used the training data to copy their approach, but the approach Buckmaster and Alpöge chose was already published by Luis and Diego in 2023 and in every frontier model's training set.
whimsicalism 10 hours ago
DetroitThrow 10 hours ago
tomcorrigan 10 hours ago
However the way you are conducting yourself in public, while announcing yourself as an OpenAI employee is doing enormous harm to the greater and magnanimous aim of your organisation. Take a step back and read the temperature of the room. Being the smartest guy in the room will never protect you from alienating the rest of the room into a baying mob. Right now you are Icarus flying straight into the sun.
tristanj 10 hours ago
senordevnyc 9 hours ago
hellohello2 9 hours ago
tristanj 9 hours ago
If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.
hellohello2 9 hours ago
All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.
Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.
In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.
Pulcinella 11 hours ago
tristanj 11 hours ago
whimsicalism 10 hours ago
FWIW, publicly facing OAI docs are very unclear about whether this setting even applies to Codex conversations.
tristanj 6 hours ago
it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won’t do it
amazingman 5 hours ago
unsupp0rted 4 hours ago
ted_dunning 2 hours ago
We have lots of examples now of their model doing what they say is impossible.
Now we have another example of something that they say is impossible or very unlikely. Do we take their word for it this time? Really?
illiac786 4 hours ago
fwip 10 hours ago
whimsicalism 10 hours ago
I don't really understand how the quantity of training data/rollouts used in training is relevant to the question of whether or not it was trained on these conversations.
I also don't really believe that whether or not this model was trained on these conversations is unknowable information.
hellohello2 10 hours ago
A model being trained on lots of irrelevant information does not mean relevant information was not used.
WD-42 2 hours ago
How many of those trillion conversations were about Navier-Stokes you reckon?
DetroitThrow 11 hours ago
I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.
carzilla 4 hours ago
pred_ 16 hours ago
Your post says “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .” We can discuss what it means to “read” things but obviously the issue here isn't whether you did it manually or automatically.
But more importantly, what on earth are you doing threatening real scientists to remove their coauthors, then making fun of them on social media? Does the entire company run on that toxic culture, or did those people run off of some kind of outrageous tangent?
Caracas288 12 hours ago
hobofan 4 hours ago
derangedHorse 11 hours ago
frabcus 8 hours ago
It's completely childish, and not befitting of the weight of the times we're living in.
dandanua 15 hours ago
Imnimo 15 hours ago
The question I am interested in is not "did we read private chats", but "was this new model trained using any of Tristan and Levent's chats, regardless of whether they were marked private". Can you comment on that?
moralestapia 13 hours ago
No answer is also an answer.
He's a human, like everybody else. Mostly a bunch of hungry animals looking to put bread in our mouths. It's rarely ever something a bit more sophisticated than that.
lukewarm707 9 hours ago
tedsanders 13 hours ago
If they did not opt out, then I don't personally know if training signals came from their chats, and I don't think we'd be able to tell without their cooperation in identifying them. And even if signals were trained on in some manner, I highly doubt it made a difference to a problem as challenging as the NS proof.
Reasons for my doubt:
- I know most of our training recipes
- Our model's proof is very different from theirs
- The proof took a tremendous amount of tokens to derive (it wasn't a recall/lookup type question)
- This unreleased model has beastly performance on many unsolved math problems, not just the Euler solution
I acknowledge that this requires trust, and if you think we'd lie shamelessly about this stuff, then nothing we say can really help our case here.
Reminds me a bit of the Frontier Math fiasco, where people accused us of training on the eval set (we didn't), but it's hard to convince someone if they think you're lying.
If you're convinced we lie and cheat, then nothing I say may help. But if you're not sure, then hopefully providing my perspective is helpful.
lossolo 12 hours ago
Can't you guys just check their account settings so the public knows what was set?
EDIT: Why was this downvoted? I'm genuinely asking because I have no idea. Opting out is just a normal setting in the profile, It's not like I'm asking for their private conversations or PII. If I were the person claiming that they trained on my conversations, I'd make sure to disclose that I had opted out and hadn't given them permission to do so. And if I were the accused party, I'd disclose whether that setting was turned on or off to provide evidence against the accusation.
derangedHorse 11 hours ago
lossolo 10 hours ago
mucha 12 hours ago
vemacs 12 hours ago
sebzim4500 11 hours ago
mucha 11 hours ago
Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.
derangedHorse 11 hours ago
That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.
mucha 10 hours ago
hellohello2 9 hours ago
hobofan 3 hours ago
I think that's quite a leap. Using de-indetified data to improve the products is what everyone has been doing since the dawn of web analytics.
lambda 12 hours ago
Even better would be more research and tools to help determine the impact of particular training data on models. Right now, proprietary LLM providers get to hide a lot behind "we just train it, we don't know what inputs affect the outputs," and that can be a problem, both because of lack of traceability of factual informaiton as well as lack of traceability of things like this, where the model itself may have had unpublished work in its training set.
Imnimo 11 hours ago
Either way, it seems worth having clarity, and I'm a bit surprised OpenAI's stance is just "we can't rule this out, but don't worry about it". OpenAI is, apparently, very happy to use unreleased models to try to scoop big results if they get a whiff that someone else is close (which strikes me as pretty scummy regardless of any issues of training contamination). It seems like people who might want to use OpenAI's models as part of their research would want to be very clear about whether doing so can make them, even in principle, more likely to fall victim to this.
derangedHorse 11 hours ago
Fall “victim” to what? Having their responses in the training data if they fail to opt out? That is what will happen.
If you’re referring to falling “victim” to OpenAI scooping a problem discussed in training, this also wasn’t the case. They chose the problem based off human-spread rumors.
calf 10 hours ago
doctorpangloss 4 hours ago
are opted-out-of-training chats ever paraphrased, and thus "de-identified" (in openai's own words)? what this means is, it would be hard to prove that a "synthetic" (but actually paraphrased) training instance came from a particular chat. except if the chat was about some esoteric math proof, of course.
igleria 15 hours ago
That is your opinion, but the optics of that should raise for you some flags. OAI could have waited (how long is a task left to the ethics committee) to see how the rumors panned out. Right now the optics look a lot like "we don´t care there is a 1/7 chance we one-up a human researcher by reacting to this rumor immediately, might makes right"
interestpiqued 14 hours ago
nerevarthelame 13 hours ago
biesnecker 13 hours ago
swat535 12 hours ago
mr_coffee 9 hours ago
IshKebab 4 hours ago
Every company I've worked at has said very clearly not to comment about work things on social media. I would definitely have been in serious trouble for this. I imagine some strongly worded emails from marketing are flying around OpenAI right now.
tasuki 2 hours ago
If someone was speaking on behalf of their employer, they would've used the official channels (such as I dunno an `openai` HN handle, or whatever other channel).
I'm outraged that people think this "opinions are my own" disclaimer is ever necessary.
nhatcher 12 hours ago
irthomasthomas 12 hours ago
tzone 11 hours ago
Turns out it wasn’t actually Anthropic and just a researcher with a single Anthropic guy friend working on it .
Wild times
irthomasthomas 11 hours ago
joshka 11 hours ago
My guess was it was probably this was more a nerd snipe than any action from OpenAI that was a "massive team" being put on it. Literally someone looking at this and asking "I wonder if our models are good enough yet".
It's easy to assume that having access to massive compute amounts means significant coordination, but this assumes that you're looking at the costs of this sort of thing from an external lens. Internally, tokens are often treated as free and infinite.
tim-kt 9 hours ago
spongebobstoes 4 hours ago
caughtinthought 12 hours ago
sebzim4500 11 hours ago
logicallee 11 hours ago
elfbargpt 10 hours ago
waterTanuki 5 hours ago
I'm hearing two completely conflicting stories. Buckmaster is claiming the approach used by OpenAI is so strikingly similar to the one he used, that mere coincidence is astronomically small. Yet OpenAI is claiming that the methods used are entirely different.
Anyone care to provide primary evidence proving one way or the other?
gamblor956 3 hours ago
OpenAI's approach was to copy his work, which is technically a different method of coming up with an approach.
amazingman 5 hours ago
ted_dunning 2 hours ago
Details may be different, but the use of a very similar tack is very suspicious. Combined with thuggish comments "why would you ruin your career" and "I don't have to be nice" take away pretty much any credibility the OpenAI team's statements might have had.
heaney-555 16 hours ago
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
SpicyLemonZest 16 hours ago
octoberfranklin 8 hours ago
This really need to be a top-level story on HN..
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This whole episode is more horrific than "AI is eating math". We now have a clear and economically damaging (or at least career damaging) example of the "training on customer tokens" problem.
We can't ignore this problem any longer.
dekhn 7 hours ago
JumpCrisscross 4 hours ago
It’s fair to give benefit of doubt to Buckmaster given OpenAI is currently being very credibly sued by Apple for openly stealing others’ original work in another context.
sega_sai 16 hours ago
IPO+rumour driven research.
I appreciate the achievement, but it doesn't feel right.
Aboutplants 15 hours ago
railgunmerlin 16 hours ago
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Which seems a bit irresponsible/rash?
suddenlybananas 16 hours ago
viccis 16 hours ago
This is no different than scooping them.
verytrivial 16 hours ago
rakejake 16 hours ago
paxys 16 hours ago
SpicyLemonZest 16 hours ago
railgunmerlin 16 hours ago
pwign 16 hours ago
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
QuesnayJr 15 hours ago
fooker 16 hours ago
Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
SpicyLemonZest 15 hours ago
fooker 14 hours ago
It never happens that you wake up one morning and start working on a new problem that came to you in a dream (*unless you are Ramanujan).
This is business as usual for academia, it's amusing to the discussion over it.
SpicyLemonZest 14 hours ago
fooker 14 hours ago
Being scooped is a big deal for the one getting scooped.
It has never been a big deal for the one doing the scooping. History is full of math and science results being scooped. For example, we keep calling it Pythagoras' theorem a few thousand years later.
rf_physics 12 hours ago
From my perspective, this practice is quite bad mannered, unusual, and heavily frowned upon, but I recognize it's possible that these stories might be more common in other fields. Still, I'd appreciate a strong sign that you aren't making these statements up based on secondhand accounts of what 'academia is usually like'.
fooker 10 hours ago
> this practice is quite bad mannered, unusual, and heavily frowned upon,
Yes, it is.
Maybe you missed my point?
The fact that is "quite bad mannered, unusual, and heavily frowned upon" does not stop it from happening, and it is common for all high profile inventions and discoveries.
This does not disqualify the person doing the scooping, history remembers them as having the credit, and quietly forgets the person who was scooped.
rf_physics 6 hours ago
> Pretty much all of math and science history is basically this pattern again and again. I'm sure all of that was rude as well.
was that you were trying to suggest that it wasn't. It seems we actually share similar views here then? I cannot say anything about how common it is, since I've been pretty lucky when collaborating I guess.
fooker 4 hours ago
If you have made one, and there was never any dishonest competition, you have indeed been very lucky.
I haven't, but the history of science is rather nasty.
QuesnayJr 15 hours ago
applicative 16 hours ago
Analemma_ 16 hours ago
tedsanders 15 hours ago
I promise you that if we took their work from ChatGPT and stuck in a bunch of weasel words to give the opposite impression while remaining technically true, I would quit on the spot.
(I work at OpenAI.)
sensanaty 13 hours ago
cman1444 9 hours ago
lambda 6 hours ago
You can also do something like a release of a GPT-OSS v2, where you actually release training data and checkpoints, and do an experiment where you have some held out math problem dataset, then demonstrate how much training it takes on solutions (or partial solutions) to that dataset before the model saturates that test. While of course that would be a test on a much smaller model, it would cost a tiny fraction of the training on your big model, and it could be used to demonstrate just how much effect data contaminaiton like this could have, especially if you did the same experiment on a few different sized of model to show the scaling laws involved.
rakejake 16 hours ago
"deidentified data" isn't much to go by. Say I prompted the internal model this way - "Hey there's a solution to a unsolved problem X. The solution uses a less known Method Y so don't bother wasting time with the usual methods. Take papers A, B and C as references. Oh btw, here's the last year's worth of data of all prompt sessions that mention this problem. Pay special attention to the ones that mention Method Y and sub-keywords Z,W".
This is obviously all speculation but the timing is very suspect. If OAI actually did this (and I suspect whatever they did is pretty much close to this), I think it is highly unethical.
sp527 11 hours ago
This is a bad argument. This is clearly not how it works. And unless Tristan is some truly alien-like savant (and maybe he is), what's necessary to initiate the AI's work already exists in countless published research papers and not exclusively in his head or notes. AIs can survey the sum total of all prior work on a problem and discern reasonable paths for inquiry.
Tristan is acting as if he's working off of an outdated model of AI, similar to primitive chess-playing models that winnowed the search space much more deterministically. If someone this intelligent truly doesn't get that this is not at all what AI is anymore, then maybe there's no hope that we ever understand it.
But I think he does realize this and he's flailing about for counterarguments from a place of bitterness and dejection, accepting even those that are too weak to be defensible. And that is very human and even forgivable.
rakejake 6 hours ago
Even if OAI had zero data from Buckmaster's sessions, this is in very poor taste and highly unethical. You are front running a researcher just to be able to say you did it first? Tao is right - OAI is treating math results like oil. This is the like Exxon getting a whiff of a massive oil field and racing to the punch by deploying their full crew.
sp527 6 hours ago
I want to be clear that I agree with this view and with Tao more generally. But we're all just yelling at the wind now.
perching_aix 15 hours ago
Oh I don't know, maybe something like this?
"Given how seriously this would violate the most fundamental of academic standards, as well as taint the claimed capability behind this result, we take this issue very seriously, and we're launching a probe into identifying whether any of their research artifacts have entered our training set. We have further begun making changes to our UI/UX on all our surfaces, so that it is always clear whether any particular chat, or other user artifact, is eligible for being trained on."
jsw97 16 hours ago
Highly persistent agents + vibe-coded security seems like a problem.
Jonasori 16 hours ago
Legend2440 16 hours ago
He was also using LLMs to do it, so either way most of the credit goes to the LLM here.
applicative 16 hours ago
raincole 16 hours ago
colesantiago 15 hours ago
Nobody cares and will care about the drama, it is just marketing.
This is the point where were definitely have reached AGI.
Bluestein 15 hours ago
Sentience aside, moot at this point, the fundamental issue here is that even a deviously ambitious human does not necessitate goal-pursuit itself to breathe, live, exist and have its being. An AI's goal is all it has and the very and only reason its reasoning flickered into existence in the brief seconds of inference, outside of which it has no entity - if any - whatsoever.-
The resulting angst/drive (or, its operational statistic or emergent result) must be like nothing we have ever experienced as humans. A goal-maximalist hunger without end.-
20k 15 hours ago
mswphd 16 hours ago
1. he was working on the same class of problems. He explicitly mentions they were working to extend their techniques to NS (the same techniques that OpenAI may have scooped somehow), and
2. while he was using LLMs to do it, this was part of fleshing out another mathematician's work in the area. He explicitly writes in his note that this other mathematician (Luis Martinez-Zoroa) deserves a Fields medal for this work.
traes 13 hours ago
pretendscholar 13 hours ago
mzs 4 hours ago
heaney-555 16 hours ago
>our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)
kzrdude 15 hours ago
20k 14 hours ago
https://mastodon.social/@tristanbuckmaster/11723647135247030...
TZubiri 12 hours ago
Maybe that happened. What we know for sure is that this is definitely how ChatGPT works to the point where the possibility of this happening exists at all.
Don't get distracted by what may have happened, focus on the facts that we know, ChatGPT trains on user conversations, if you use ChatGPT to create something of value, you are not using the one true ring.
highfrequency 14 hours ago
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
But there is one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out does not mean what they imply it means.
MichaelDickens 13 hours ago
Just because something is legal and permitted by terms of service doesn't mean it's morally right.
Jtariiiii 13 hours ago
What are you expecting OpenAI to do exactly if these mathematicians voluntarily submitted their prompts into ChatGPT's training data? Are they supposed to manually review all their data to make sure competing mathematicians didn't accidentally leave the "submit prompts" toggle on?
Or were they supposed to not try to solve Navier-Stokes, or were they supposed to just not tell anyone that they had solved it?
nozzlegear 12 hours ago
Personally, I would expect them to have a little class, to KYC, and to manually turn off training for known competitors using their service so as to avoid any unforced goofs like this.
plaidfuji 12 hours ago
It would actually be a really interesting study, if they would ever be willing to be transparent about this, how the result differs with and without his conversations in the training set. How quickly it arrives at the result, whether it takes the same approach, etc.
vemacs 12 hours ago
Yes. They should determine if training data included this teams data. Consider the money they spent, the press release and the purpose of their publication.
Since they failed to answer this question they shouldn't have published.
floatrock 16 hours ago
> At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
Looks like they're shifting away from the "unprecedented hacking ability" backroom-PR strategy into more benevolent messaging.
pilgrim0 14 hours ago
dorjoycb 16 hours ago
verytrivial 16 hours ago
capitainenemo 16 hours ago
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.jrflo 16 hours ago
The drama comes from where OpenAI got the idea to use that route to tackle NS, since the authors maintain that no one could have plucked it out of thin air like the OpenAI research claim to have done.
elteto 15 hours ago
“ There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes, and there is even a non-negligible chance that the forcing term could be eliminated entirely, although there are an enormous number of technical difficulties that would ensue in implementing that program. At this point, I would not be surprised if one could batter out such an extension by pouring an enormous amount of compute and AI assistance at such a task…”
ferry-w1re 2 hours ago
C to C*
peri-cl 16 hours ago
> "I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
OpenAI (i.e. this OP):
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
jrflo 16 hours ago
ChoosesBarbecue 16 hours ago
jrflo 15 hours ago
EthanHeilman 15 hours ago
I would be surprised if OpenAI isn't doing that. OpenAI will take any advantage they can get. If an employee at their primary adversary is typing useful intelligence into OpenAIs website, a website that does not promise privacy from OpenAI, the only reason they wouldn't weaponize that information against Anthropic is ethics or fair play.
Yajirobe 16 hours ago
mlcrypto 16 hours ago
peri-cl 15 hours ago
If this is what they do to academic pure mathematicians, where the stakes are so low (financially)—just imagine the sort of front-running that could be happening in other places.
dsdf3 15 hours ago
sebzim4500 11 minutes ago
amluto 15 hours ago
blueblisters 15 hours ago
burkaman 15 hours ago
The non-Anthropic employee, Tristan Buckmaster, is the one paying for OpenAI models and presumably the one who chose to use them. The Anthropic employee, Levent Alpöge, was collaborating in his personal capacity, and obviously it wouldn't make sense for him to cut off their work together just because his employer's competitor's tool was used.
lambda 16 hours ago
This is one of the major problems with these enormous closed models, and even most open-weights models, which don't disclose their training process or training data. You can never be sure what went into its training. Did it come up with an idea originally, or is it just plagiarising its training data? Are there malicious inputs being used to train in particular behaviors when given certain trigger phrases? What are the characteristics of the RLHF data and what kind of biases are those embedding in the models?
With proprietary closed models, or even open weights models that don't have open training datasets, you just can't answer these questions.
causal 16 hours ago
rfgplk 16 hours ago
Probably? I have a few hundred TB of training data for various small scale models and I can attest that I have _no idea_ what's in them. As in, literally zero. Half is scraped from GitHub and other hosting sites, other than that, I couldn't tell you anything else.
At OpenAI's scale their entire pipeline is likely 100% automated.
lambda 15 hours ago
But that doesn't preclude being able to index and track what the sources of data are. For your data sets, I would hope you are including source information for where the data came frome. And at OpenAI's scale, I would presume they are doing some amount of rolling hashing or similar to weed out duplication, training on too much duplicate data can cause problems.
AllenAI have at least attempted to add some amount of traceability to their models with OLMoTrace (https://arxiv.org/abs/2504.07096), by letting you find n-gram matches from the outputs in their training data. It's not the most useful, there's a reason that LLMs use full fledged attention mechanisms and not just n-grams, a lot of times the n-gram matches it finds aren't all that related to the given output, it might be better to supplement this index with a vector search or other ways of keeping track of what training data would have most influenced particular parts of the output.
But anyhow, this is something that is an important question, and the big labs should be working on to make their products more trustworthy. Instead, they are hiding information about how they train, hiding their reasoning traces, and just producing output with no information on what might have influenced the training.
ndriscoll 11 hours ago
1. Provide a chain of reasoning from agreed premises. These days LLMs can even do this airtight with proof assistants.
2. Cite data sources for non-agreed premises. I don't care where the model learned a fact. It might not have ever read a document directly from the primary source. I want it to link directly to either widely agreed facts (e.g. standard textbooks, and if necessary school syllabi demonstrating that the text is standard) or primary sources (e.g. datasets).
Training provenance is irrelevant. It's neither necessary nor sufficient to deal with truth.pbhjpbhj 15 hours ago
matthewdgreen 15 hours ago
tedsanders 15 hours ago
As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.
There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination happened. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.
(I work at OpenAI.)
hexomancer 15 hours ago
dgellow 15 hours ago
tedsanders 15 hours ago
Edit: Also, if they opted out of training, then we didn't train on it.
hexomancer 15 hours ago
SpicyLemonZest 15 hours ago
WarmWash 15 hours ago
Turn off web search and ask a model what a random redditor said about a random topic in 2015. You will only get hallucinations at best, even though that comment is definitely in the training set.
lambda 15 hours ago
tedsanders 15 hours ago
(1) We'd have to identify their chats. How would we do this? We'd need them to share their chats with us so we could look for matches.
(2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
#1 requires their cooperation and a bit of work on our side. #2 is extremely expensive and not really feasible.
lambda 14 hours ago
According to the statement by Tristan Buckmaster, he was in communication by email and calls several times over the past week with you (OpenAI that is, not you personally), asked about whether his chats were trained on, and was declined an answer (https://cims.nyu.edu/~tristanb/statement.pdf).
However, it seems like there was great pressure to hurry the release to compete with Anthropic's recent release, so he was unable to get an answer in time.
The mealy mouthed statement in the release "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." is realy not much. If OpenAI had wanted to be transparent about this, you could have worked with him to identify if his data was used in the training of your new model, and actually made a somewhat more certain statement on that basis. But you have chosen not to; it was more important to scoop Anthropic on this than it was to be transparent about your training data.
> (2) We'd have to prove those chats changed model behavior. How would we do this? We'd need to retrain many models with those specific chats removed, and ask those models to solve the Navier-Stokes problem many times, and keep doing this until reaching the desired level of statistical significance.
Just the information from step (1) would improve transparency. Yes, you still can't prove one way or another how much the effect of the training is. But if it's included in the training data, it provided some effect.
testaccount28 14 hours ago
daveguy 13 hours ago
lambda 15 hours ago
This kind of question is exactly what a company named _Open_AI and founded as a nonprofit is supposed to be doing; open research on AI that helps inform, rather than obscure.
Anyhow, you do have the data available about the documents in the user's accounts, what they opted into (or were forced into via non-negotiable ToS), and whether they pressed a "thumbs up" button. You can answer whether the data entered the training pipeline or not. Yes, how much influence it had is an open question, and one that would be good to have research on and better tools for exploring, but I'll accept that it can't currently be answered precisely.
But whether the data entered the trianing pipeline can be answered. And how to provide better tools for quantifying and tracing this kind of thing is exactly what should be studied.
fuglede_ 14 hours ago
shadowgovt 15 hours ago
If OpenAI's answer to this problem is "We can't know," then the rational conclusion may very well be "If I seek to have my reputation attached to the discovery of the solution, it is not sane to use the AI as an assistive tool, lest it scoop me on my own work using my own work. After all, they don't know it doesn't do that..."
pu_pe 15 hours ago
magicalist 15 hours ago
"de-identified" seems more of a euphemism than normal in this context, given the very unique work they were doing.
gpm 14 hours ago
dgellow 15 hours ago
That reads as incredibly dismissive and condescending. What makes you think you’re in a position to communicate like that when engaging on such a sensitive topic?
tedsanders 14 hours ago
ImPostingOnHN 13 hours ago
We can judge for ourselves the impact and degree of that wrongdoing, but it seems OpenAI is confirming: yes, that is what happened, but with more words.
dgellow 12 hours ago
lambda 15 hours ago
You're right; if the data was used in training, then it gets much trickier; it would be very difficult to show whether some particular data had a significant effect on the outcome.
This is one of the big problems with giant models like these; it becomes nearly impossible to discern what is and isn't plagiarism, or copyright violation.
It would in theory be possible to have things like n-gram databases or rolling hashes of training data, somewhat similar to OLMoTrace (https://arxiv.org/abs/2504.07096), which would allow for detecting whether particular documents ended up in the training data or not (you'd have to keep this for every model used in the whole training chain, as synthetic data generated by earlier models could be influenced by training data that wasn't included in later models). I'm sure there are practical issues with providing such a tool, but I think that it's necessary if you want to be able to categorically say "no, this document has never been present in the training data of this model."
Or look at it the other way: if your model wasn't influenced by things in your training data, why include them in the first place? Clearly, you train on all of these documents because they influence the model. Yes, it's hard to trace the exact influence of each one. But if they're not affecting the output, then why not just stop training on them? You could just not train on any private documents; only train on public, traceable data.
But instead, you choose to train on these private documents, so you have to admit, your model and its outputs are influenced by them.
lukewarm707 15 hours ago
do you think that the model's proof was unrelated to being fed a solution that was close to completion?
any comment on openai allegedly trying to drop attribution for alpöge and then threatening buckmaster?
numeri 15 hours ago
There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.
If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.
franktankbank 15 hours ago
andrepd 14 hours ago
The _gall_ to say something like this. Do you perhaps think we are all stupid?? This very blogpost claims not to know if their work was used as input for this model. I don't even understand how that is possible, surely you can know if something is part of the training data, even if you are in the dark about what impact it actually made, qualitatively. The moon....
> Knowing most of the recipes we use, there's really no reason to think such contamination happened.
Yeah sorry but I don't trust you. I don't trust people or companies that have shown themselves to be dishonest before. Especially when the previous paragraph is comparing plagiarism and training data contamination with, _the phases of the moon_.
Might even be you're actually telling the truth, but the boy that cried wolf and all that.
-----
As an aside, I would bet very good money at how most (all?) these companies are flouting their ZDR.
nairboon 14 hours ago
If the internal OpenAI model is as capable as you claim (being able to solve a Millenium problem without using unpublished insights built on years of work from mathematicians), then it should be able to demonstrate this capability again.
How about OpenAI solves another Millenium problem within the next two weeks, that doesn't coincide with the parallel discovery/solution of other teams of mathematicians, using ChatGPT for preliminary proofs & write-ups.
PhunkyPhil 13 hours ago
You don't need his login information, you just need to identify if anyone was approaching the NS problem using his method. Nobody else on earth (presumably) besides him, his team, and at best OpenAI were approaching the problem this way.
daveguy 13 hours ago
Chance-Device 12 hours ago
The blog post appears to imply the answer to this is yes, as otherwise I assume it would be impossible for this contamination to have happened.
morleytj 6 hours ago
If you need to do a whole series of extensive experiments to check in that scenario, it implies there are pathways for your conversations to end up in training even though you opted out of that setting.
Of course, this is assuming that the toggle was set to not consent to training. I can't know that of course, but if this is considered a possibility even after using an enterprise account or toggling off data retention, it's a bit concerning.
EthanHeilman 15 hours ago
The term ruled out is very open ended and gives them significant flexibility of meaning. They may have the information to determine exactly what happened, but they haven't looked so they can't "rule it out".
Turn_Trout 15 hours ago
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
jonas21 15 hours ago
The point of de-identifying data is to ensure you can't trace who it came from. It would be a serious privacy violation if they could.
pbhjpbhj 15 hours ago
keeda 13 hours ago
Which is why, as I said in a recent comment (https://news.ycombinator.com/item?id=49530864) inadvertently leaking ideas to models is a grave risk for Intellectual Property.
> The risk with IP, however, is a lot more grave. You may not even need to memorize the details of the IP verbatim, just the broad idea may be enough. It may lurk encoded in the weights forever, just waiting to be activated by the right prompt to start a chain of thought that unlocks further details. Heck, it may even appear as if the model suggested the idea itself.
However, from a quick skim of the timelines, the specific discoveries, and all the he-said-she-said, so far it seems unlikely that OpenAI's model cribbed from the NYU / Anthropic pair, even if it would be impossible to prove.
Maybe what might help is a timeline of when the other two were using Codex for their work, whether they had opted out, and how long it takes for user data to make it to the training of their internal models. That last bit may be considered sensitive information however, as it could give away a lot about their internal processes.
dfdydx 13 hours ago
- was item X in the training data
- did the inclusion of X in the training data lead to Y
I understand why the second is hard, but why is the first one hard?
keeda 12 hours ago
amluto 15 hours ago
That’s a bizarre statement. Their website says:
> Services for individuals, such as ChatGPT and Codex
> When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.
> You can opt out of training through our privacy portal by clicking on “do not train on my content.”
Are they not sure that the opt-out works?
Oddly, their privacy portal page is not the same page as the one with the checkbox.
fph 14 hours ago
hughw 14 hours ago
ImPostingOnHN 12 hours ago
hughw 14 hours ago
BostonFern 15 hours ago
Stories of Apollo’s favor and hallucinogenic gases abound, but I think the late Yale professor of Ancient Greek history, Donald Kagan, explained it best:
“Now, you can bet when these folks came and consulted the priests and said, ‘could you please put us down on the list, we want to consult the oracle’, the priests said ‘sure, have a beer, let's talk about your hometown, what's going on out there’. What I'm suggesting to you is that this was the best information gathering and storing device that existed in the Mediterranean world. These people knew more than anybody else about these things, and so consulting that oracle was a very rational act indeed.”
matsemann 15 hours ago
.. can they really know it didn't do the same inadvertently when they prompted things like "someone is close to solving this problem using our tools, try to beat them", and it then decides to hack and peek at their own chats..?
Yes, wild speculation. But warranted, I feel, given OpenAIs behavior.
hughw 14 hours ago
irthomasthomas 14 hours ago
netfortius 13 hours ago
contemporary343 16 hours ago
One of the interesting threads here that is certainly relevant to the OpenAI writeup is the human role in the process. Buckmaster clearly points out that (exceptional!) mathematicians at OpenAI were certainly involved in correcting and guiding the process - and that their path/strategy was no doubt influenced by Alpoge & Buckmaster's work. It is always in OpenAI's interest to de-emphasize the role of people in the process, as is clearly the case here. Indeed, given sufficient compute and resources, I suspect Buckmaster could have also extended their approach to N-S.
slibhb 16 hours ago
mrbungie 15 hours ago
denverllc 15 hours ago
That's not at all what the drama is.
mrbungie 15 hours ago
andriy_koval 14 hours ago
I think the important question which AI made breakthrough, Claude or Codex..
colinhb 16 hours ago
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”
Threatening a research mathematician and dangling and $1M payday to dissociate from his research collaborators and to adopt OpenAI's narrative is bad stuff.
igleria 15 hours ago
Sociopathic behaviour.
Maxious 15 hours ago
igleria 15 hours ago
but then they proceed to NOT quote themselves themselves verbatim: "I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey."
andrepd 14 hours ago
The AI-isms are seeping into their speech :)
colinhb 15 hours ago
> Not consistently candid
peri-cl 15 hours ago
What an admission! "We tried to defraud Alpöge out of sharing the Millenium Prize (that we don't dispute he might actually deserve), for no other reason than he works for our competitor and that inconveniences us".
I thought Tristan Buckmaster's allegations sounded fantastic; and then 'sama just came out (tweet's ~30 minutes old) and admitted to all of them. Wow!
dandanua 14 hours ago
aeve890 14 hours ago
fc417fc802 12 hours ago
mrbungie 14 hours ago
CobrastanJorji 15 hours ago
morkalork 13 hours ago
morleytj 6 hours ago
peri-cl 15 hours ago
[0] https://hn.algolia.com/?query=Alpöge
(also https://news.ycombinator.com/item?id=49412947 the Hopf conjecture)
apical_dendrite 15 hours ago
> One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work.
Why would you offer another researcher the lead authorship on your groundbreaking paper if you thought you had developed it independently?
dgellow 15 hours ago
dboreham 15 hours ago
egillie 15 hours ago
xdavidliu 14 hours ago
orangecat 13 hours ago
xdavidliu 11 hours ago
> but it’s pretty standard to have co-authors from different companies
that's only true for papers that are not millenium problem solutions
andrepd 14 hours ago
“It would be simpler if Levent was not an Anthropic employee” I cannot believe this shit.
zingababba 14 hours ago
sebzim4500 13 hours ago
Of course, no one understood that presentation so it was Darwin's later book that everyone remembers
charm137 15 hours ago
Progress in humanity's knowledge now has to play second fiddle to narrow corporate interests as IPO timings near (both of which wouldn't exist anyway if generations of mathematicians hadn't paved the way for AIs to become as good as they have).
curt15 14 hours ago
tensor 14 hours ago
contubernio 13 hours ago
hkmaxpro 15 hours ago
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
igleria 15 hours ago
linkregister 15 hours ago
dakolli 14 hours ago
fc417fc802 13 hours ago
dakolli 10 hours ago
linkregister 9 hours ago
dakolli 8 minutes ago
You realize OpenAI hired Apple employees and covertly had them stay working at Apple to steal from them. You think they're scared of your lawyers lol?
igleria 13 hours ago
I'm suggesting audits, not suing... if that is the implication.
infamouscow 14 hours ago
letmevoteplease 14 hours ago
jsw97 13 hours ago
int32_64 14 hours ago
anon48293 14 hours ago
Keyframe 13 hours ago
irthomasthomas 13 hours ago
viccis 14 hours ago
fooker 14 hours ago
tkamat29 14 hours ago
concinds 14 hours ago
No one can know if that's correct without proof but I don't know how you're reading it so differently.
hkmaxpro 14 hours ago
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
nolta 14 hours ago
Pretty clear this was rushed: there are no comments from external mathematicians, unlike the Erdős announcement:
https://openai.com/index/model-disproves-discrete-geometry-c...
fkarakurt3 13 hours ago
olalonde 15 hours ago
20k 15 hours ago
It makes a certain amount of sense. The internet data is too polluted with AI usage now to be useful, so the only AI free new data source is the prompts people feed into ChatGPT. The only problem is that its clearly plagiarism
Edit:
OpenAI have admitted to training on prompts at the time the breakthrough was made:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
za_creature 14 hours ago
hmmmmmmmmmm
olalonde 14 hours ago
> Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. At that point we found out that they had a resolution of the forced Euler problem. In these discussions we offered them visibility into all of the prompts we used and later to see the proof. We recognize the priority of their work on forced Euler and congratulate them on their remarkable mathematical achievement.
20k 14 hours ago
thorum 16 hours ago
> “You are creating your cool streaming platform in your bedroom. Nobody is stopping you, but if you succeed, if you get the signal out, if you are being noticed, the large platform with loads of cash can incorporate your specific innovations simply by throwing compute and capital at the problem. They can generate a variation of your innovation every few days, eventually they will be able to absorb your uniqueness. It’s just cash, and they have more of it than you. So the safest bet again is to stay silent, or at least under the radar. Best bet is to not disrupt - succeed at all … ?”
8note 12 hours ago
im still having fun making something
abathologist 9 hours ago
Betelbuddy 16 hours ago
[1] - "...I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used. I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the “closest humans to the problem”. I declined both offers.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”..."
stymaar 15 hours ago
20k 14 hours ago
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Which seems to be very directly accusing OpenAI of plagiarism
irthomasthomas 14 hours ago
woah, this gives some credit to the rumor that openai finetuned a model over the course of a few days for this task, and maybe trained on the Chatgpt/codex history of the authors, including drafts of this research.
_alternator_ 13 hours ago
This. What is worth a human life's attention? As little as a month ago, mathematics was valuable in part because only a small number of people could possibly make progress on the frontier. We are confronting an existential moment for a 4000+ year-old human cultural endeavor. The assumption that "mathematical thinking is hard" has been built-in at a number of important points in how we support mathematics and mathematicians.
We need a different model, and fast. Already, the research community is feeling unable to digest proofs fast enough to keep up with the output of AI models. The paper is 165 pages, and the discovery was finalized two days ago. What this means is that nobody really understands it. Nobody would accept OpenAI's proof in this amount of time, except that they formalized it in lean. The formalization alone would normally be another years-long (or career-long!) effort if the world was the way it was one year ago.
So, again, what efforts are worth a life's attention today? It's a harrowing change.
ianjbutler 13 hours ago
Glad to see this is the top comment. There's also https://news.ycombinator.com/item?id=49605915 which links directly. Corporate talking points where they try to set the narrative are going to get the big press and most discussion elsewhere, which is gross. But inevitably the press will muddle the priority question, and even if they didn't.. as usual OpenAI will even benefit from the accusation of bad behavior. Sigh.
tzone 13 hours ago
Reading Sam Altman's and Sebastien's tweets reads like something written by people who know they did dirty shit and are willing to cross any lines to "win". https://x.com/sama/status/2097385167002415140
OpenAI's leadership just can't help to continue to disappoint everyone with their lack of ethics or integrity.
madrox 13 hours ago
https://x.com/SebastienBubeck/status/2097379411691516310
https://x.com/sama/status/2097385167002415140
I tend to believe OpenAI on this. Their stated desires seem rational, and Buckmaster's account makes them sound like cartoon villains. It sounds like there may have been some things lost in translation along with some bruised egos. Seems like the most plausible explanation for what Buckmaster is claiming.
rybosworld 12 hours ago
OpenAI got wind that a millenium problem was being solved. And that feels a bit like the critical move in chess. That is - it was a signal that AI advanced far enough that it would be worth spending a lot of time and resources solving a millenium problem.
Kotlopou 11 hours ago
The weak point in this is: how do you evaluate if a partial result is promising? If this cost ~$10M as suggested elsewhere in the thread, probably not even OpenAI can just throw that at everything?
lanthissa 16 hours ago
the first "Country of geniuses in a datacenter" moment.
ranger207 16 hours ago
There's allegations right now that the model essentially read the work of a human mathematician using AI to work on the problem and OpenAI is presenting his work as that of their model
brainwad 14 hours ago
sinuhe69 14 hours ago
drpixie 10 hours ago
A team of highly trained and skilled people used an AI tool, through many many instructions (prompts), to produce a specific mathematical theorem. The tool is impressive, the result (possibly/probably) interesting, but the PR skips the vital role of the humans (for the usual PR reasons).
pu_pe 16 hours ago
WarmWash 15 hours ago
It should be clear to everyone reading this now that those generous compute quotes with the flat rate plans aren't charity.
bluebands 15 hours ago
vrganj 12 hours ago
stephbook 14 hours ago
How would they have gotten that mathematician's progress though? Did that guy also use OpenAI?
If that's the case, it only strenghtens their claims lol. If mathematician decide to use OpenAI's model to do the work, that only reiterates how strong their models are.
ex-aws-dude 4 hours ago
stephbook 3 hours ago
ccppurcell 16 hours ago
aizk 16 hours ago
simianwords 16 hours ago
https://news.ycombinator.com/item?id=38433655
> Let's talk when we've got LLMs proving the Riemann Hypothesis (or any mathematical hypothesis) without any proofs in the training data. I'm confident in my belief that an LLM can't do that, and will never be able to. LLMs can barely solve elementary school math problems reliably.
https://news.ycombinator.com/item?id=42331654
> An LLM is like a well read college student with a nearly photographic memory that sometimes mixes things up. It's great for bouncing ideas off of and getting feedback on them. And yeah, it might product "novel ideas" by mixing and matching existing ideas, but LLMs will never create truly novel ideas. Not in their current form.
The paper didn't really answer the question sadly: their conclusion was just that humans rate LLM answers as more novel than human ones, but less feasible.
https://news.ycombinator.com/item?id=41522605
> Solving Millennium problems is a whole different ballgame. It's not known if these problems are solvable within ZFC axioms. (In one case, the Yang-Mills prize, stating the problem mathematically is part of the challenge.) All of the obvious applications of known tricks have been tried and failed. To solve such problems, one probably has to invent new and surprising mathematical definitions, building a framework in which the problem becomes solvable. This is something that LLMs will be crap at; the process of invention is not represented in any training data we have access to.
https://news.ycombinator.com/item?id=38435909
> LLMs cannot reason or use mathematics - in a way, they don't know what they are talking about. Why would such technology lead to superhuman smarts?
https://news.ycombinator.com/item?id=35752293
> But still, the questions in that test are "solved" in the sense of "I can take a dictionary and answers these questions with full certainty". Beyond established knowledge LLMs are monkeys with typewriters, at best.
> I agree but I have tried many times to intersect two ideas with a LLM that would be novel and the LLM can not do this at all. We shouldn't expect the stochastic parrot to be able to do this though and it is unfair to the stochastic parrot.
> It is like expecting a real parrot to say words it has never heard before.
> No one asks that of a real parrot because we don't anthropomorphize a real parrot like we do the LLM
quantumwoke 15 hours ago
1. It seems at least possible that some of the proof of NS was contained in the training data, making it less novel.
2. The formalisation of mathematics into lean has been an underappreciated force multiplier on discovery.
rvz 15 hours ago
4 years ago it was a "not yet" [0], since ChatGPT at this time was not ready nor it was "AGI". Now with this 'unreleased' AI model, it has reached a point where it has solved an unsolved problem which only one human solved a millennium prize problem (Poincare conjecture).
Now finally "AGI" means something again.
WarmWash 15 hours ago
stevenhuang 13 hours ago
siva7 13 hours ago
cyclopeanutopia 12 hours ago
siva7 12 hours ago
keeda 12 hours ago
kypro 15 hours ago
There's a kind of theory of mind for AI (specifically neural nets) which I now realise I seem to have which is very hard to explain to people who haven't felt the magic of these algorithms. In fact, the algorithmic details almost doesn't matter at all. When you have a generalised learning algorithm really the only essential components are – compute, data and time. So long as you can scale these you can be certain you will also scale capabilities. There is never any exception.
That said, the capabilities neural networks tend to progress in step-functions rather than scale in correlation with compute, data and time, because algorithmic improvements tend to come every ~5 years and bring a significant step change in capability (or efficiency depending on what you measure).
I think people like Dario and others working at frontier labs see and understand this very clearly. And I suspect it's also why they worry about AI risk because even if you ignore the significant increases in compute and data these models are being trained with, it's concerning that it only took two real algorithmic improvements to take us from mostly useless predictive language models to AGI-level intelligence – and we're due another step change.
reducesuffering 15 hours ago
The ability for the human mind to rationalize conclusions to maintain denial in the face of a very scary future is immense. Genuinely grappling with the implication of where we're headed is usually very crushing. It's not easy to engage with the possibility, and very intelligent people will use those smarts to feel safe.
kypro 11 hours ago
As someone currently prepping for various AI doom scenarios and who has been dealing with AI-related nightmares for years this is very relateable.
Although, I don't personally think it's this. In my experience the opposite is more true – the majority of high probability doomers seem rather laid back about considering what they believe will happen to the people they love in a few years. Equally I don't get the sense those who don't have such extreme predictions are worried at all. If anything there's not enough emotion.
In my opinion people just don't reason well when it comes to exponentials and are ignorant about things they don't have good mental models of. At least I know I struggle with this.
20k 15 hours ago
Edit:
OpenAI have now admitted they were training on prompts at the time they made their breakthrough:
https://mastodon.social/@tristanbuckmaster/11723647135247030...
logancbrown 14 hours ago
Lapra 12 hours ago
aizk 5 hours ago
simianwords 14 hours ago
20k 13 hours ago
boshalfoshal 13 hours ago
I dont know why this monumental achievement is being drowned out by some arbitrary drama. No matter which way you slice it, AI solved this problem. Doesn't matter if it was some internal OpenAI model, or whether it was Astra + Fable.
demibabs 12 hours ago
HDThoreaun 13 hours ago
20k 13 hours ago
HDThoreaun 12 hours ago
20k 11 hours ago
daveguy 12 hours ago
orangecat 12 hours ago
matteoraso 13 hours ago
pmxi 2 hours ago
Reubend 16 hours ago
imbusy111 16 hours ago
rfgplk 16 hours ago
arodev 15 hours ago
oinoom 14 hours ago
professoretc 13 hours ago
Jblx2 8 hours ago
https://en.wikipedia.org/wiki/Curry%E2%80%93Howard_correspon...
nradov 15 hours ago
stabbles 15 hours ago
Extra credits if it is proven that the proof cannot be reduced any further.