bob1029 3 hours ago

> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...

I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.

I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.

adg33 3 hours ago

It's something that happened before LLMs - multiple discovery. Calculus is a classic example.

jansport123 3 hours ago

Yes but in this case, the allegation is Leibniz literally looked into newtons notebooks

skylurk 3 hours ago

> I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days.

If you read the history of major scientific discoveries, this has been the case for a long time. There are many things that were independently discovered by different people at nearly the same time. Once people know something is solved or solvable, it gets a relentless amount of focus.

eru 3 hours ago

This seems to be the norm rather than the exception.

On a tangent, the genius of people like eg Einstein is not so much that he came up with all these things: other people were close, but that he was a singular individual that did all of these discoveries, instead of five different guys all making some breakthrough here or there.

teiferer 3 hours ago

Maybe that shows how scientific discoveries come to be. It's not a genius sitting alone in their chamber for a decade and then suddenly they emerge with this huge thing. That's Hollywood fiction. Scientific progress is the colaborative effort of countless researchers over long periods of time, communicating, exchanging ideas, many of them wrong, tweaking, trying, thinking, arguing. When a breakthrough happens then it's the tip of a mountain of work that came before it. If two individuals stand on that mountain and feel there is something somewhere then it's not too strange that they take the last step at roughly the same time because conditions were right. The preconditions were in place at that time, the results required for this were available and the focus was on this specific thing.

I'd say it illustrates well that this last piece, the person celebrated for the achievement, is disproportionally overvalued and the rest of the work they are standing on is disproportionally ignored.

a_bonobo 3 hours ago

This ties back into AI, too, right?

I always like to bring up how many decades of research, how many hundreds of years of entire PhD-theses, how many sleepless nights were used up to generate all the protein structure data that made up the corpus of Protein Data Bank - that was then hovered up by the AlphaFold team, and guess who got the Nobel Prize...

teiferer 2 hours ago

Just like all non-AI related nobel prizes before them.

Every laureate stands on the shoulder of giants which is their field as a generations-spanning body of researchers. They did the last step and get recognized. The good ones acknowledge that in their acceptance speeches.

avs733 3 hours ago

If you believe the totality of the document, there was more shadiness in how OpenAI acted than just timing. Save other things they are accused of, the progression from rumors to replication would attract much less scrutiny. With those in mind timing begins to look suspicious at best.

Would anyone be surprised if major model companies had tagged the accounts of competitor employees for extra tracking? Given the concerns about distillation and bench marking it hardly seems irrational, but how it is used matters quite a lot.

sodic 2 hours ago

I have a similar story, but perhaps even stranger.

I work for a startup. We often bring a wooden arcade with us to conferences as a marketing gimmick.

The arcade runs a single side-scrolling video game. You're running from a monster and dodging obstacles. The goal is to survive as long as possible, and your result is measured in meters.

There are always a few competitive guys who spend the entire conference taking turns to play it. And every single time, the same thing happens.

Say the current high score is around 200m. Everybody fails somewhere around that number: 190m, 186m... Maybe someone manages 210m. And the high score moves up at a snail's pace.

Then, a new guy shows up and gets something like 500m on his third try. From their next turn on, everybody easily does 450 or more, even though they were struggling to get past 200 just one turn ago.

What makes it stranger is that the game is dead simple. It's not like the new guy discovered a move that unlocked this capability. And it wasn't a lack of motivation either - they'd all been playing for an hour already. They just started performing better after seeing it was possible. There has to be a name for this phenomenon.

KeplerBoy 2 hours ago

I wonder if it's the same with the sub 2 hour marathon which was broken this year (it certainly was with the 4 minute mile).

aswegs8 2 hours ago

Reminds me of amateur table tennis. There is a strange dynamic where you down regulate your performance unconsciously when the opponent is playing worse and vice versa.

Could be described as some physical form of this effect: https://en.wikipedia.org/wiki/Asch_conformity_experiments

The term would be conformity / normative social influence.

Mawr an hour ago

Surely this has nothing to do with specifically table tennis, nor that it's amateur.

A better example would be mixed boys/girls sports classes in school, where the boys deliberately hold back as to not injure/scare the girls.

It's a pretty obvious and human thing not to go out and completely destroy a much weaker opponent. We're social animals after all.

There also may be an element of energy conservation, there's objectively no need to put in any more effort than necessary. Inefficient.

suls 2 hours ago

Reference-dependent effort ("bunching" around a target) [1] is probably the closest thing. The existing high score was acting as everyone's reference point. Anchoring [2] is the more general cognitive version.

[1] https://www.nber.org/system/files/working_papers/w20343/w203... [2] https://en.wikipedia.org/wiki/Anchoring_effect

pred_ 2 hours ago

That's the other thing. Even if it somehow magically turns out that they didn't plagiarize Buckmaster and Alpöge, that an independent audit goes through all their processes and finds that they could not possibly have stolen anything, and ignoring their outrageous attempts to force a coauthor off the author list, and ignoring that their employees act like small children on social media, this is a company whose culture is “we heard of a breakthrough being possible, let's scoop it”.

Normally, when you tell a coworker that you're wrapping up a result, unless they're some kind of sociopath, their natural inclination would not be to try to steal it from you.

harhargange an hour ago

Not the hope, but telling them that an answer does exist. This ofcourse means now we are going to go into an even more darker cave next time we are looking for some gold and pathbreaking discoveries will become even more rare. Add to that, the fear of people not winning against ai and having fewer rewards, then fewer people even enter those fields or attempt problems over the next generation. AI erodes skills not at individual but at civilisational level.

burrish 3 hours ago

>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?

But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."

feverzsj 3 hours ago

It's almost as if it was actually found by manually written brute-force algorithm running on OpenAI's massive computer cluster.

jansport123 3 hours ago

I’m not a mathematician so take this with a grain of salt. Apparently terry tao commented that the approach used for the Euler paper can “probably” be used for solving NS but it’s still technically challenging and can probably be done with an LLM with a lot of compute. To me the crux of the issue is whether the insight were stolen so that the problem becomes something that is in the domain of LLMs. This is much different than LLMs coming up with the insight. OpenAI wants everyone to think the LLM came up with the insight and solved the thing by itself even though they have perhaps an army of researchers.

CSMastermind 3 hours ago

Having the chat logs enter the training data and having them have a meaningful influence on the ultimate result the model produces are very different things.

The text for all the Goosebumps books are certainly in the training data and to some small amount influenced the solve. But their contribution was so vanishingly small it would seem absurd to say R L Stein should have recourse for contibuting to the solve.

pbmonster 2 hours ago

But this is different, right?

The equivalent would be taking a (fully offline) LLM and asking it about the ending of one specific Goosebumps book, and it revealing the twist. And although that specific book was (probably) only once in the training data, a high parameter LLM can usually "remember" the twist.

pietz 2 hours ago

I find the claims from OpenAI somehow more relatable and reasonable.

- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.

- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.

- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.

- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.

piker 2 hours ago

It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.

[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]

pietz 2 hours ago

Why in the world is that fishy?

Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.

piker 2 hours ago

If they were willing to spend 105 million it, perhaps. But that would be surprising. It’s fishy because they spent something like 15 million on the right one.

aswegs8 2 hours ago

Again, that is the point. There is rumours that this one thing would be solvable, so they focus on this, spending 15m on one specific thing, instead of 105m on many. How is that fishy?

piker 2 hours ago

I understood the rumors (as described by OpenAI) to be that “one of” the prizes was solvable. If they heard NS specifically was solvable and aren’t saying that, they’re intentionally obfuscating that.

pietz 2 hours ago

Because that’s the one Anthropic was rumored to have solved.

piker 2 hours ago

"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."

pietz an hour ago

That doesn't reject my claim. They just didn't name them in this post. It feels like, you're going through great lengths reading something into this.

piker an hour ago

No, sorry I was reading it correctly but they addressed the issue in the post itself. They make it clear they were aiming at all 7 problems looking for the "two" that were rumored to be solved. But they didn't go hog on NS until they made progress. See sibling comments. I misread it the first time to be a claim that they "heard one of 7 human intractable problems are solved and spent 15 million on the right one".

MrToadMan 2 hours ago

Didn't they say that they launched an effort to evaluate their model on all open Millennium Prize problems, and then narrowed down to Navier-Stokes as the most promising after seeing results on simplified versions of the problems?

piker 2 hours ago

Yes, you're right. It seems like they claim to have started broadly and narrowed it down based on some progress. That leads to different questions but does answer my initial "fishy" point.

civvv 3 hours ago

LLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.

If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?

krona 3 hours ago

Extreme temperature levels (>2.0) can push a GPT of its manifold, essentially producing predictions barely distinguishable from random noise (it flattens the probability distribution of the next token). In theory this could predict anything including the next field of mathematics (infinite monkey theorem) but realistically that would never happen.

However, how to we know the next field of mathematics isn't a novel combinations of several other sub-fields? That level of mathematics would be indistinguishable from magic to most people and so in their eyes the GPT did something truly inventive.

eru 3 hours ago

A lot of what humans do is combining old ideas.

And an LLM could in theory also stumble upon entirely new ideas: there's randomness in how they generate their reasoning and answers after all.

I suspect that we are seeing a lot of advances coming from the combination of existing but somewhat obscure knowledge coming from LLMs at the moment, because LLMs are really good at this. At least compared to humans.

Even before our AI friends became good, they were already known for having read approximately every paper and every textbook published in any language. You only need to increase intelligence a fairly small amount from there to get to something like the 'convex hull' of human knowledge.

Compare https://slatestarcodex.com/2016/11/17/the-alzheimer-photo/

The gist is that basically whenever anyone comes up with a new method you get a big burst of activity of picking up all the now lower hanging fruit, that was previously out of reach.

lhd1 2 hours ago

This is also what I've been thinking. The result itself is amazing but it's not like this was completely unexpected. There has been a huge amount of progress on the problem in the last 10 years without which it seems unlikely today's full resolution would have been possible. It is not clear what strategy was taken but it sounds like it borrowed heavily from the two spanish mathematicians. Experts will scrutinize the proof and it will be interesting to see if anything truly original or unexpected was done, outside of known techniques, a move 37.

bonplan23 an hour ago

If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt a human mathematician would be useful at all on their own.

civvv an hour ago

How have we got to the place we are today then? Someone must have made the first steps onto uncharted territory, otherwise we would be in a homogeneous state frozen in time.

I am not saying that LLM intelligence can not be the same, that they are uncapable of dicovering new fields/problems that they have no training on. I am just pointing out that historically it kind of "must" be true that humans are capalbe of this, but we have yet to see an LLM do something like this, something radically "new" in a sense. All of these breakthroughs appear to me (not a mathematician) to be more a case of "digging" through millions of existing attempts/work, patching it together into a result.

This would already make LLM's one of the greatest tool mankind has ever made, but it has yet to display what I would consider a necessity for human level intelligence, which is this ability to discover entirely "new" things.

Would an LLM, given enough time and only the currently available trainingdata with no further input from humans, be able to solve something that was discovered tomorrow?

For humans my answer would be: maybe, probably, because this has been done historically.

For LLM's I would not be comfortable in claiming that they could. I think they would not be any better at this than traditional computational bruteforce.

tosh 4 hours ago

> My two favourite hypothetical questions regarding this used to be:

> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)

> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?

> My new preferred hypothetical for this is:

> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?

weinzierl 3 hours ago

Maybe just a rumor of a high value target having their API keys accidentally consumed in the context...

feverzsj 3 hours ago

LLM can't be trained that easily. More like actual human are checking your logs and stealing valuable things from you.

grey-area 3 hours ago

Or searching anonymised logs for mentions of this problem and using that as part of the context or training.

This would work just as well and have plausible deniability.

rakejake 2 hours ago

Exactly! This is the real Occam's Razor explanation.

rzzzt 3 hours ago

They wouldn't appear in weights but could be added to the context. My conversations regularly go "regarding your Java problem"... which was a separate item in the history from earlier. As long as I only see these (and nobody else sees mine), it can be helpful.

FeepingCreature 3 hours ago

LLMs can learn from one sample.

geraneum 3 hours ago

Having in mind the allegations by Apple against OpenAI, I don't find it unthinkable that there could've been some form of misconduct happening there.

shellfishgene 4 hours ago

What's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?

dboreham 3 hours ago

They weren't working in secret?

tyre 3 hours ago

People talk. If you have a breakthrough solving one of the most famous problems outstanding, you're going to tell people.

You'll say, "Don't tell anyone", which they will ignore because they get a rush and perceived status by sharing it. So then they tell someone, along with "Don't tell anyone", etc.

It's a small enough world (both in academic math, one at Anthropic) that you get to OAI in very few hops.

dguest 3 hours ago

The reality is also that most of your colleagues have no interest in stealing your work: they have their own work to do anyway, and having a colleague effervescing about whatever they are working on is kind of the norm in pure research. Just because they are making progress it doesn't mean they are about to do anything interesting.

Also it's not like you're looking for a lost pair of car keys: just getting to the level where you can understand a problem well enough to "steal" it takes a huge amount of work. People are going to know if you're at the level where you could be a competitor.

So in general it's pretty safe to talk generally about whatever you're doing.

20k 3 hours ago

"The route to the Clay problem through a smooth force, options c and d in Fefferman’s statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag."

We know that OpenAI trained on their prompts, plagiarism is incredibly likely. The only thing we don't know is whether or not it was deliberate plagiarism yet

irthomasthomas 2 hours ago

And deliberate or not it is still plagiarism by the sound of it.

jwr 3 hours ago

I always thought it was enough to switch off the "Improve the model for everyone" setting on chatgpt.com:

"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."

But apparently there is also an entire completely different route "Do not train on my data"?

Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?

We are getting to facebook/meta-levels of privacy settings obfuscation.

teiferer 3 hours ago

That's because what people enter into LLMs is the last gold there is out there. Everything else is already scraped or ensloppified.

Maybe next step is to filter your input client side through an unknown number of obfuscators where you ask LLMs to rephrase your question (onion router idea) such that no single provider can be certain that this is human input and not some slop feedback loop.

jsw97 2 hours ago

The last time I checked there was a loophole — if you provide feedback in-session (responding to “how are we doing” or “which prompt is better”) then they can use that feedback + relevant context. Relevant context for chatgpt might include memories / other sessions. That may not be the only loophole.

That in itself is a dark, dark pattern. There should at the very least be explicit warnings for users who have checked “do not train”; or they should not be presented with such dialogs.

sdcfgy 3 hours ago

My take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.

I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.

junofan 3 hours ago

Their privacy policy for normie subscribers says in plain English they use your Personal Data for research. I think it’s pretty unreasonable to use the service and expect otherwise.

ZeWaka 3 hours ago

>implying most users read them

andersmurphy 3 hours ago

You'd think theft would still be illegal regardless of what a privacy policy says.

georgemcbay 3 hours ago

> You'd think theft would still be illegal regardless of what a privacy policy says.

I wish that were true, but I live in the United States and it is 2026.

The President of the United States rug-pulls memecoin crypto and regularly pardons people like Paul Walczak (who was convicted of massive payroll fraud) in exchange for large donations.

I wouldn't make any assumptions about what is considered theft anymore, at least not when it is being committed by people who have enough money to be above the law.

andersmurphy 2 hours ago

I mean the fact that comments like this get downvoted is wild.

sdcfgy 3 hours ago

If that is the case, why on earth would you use it in any professional setting?

eru 3 hours ago

Depends on your profession? I sometimes work on open source code as part of my professional duties. Nothing that goes on there is necessary to keep private.

sdcfgy 3 hours ago

Vulnerabilities, responsible disclosure etc?

eru 2 hours ago

You can use 'Daybreak Blue' from OpenAI for that, if you want to.

Simran-B 3 hours ago

Aren't business consulting firms even worse? They are explicitly for business and there are known cases where they shared confidential information of one of their customers with another one.

ragebol 3 hours ago

Because it's cheap and easy. Hold for many such questions...

Why is all of the world dependent on tech an ever more hostile US? Same answer.

jonathanstrange an hour ago

I have the setting turned on in Gemini Pro even though I work on proprietary code because 1. the setting allows for some (very limited) "memory", and 2. I consider my source code almost public even when it's not open source because I don't work on programs that involve extremely high level of know how or proprietary algorithms. It's mostly CRUD that can be copied in a myriad of ways, whether people use my methods or other methods.

If Gemini can improve based on my code and sessions (maybe doubtful but who knows) and others can benefit from it, that would be a welcome side-effect.

hansvm 3 hours ago

My doctor's privacy policy is a bit more abusive than OpenAI's. It exists mostly because of a $%^&&* legal framework rather than malice, and I've grown accustomed to "if I don't want to die then I sign away these rights." Despite my having theoretically signed my soul away, my doctor isn't selling personal information to my exes or to life insurance companies (though they could in the US; that extremely personal information is no longer mine). OpenAI is engaging in the "technically legal maybe we'll see but obviously unintended" side of this transaction, and maybe that works out for them, but I wouldn't personally choose to be a shill for "it's unreasonble to expect somebody with 'legal' permission to do something other than the maximum 'legally' permitted" if I were in your shoes.

grey-area 3 hours ago

I think those are pretty unreasonable terms and it’s sad that you’re defending them.

Does google docs own the content of docs you make with it? Does Apple claim ownership of discoveries made using their tools?

shiandow 3 hours ago

I think it’s pretty unreasonable to use the service.

AlphaSite 3 hours ago

There’s an opt out so I’m blade for now.

walrus01 3 hours ago

> as they all seem to be run by assholes.

Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.

You know this metaphor? https://en.wikipedia.org/wiki/Turtles_all_the_way_down

But instead of turtles, it's assholes.

But more seriously, no, none of what I wrote above is an attempt to excuse or play down the specific role of assholes in large AI companies.

sdcfgy 3 hours ago

Of course. I work for assholes. The thing is they’re getting out assholed by several orders of magnitude here.

dotancohen 2 hours ago

  > Wait until you learn how many other SaaS and web 2.0 and cloud based things are also run by assholes.
Though it's not the real acronym, Larry Ellison himself has stated that Oracle stands for One Real Asshole Called Larry Ellison. He famously prides himself on it.

johanvts 3 hours ago

Or use Lumo from Proton. Are there any other privacy first companies offering LLMs?

Cider9986 2 hours ago

Lumo seems the most reputable. There's confer.to, duck.ai, and nanogpt.

irthomasthomas 2 hours ago

Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.

Blikkentrekker 3 hours ago

That they do it is just concerning to me in that it says that home-ran models just aren't good enough. Surely researchers like this have the processing power to run them at home, they just don't have the processing power to train models of comparable level.

This is something I feared would happen and where open source would be left behind. Maybe they can do something with crowd-sourcing computational power from volunteirs. They were after all able to get Leela Chess Zero to be comparable to AlphaZero by training from volunteer processing power but it seems to me we live in a world now where the best models keep their stuff closed.

In imagine generation too. I'm not sure how well Stable Diffusion can compete in following instructions with all those advanced models that are kept secret.

isaacfrond 2 hours ago

> Surely researchers like this have the processing power to run them at home,

Nobody has that power. Certainly not mathematicians.

pansa2 3 hours ago

> one should not use LLM services for confidential or proprietary information

That’s obvious, isn’t it? Just like you wouldn’t upload your confidential documents to an online spellchecker, or your proprietary code to an online compiler?

sdcfgy 2 hours ago

Everyone, including the security services in my country, uses OneDrive and O365.

I don’t get it.

anticodon 2 hours ago

I worked in a EU company that was developing a product competing with one of Microsoft offerings. All the code was hosted on GitHub, we used Azure for hosting and all the internal communications were in Microsoft Teams.

To me it looked absurd. But I was almost the smallest cog in the corporate structure, so I never asked what were the reasons for all these decisions.

palata 2 hours ago

Or you wouldn't write confidential emails in Outlook and have confidential conversations in Teams. Obviously. Right? Right?

hn993302 2 hours ago

Or if you're going to trust one of them, maybe it shouldn't be OpenAI

sdcfgy 2 hours ago

Why would I trust any of them?

hn993302 4 minutes ago

It's either that, not use an LLM, or local-host an LLM. If you can local-host one then great, this is a compelling reason to do that too. A lot of people aren't in that position.

feverzsj 3 hours ago

It could be much worse.

OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.

No LLM is even needed.

grey-area 3 hours ago

Or, using anonymised data, search for anyone seriously trying to tackle this problem - probably about 10 people in the entire world and use their ideas as a starting point. They wouldn’t even need to be watching specific accounts or using de-anonymised data if they know what they’re looking for.

_bobm 2 hours ago

People are focused on the drama but the problem showing is the data. This is the elephant in the room and I am surprised that openai can be that stupid with it.

How can openai do this, what is being claimed, at the scale of their entire userbase? If they do this only for particular sessions then how do they sieve through sessions for the good stuff?

How are sessions stored, how are they processed, how much storage and how much compute is used in these pipelines, how economical is it, how fast are the requirements on the storage on the compute growing as userbase grows and generated data grows.

All these questions are far more pertinent than the navier-stokes, but i can only imagine all at openai doubling down on this "very important" mathematical milestone.

jonathanstrange an hour ago

It seems completely trivial to feed sessions to their own LLM and ask it to look for various things in them, from detecting problematic use cases to finding interesting mathematical work.

keremk 3 hours ago

Occam's Razor says: "They heard this problem is solved or about to be solved amongst the rest of the other problems. They prioritized this and put substantial compute with their newest model and solved it." I know everyone loves juicy rumors, theories etc. but honestly that is the simplest and most plausible explanation given the state of AI improvement now. Obviously spending 15 million on a problem is not a slam dunk decision even for a company like OpenAI but if it has a significantly high chance of solving it and their competitor will be claiming they solved it, then it raises the stakes and they go after it. In fact this is the most rational and also curiosity-driven thing to do and totally what I would have expected from any frontier lab. Of course if one wants to prove their confirmation biases that they train on sessions or be able to identify individual users, the non-zero chance of that being also another explanation is attractive enough to wet their appetites.

npiano 3 hours ago

Why is that simple or plausible? Why is simpler or more plausible than lifting an almost-finished solution from a researcher's account?

ghshephard 3 hours ago

Because they solved different problems, and where there was overlap, the solutions look different?

karmasimida 3 hours ago

Because it is not finished. Their follow up claim is that OpenAI’s approach looks like another proof they had been working on the side, but hasn’t published yet

20k 3 hours ago

OpenAI have admitted their new model they used was trained on prompts at around the time that researcher was working on it, so it seems self evident that it was used as part of the millennium solution

spwa4 2 hours ago

1) Because AI models are 1000x better about following a problem to its conclusion than coming up with a genuinely new idea.

2) Because if we accept the facts ChatGPT only came up with its "new" idea after being told exactly what the new idea was by a mathematician (OpenAI doesn't dispute this btw). And OpenAIs story comes down to the usual "We didn't look at it, trust me bro", which is made more hard to believe because OpenAI only started their efforts after receiving news of what the researcher was doing.

Oh and OpenAI emphasizes that part of the researcher's progress was made ... on OpenAI.

3) And, probably, the researchers were likely stopped by token limits, and that's the only reason they were slower than OpenAI themselves, which is very, very unfair.

4) OpenAI's story "smells" (like so many AI stories lately). Supposedly the company's team asked ChatGPT about solving millennium problems, and out of all millennium problems it just happens to pick the one where a solution can be found in its chat logs?

5) Yet again it would be in good taste for these AI companies to just give this to the researchers (no shortage of difficult unsolved math problems, so if AI can solve them all, just find another one). But instead, yet again they're fighting about it.

6) OpenAI admits they only went after this problem, with a team, no less, after finding out which researcher went after what problem, because of how they thought it would affect ChatGPT's PR. They are demonstrating, in other words, their willingness to destroy human researcher's reputation for PR wins.

That's getting close to big tobacco level morals right there.

7) If anyone wants to verify how much OpenAI cares about the truth, just ask on ChatGPT about the copyright lawsuit outcome and how it applies to OpenAI.

You'll get EXACTLY the sort of responses you get from Qwen about Tiananmen ("we didn't do it, you have no data, everyone's lying and if we did do it, it was perfectly reasonable because " style argument. Try it)

karmasimida 29 minutes ago

I don’t know much about NS problem or Euler problem anyway. Can’t tell how significant it is to go from Tristan approach to OpenAI’s results.

It is 10k agents after all with a model that is 2x-3x more powerful than Astra. So claims some intuition is his alone and the model can’t find it independently is Tristan’s belief not the reality, which I don’t find particularly convincing.

I do think that Levent guy isn’t independent and indeed does have an enormous amount of tokens to spend and also on an internal model from Anthropic.

All things considered, I think OpenAI people are competitive, even aggressive. But I don’t think they steal anything. Ultimately, the model’s capability is the real surprise factor here, the fact they could get a solution after all in 3 days, that is the cause of drama, if it is 3 months, then it is a nothing burger, because by then the original results had already been published.

l5870uoo9y 3 hours ago

Given that the solution took a somewhat “unusual” approach, I find it even more unlikely that an AI model would have come up with this on its own.

pietz 2 hours ago

Isn't that *exactly* the type of solution you'd expect from AI?

Move 37 comes to mind.

iLoveOncall 2 hours ago

Only if you understand nothing about the difference between LLMs and AlphaGo.

alangibson 3 hours ago

This is not the simplest explanation.

The most direct line from problem to proof is OpenAI building off of conversations the mathematicians had with their AI.

PowerElectronix 2 hours ago

Funny how they did not solve any of the other problems, just the one where there was already solutions to the NS with some restrictions in their chats, and their solution seems to derive from those.

jojva 2 hours ago

They explain in the article that they eventually redirected all their resources towards NS.

consp 2 hours ago

So instead of dictating research to unsolved, or largely unsolved, problems, we are now as "a society" directing compute power towards sniping research outcomes.

I thought this was an ebay thing for people with too much free money, but it seems a bit larger.

OtherShrezzing 2 hours ago

Considering how language models work, you'd expect two distinct conversations on the same mathematical problem to have enormous crossover.

ozgung 2 hours ago

I agree with Occam.

That 100x step up from using 100 agents for Euler to 100_000 agents for Navier-Stokes, in a single day seems a bit sus.

I also think this is not ethical behavior. This is at least academic dishonesty, kind of a plagiarism or intellectual theft.

That’s why they wanted to credit Tristan and to give $1M award to him. But again, they acted unethically in that process as well. They wanted him to remove Levent (Anthropic affiliation) as co-author and threatened Tristan to “end his career”. Their behavior is actually telling, their work was not completely independent from Tristan&Levent’s unpublished work.

Bluestein 2 hours ago

There's also an elephant here in this room: Quite "coincidental" one of the coauthors worked for the competition. Of course GPT would know this fact. And OAI has every incentive to particularly monitor those accounts.-

davesque 2 hours ago

Isn't it also a simple idea that a model designed to recall relevant information from its training data, which is also known to have been trained on data from user transcripts, would, in fact, reproduce directly relevant work by leading experts in the field? Seems like Occam's razor would apply to that situation as well.

We know that LLMs are trained to recall relevant info. We know AI vendors are using user transcripts to train models. Two plus two equals four, right? I mean, an LLM that failed to recall the transcripts of those researchers would be a bad model.

aadyachinubhai 3 hours ago

LLMs can't contribute good code to some of the good OSS math libraries, How is it even solving these problems?

matrix2596 3 hours ago

search, verifiability and compute

emil-lp 3 hours ago

That's a good question.

It is able to contribute code, but maybe not good code.

It's the same in math: it's able to solve problems, but not necessarily in a good way with a human readable code.

Math papers are a lot like software:

- theorems are like API

- lemmata like internal/private function API

- definitions are like types

- the proofs are the implementation

The proofs of ChatGPT are not necessarily readable or maintainable.

linkgoron 3 hours ago

You don't need a GOOD proof, just A proof.

feverzsj 3 hours ago

No one knows if it's actually LLM doing the heavy weight. It could be just human written brute force algorithm running on their massive computer cluster.

vbarrielle 2 hours ago

Math problems are often stated in a way that makes it possible to automatically verify if a solution is correct. Which means a loop that speculates an approach (LLM and/or prompts), implements it (LLM), then checks (automated) can work. You still need to have a very good LLM, and probably very good prompts with interesting research directions otherwise you can probably loop forever.

recursivecaveat an hour ago

Notably all the major announcements so far are counterexamples or formalizations of existing results to my knowledge. Not necessarily something you can just brute force, but areas with high return on elbow grease.

caughtinthought 3 hours ago

Basically no new info here, not really sure why this post needed to be written tbh.

fimi 3 hours ago

Yes, you have ability to get information very quickly, but not everyone does.

kzrdude 3 hours ago

On the contrary, a level-headed summary that gathers information from all the different sources is necessary.

caughtinthought 3 hours ago

Sounds like a great use case for an LLM

teiferer 2 hours ago

They are busy generating pelicans on bicycles. The summary therefore needs to be written by a human.

(This was a joke. I value Simon's role in the community.)

kzrdude 43 minutes ago

It's got 'max' and 'ultra' but I don't see a setting for level-headed.

Simran-B 3 hours ago

I don't get the sales pitch, spend 15 million dollars to win a 1 million dollar price?

Showing of the model's capabilities - okay, but it's not like it solved the problem on its own, and apparently not particularly efficient. Are there practical applications that justify the investment?

grey-area 3 hours ago

The answer to this is obvious: the effect on the multi-billion dollar valuation in the imminent IPO.

teiferer 3 hours ago

How is that any different from any other academic research? Every PhD candidate solves problems essentially nobody cares about. They don't even get $1M, they get nothing.

There are two benefits though.

One is recognition. Cred. The PhD candidate gets to put a ", Ph.D." behind their name, opening doors to future academic employment or other endeavors where people value titles. The AI lab gets to say their tech solved sth that humanity wanted bad for a long time. Both cases with substantial financial upside (higher income for Mr. PhD and higher company valuation for the AI lab).

The other one is that this is how scientific progress works. $1M or not. That number was just a PR campaign by the math community to point to some goals. It's clear that it would cost more than $1M to get there.

tu26muwu 3 hours ago

> The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster [...]

I think you should not say that. Buckmaster did only state his version of events and was very clear on that he did not make any accusations at all.

To quote from his statement pdf:

> I am not accusing anyone of anything.

Aeolun 2 hours ago

But we all know that is code for 'something wonky happened, and it wasn't me'.

rao-v 3 hours ago

It’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).

What I cannot reconcile is the timeline and the concern in this specific case.

I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)

KeplerBoy 3 hours ago

I would assume a lot of codex data goes back into training. A well steered session is extremely valuable data.

_bobm 3 hours ago

Hah, what is the infrastructure which takes user sessions (chats with API keys, directions, navier-stokes math/progress) and regurgitates this into pre-training, RL, fine-tuning data? Or better, in-context data?

People talk about the "compute" but what about the "storage"? Is storage exponentially greater, or soon to be, than the compute? Is the storage going to slow down growing to some constant rate, i.e. all people on earth using chatgpt, or no, on the contrary, it will keep growing?

If there were any shady business, I do not condone it, but technologically we are not there yet for said shady business to happen.

awestroke 3 hours ago

AI companies use heuristics to filter sessions, then llms to further filter, then use various techniques too anonymize the session, then process it and add it to various datasets for further selection and refinement. they don't need huge storage for this.

_bobm 2 hours ago

what is behind "process" it and "further selection" and "refinement" and how big are these "datasets"? These companies ship the encrypted session to you not because they want to.

I agree that they have pipelines for what you are describing but how effective they are at scale and at focusing is the question.

vb-8448 2 hours ago

> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models

AKA everything you send to them (and I bet it's the same for any other lab) will be used, no matter what are the TOS, the law or what they publicly say.

unified101 34 minutes ago

Put down the pitchfork. It's a toggle in their settings.

protoman3000 3 hours ago

If they wanted to solve a Millenium prize problem so much, why did they not try to solve P-NP instead? It boggles the mind.

cammikebrown 3 hours ago

That one is way more difficult.

eru 2 hours ago

What makes you think they didn't try?

throw-qqqqq 2 hours ago

Are you kidding/trolling? The P=NP problem is FAR more fundamental, and if proven true, would basically be a proof that e.g. public key crypto can be broken (NOT a description of how to though).

Basically, it would be a proof that all the REALLY hard (combinatorial) problems out there, have a much simpler solution, if we were able to find it.

EDIT:

NS is used daily in engineering and gas/fluid modeling. We sort of “know it works”. The smoothness proof is “just” formalizing what practitioners assume is true (very coarsely said, no intention to diminish the result!)

It’s a bit like the Collatz function IMO, empirical evidence isn’t proof, but we’ve got a huge amount of evidence for the behavior we’re trying to prove.

I believe P vs NP is a different beast entirely. We don’t even know which way the answer should go.

lucfranken 2 hours ago

Outside of this discussion about what is fair and not. This is so highly interesting to think through.

The amount of millions available to do those kind of research cases is practically unlimited.

There are an unlimited amount of cases to work on.

What a huge development would this give to both humans and the world in general. Because in the end better understanding gives new options.

It's deeply interesting that those things now get a concrete economical price which seems to be viable to extrapolate. The enormous additional "production" of knowledge will inherently increase the speed of all pieces of research and development.

Taken into account that it's used wisely, the risks with a strong force are always huge as well.

maciejzj 2 hours ago

I know that this may be somewhat dramatised and even infantile, but my reflection is that in the world run by these reckless AI companies everyone looses. Navier-Stokes is solved but it feels like no one has won anything, controversy prevails, there is no glory in the math breakthrough. There is hardly anything to cherish, and even the guys at the top of it in OA who sit on the (supposedly) superhuman intelligence come across as massive losers and frauds.

RhysU 21 minutes ago

I keep looking for a technical article to appear on HN discussing literally anything about the mathematical result--- Not fluff, not marketing, actual content.

Instead, all I read on HN about N.-S. is human soap opera, told from every possible angle.

In 100 years we won't care about the soap opera. The N.-S. result itself will still matter.

Someone, anyone, please, submit articles on the result itself.

cs_throwaway 3 hours ago

Maybe someone on the NYU team forgot to opt out of “improve the model for everyone”.

karmasimida 3 hours ago

Use Bedrock or any kind of big tech hosted version of the frontier labs can be a solution. I think if secrecy is of utmost importance to you, then do not send data to first parties