>But an LLM has no mind to feel bad if it cheats without getting caught
All the interpretability research we have would not indicate that "LLMs have no mind". It seems to me you have a conclusion and are working backwards to justify it. I guess I just don't see where 'they have no mind' would logically follow 'they sometimes cheat'.
It doesn't ring true and it never rang true. He was wrong in 2022 and he'd be wrong today. He had (and likely still has) a wrong model of LLMs. Not exaggerating 2022 capabilities and having a model so wrong you're out of whack 4 years later are 2 different things. There are people who had the right idea from the start.
They aren't going to sit on millenium solutions regardless (so if Hodge and /or BSD is really done it will get announced especially because of the baseless accusations), and they aren't going to stop trying to solve P/NP and Riemann. The letter doesn't really matter. It's not the first time, and I don't think AI's dramatic ramp in capabilities ever had positive reception from the bulk of mathematicians anyway.
Yeah, I don't expect them to stop, but I do think they are probably hurting themselves, as well as mathematics, by continuing do to this.
Imagine if they had handled this differently and these results - still using OpenAI models - were coming from the math community. How much better the PR would have been - AI helping math/science rather than yet another story of AI harming society in some way.
No doubt this is what they were at least partially aiming for - not just shooting a trophy animal to brag about, but also being seen to advance math/science, a la AlphaFold, not just take all our jobs and enshittify society with deep fakes and AI slop. But, they heavily misjudged.
If you're in OpenAI's position, the PR from the group you're disrupting is rarely relevant. People from the outside will (correctly or not) look at this as "Mathematicians don't want OpenAI to solve problems to keep their status/jobs", and you'll have as much sympathy from them as every other replaced profession in human history - very little to none.
Unlike AlphaFold, this technology has the potential to wholesale replace the entire profession. You're never going to get anything more than bad PR from that group as the threat looms.
Over 3 Billion images gets generated per week via OpenAI chatgpt image models. None of the poor PR from artists on AI generated images even remotely matters.
Most research mathematicians are employed as academics, and it's hard to see universities replacing teaching staff with AI even if that were possible.
IMO giving a hypothetical AlphaMath to mathematicians, the same way Google gave AlphaFold to research chemists/biologists, would have resulted in far better PR, and the profit opportunity of attempting to replace the jobs of either research group is minimal.
It's downright bizarre the way companies like Anthropic (primarily), and to a lesser extent OpenAI and anyone else, are themselves pushing the narrative of this tech may kill you, will take all your jobs, etc. That may all happen, unfortunately, but being aware of that possibility you'd think these companies would be a bit more mindful of their messaging and behavior, not just go out with a scorched earth approach of "well, they are going to hate us anyway".
>Most research mathematicians are employed as academics, and it's hard to see universities replacing teaching staff with AI even if that were possible.
If universities come to only need research mathematicians for teaching ability, then it would still gut the profession. You would presumably need far fewer of them for their research ability, and hopefully hire more people who can actually teach. It's strange that you don't see that as a threat to the profession. Lots of research mathematicians would lose their jobs even in this scenario.
>IMO giving a hypothetical AlphaMath to mathematicians, the same way Google gave AlphaFold to research chemists/biologists, would have resulted in far better PR
GPT-N isn't AlphaFold. I'm telling you AlphaFold is a bad analogy! AlphaFold has superhuman capabilities at one specialized subproblem - protein-structure prediction - embedded in a much larger biology/drug discovery pipeline. It doesn't replace the biologist or drug researcher, it just gives them a better instrument, and if you're not in the relevant professions it's essentially useless to you.
GPT-N is general purpose intelligence machine that can increasingly do chunks of work you would otherwise employ the mathematician or software developer to do.
Shanmu Jin is a perfect demonstration. A neurosurgery resident, not a research mathematician, who encountered the Crouzeix conjecture through his transcranial ultrasound research and used GPT-5.6 Sol to solve it (20+ year old longstanding problem).
In the AlphaFold story, the biologist gets a powerful new tool. In the Jin story, someone who isn't a mathematician can obtain research level mathematics and incorporate it into their research/code/business whatever without talking to a single human.
You're asking why they don't market GPT as "AlphaFold for Mathematicians". They don't because it's not.
>It's downright bizarre the way companies like Anthropic (primarily), and to a lesser extent OpenAI and anyone else, are themselves pushing the narrative of this tech may kill you, will take all your jobs, etc.
It's not that bizarre. It only seems strange if you think it's all a facade, "marketing" or whatever nonsense HN is convinced of. These are people who believe very strongly in what they are doing and in the potential of it. For these people, what you are asking them to do is actually incredibly slimy. And it might win some brownie points, but not for long.
> You would presumably need far fewer of them for their research ability, and hopefully hire more people who can actually teach
I suppose you are suggesting that research mathematicians are either over-qualified mathematically and/or under-qualified in ability to teach, but it seems pretty clear that universities prefer to hire domain experts whose reputations and long publication lists attract students and raise the perceived academic standards of the school.
Your "scenario" of much math research soon being done by AI (supervised and paid for by who, one might wonder), while universities hire people chosen for their teaching skills not academic reputation, seems a bit of a stretch ...
> You're asking why they don't market GPT as "AlphaFold for Mathematicians". They don't because it's not.
No, I am not asking that, nor asking anything for that matter.
I was pointing out that giving a free research tool to mathematicians would likely be better received, and receive better press, than doing something that most top-tier mathematicians are opposed to.
My hypothetical "AlphaMath" certainly could just be free GPT-N access for academics/researchers (as they are also doing to some extent), or it could indeed be a custom system that OpenAI built as a gift to the math community.
As far as the AI companies acting in slimy fashion goes, the most slimy behavior of all is to strongly believe you are doing something that will kill people and cause massive societal disruption ... and still keep doing it.
>I suppose you are suggesting that research mathematicians are either over-qualified mathematically and/or under-qualified in ability to teach, but it seems pretty clear that universities prefer to hire domain experts whose reputations and long publication lists attract students and raise the perceived academic standards of the school.
I'm saying there's little correlation in how brilliant a researcher you are and your teaching abilities. Yes universities optimize for the former. That's not necessarily a good thing for students even if it increases the prestige for the university. If there's some future where research and teaching are decoupled, I don't expect things to remain the same, but who knows I guess.
>I was pointing out that giving a free research tool to mathematicians would likely be better received, and receive better press, than doing something that most top-tier mathematicians are opposed to.
At least currently, it's not possible for this to be a free tool. Both OpenAI and Anthropic have discounts for Universities/Education for such use cases as far as I'm aware.
>As far as the AI companies acting in slimy fashion goes, the most slimy behavior of all is to strongly believe you are doing something that will kill people and cause massive societal disruption ... and still keep doing it.
This is pretty bad, but I wouldn't call it slimy. And it's no longer up to any one person or company anymore.
Nah the AI won't replace coders or mathematicians until it can maintain codebases/knowledge long term.
I actually don't think that current AI can replace any profession that requires human interaction over many weeks. This is because imo they lack long term planning abilities
>Nah the AI won't replace coders or mathematicians until it can maintain codebases/knowledge long term.
Okay...you understand that are training for this and it has gotten much much better at doing this over the years ? You should probably also understand that it doesn't need to be able to do this to cull the profession ?
I also think they overestimate how much of most jobs can be done by an LLM (a text generator!), and how much importance and reliance most companies put into soft skills, and non-linguistic understanding/feels, both during interviewing (are they Googley-enough?) and afterwards.
For example, how do you reconcile return-to-work mandates with the idea that companies are going to be happy with faceless remote workers? What does the boss do when the shit hits the fan and he would have yelled at people about the need to work all weekend, but instead all he has to yell at is an LLM that tells him he's "right to push back", that it promises not to delete the production database next time (except it will, because it can't learn), and that it could care less about being fired because it's just a calculator?
Yeah. Not to mention the paper-clip maximizing that RL induces in them. I had 5.6-Luna use a parser combinator lib in order to find out that it imported it but wrote its own parser, so it technically followed my instructions.
Imagine that paper-clip maximizing happening over millions of tasks.
In a related vein, not too long ago I asked Sonnet how many states and non-terminals were in an a YACC parser for ANSI C, something that could easily be googled for, but it instead chose to go off and downloaded and build bison from source, downloaded a grammar, built the parser ... but then failed to give the answer since I'd hit my daily free limit.
It's really not that big. Yeah Navier-Stokes was easier than Riemann but that's not really the issue.
AI has and will improve at a much greater rate than human mathematicians. So it's really a question of if AI gets good enough to tackle it before any human does. It doesn't look like humans will be solving it anytime soon but where will AI be in 2 years ?
Hell, it looks like at least one other result will be announced soon too.
The thing about mathematics is that it can be arbitrarily hard, including impossible to prove a given theorem.
I don’t know the details of RH, it might very well be solved soon, but it could also be impossible or just so difficult that even orders of magnitude more intelligent AI can’t solve it even.
If it is impossible to prove, it might be possible to prove that it is impossible to prove, or that itself might be difficult or impossible…
Has and will. Are you going to back that assertion up at all, or just repeat it like that other viral thought-terminating cliche: ‘this is the worst the models will ever be’?
No it isn't. Best and worst and ill-defined anyway but the chess ELO score of various LLMs has fluctuated up and down, it's not been montonically increasing. What is the best answer to "how do I make cocaine"? The models are getting larger, with more compute and RAM backing them, but that doesn't automatically make them better if you don't define how you're measuring better-ness.
None of the frontier labs care about Chess as it's already a solved problem. If they did, the models would be much better. It's really not that hard. Google has a paper on grandmaster level chess without search from transformers.
Better obviously means better, like how they became better than they were 6 months and a year ago.
"Better" is not one dimensional across all use cases even if model capabilities are improving in aggregate.
e.g. If someone said "this is the worst they'll ever be" in response to some writing with obvious LLM cliches in 2024, I'm not convinced that prediction was actually correct.
The focus of OpenAI/Anthropic pivoted aggressively to the agentic performance arms race instead of making a more human sounding chatbot so regressions in writing ability aren't really a concern anymore if agentic benchmarks improve.
The first time I heard a recommendation to use Claude was specifically because it sounded much more "human" and natural than ChatGPT. Fast forward to now and idiosyncratic Claude-isms repeated every other sentence and its convoluted verbosity has become a widely mocked meme.
Right but we're talking about a single subject here - mathematics that labs are incentivized to keep improving for some time.
>e.g. If someone said "this is the worst they'll ever be" in response to some writing with obvious LLM cliches in 2024, I'm not convinced that prediction was actually correct.
2024 creative writing prose was...the last few versions have stalled, but I think they're still better than 2024.
I would describe better as how much of my work I can delegate to the agent. Right now I'm delegating much more to Astra high than 6 months ago to Opus 4.6. Every dev has this feeling, it's weird to even argue what a better model/harness means.
Chess is not solved in any meaningful sense of the term. Computers have been better than humans since the 90s, but better chess programs are released all the time.
It's solved in that we have had grossly superhuman capabilities for some time. It's not interesting for frontier labs. I suspect you understand this and the greater point so why be needlessly pedantic ?
They aren't going to stop at one, that's for sure. They already claimed they have "made substantial progress" on another millenium problem. Let's say they bag another one (Hodge and/or BSD according to the rumors), if it looks like their internal model could solve P/NP or Riemann Hypothesis, you think they wouldn't take that chance ?
If it's a counterexample to BSD, that would be pretty surprising.
It would also be a considerably more impressive achievement, because experts had mostly shifted to Navier-Stokes regularity being false, while as far as I know almost everybody thinks BSD is true. Hodge people seem less sure about.
If either conjecture is true and they prove it, that would be an even bigger success, since the techniques might unlock any number of other theorems.
>The Wednesday evening statement from OpenAI was more emphatic: “We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training.”
>The statement added, “After investigating, we can say with full confidence that no user inputs past July 3rd could have influenced this system in any way.”
Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months
two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem with 10,000 agents?
Just knowing that there had been progress is enough to have an idea that throwing more compute at it might work (OpenAI had previously tried all the Millennium Prize problems with somewhat limited compute and failed).
It's comparable to Magnus Carlson saying that if he wanted to cheat, all he would need would be for someone to tell him to spend more time thinking about a specific move (just a wink would be enough) as an indication that a computer had found something interesting.
It's as-if after OpenAI first failing on Navier-Stokes (which OpenAI had just tweeted about 2 days earlier!), someone winked at them and said "you might want to try a little harder ...".
The comment you replied to quoted "no user inputs after July 3rd" with no restriction to Buckmaster or Codex.
Obviously the result of OpenAI's investigation was that no usage data has interacted with the system after that date.
What else do you expect them to investigate?
If Buckmaster and co. provide their chats, OpenAI could potentially search for them in the anonymized opted-in usage data. Then they could say if any data has been used.
By all accounts individual usage data does not have the direct impact on the model most here fantasize about. To prove this, OpenAI would need to do new training runs to replicate the system used minus the particular usage data in question, if it exists, and then benchmark this on the problem again.
Potentially multiple times, in order to reach a conclusion.
OK, good to know (if they can be trusted - Altman clearly is a liar), but it doesn't really change the big picture much.
1) OpenAI by their own admission, only re-tackled Navier-Stokes because they heard it had already been solved (but not yet published). This isn't advancing science or helping the mathematical community, this is just being a dick.
2) OpenAI, specifically Sebastien Brubeck, then threaten to "not be nice" and "ruin the career" of one of the mathematicians whose work they had succeeded in duplicating, unless he agreed (which he refused to do) that his collaborator, an Anthropic employee, was not named. This is not only against mathematical norms of credit assignment, it is also being a pathetic human being.
OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute, and the assistance of a whole team of people at OpenAI, to replicate (then exceed) the work that just took two people, with some academic grants as an AI spending budget to achieve (a few $100K - listed below).
I'd say advantage humans this time. Better luck next time OpenAI - and if you don't want unfavorable comparisons then maybe choose to work on problems that have not been solved yet, and that humans are NOT making nice progress on.
> OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute, and the assistance of a whole team of people at OpenAI, to replicate (then exceed) the work that just took two people, with some academic grants as an AI spending budget to achieve (a few $100K - listed below).
I think you have to work pretty hard to minimize what OpenAI achieved here like this.
The Navier-Stokes equations have been around since 1850. The smoothness problem has been well known for over a hundred years and has only gained importance. It's been a Millennium Problem since 2000.
Levent Alpöge and Tristan Buckmaster did great work to solve the related Euler problem, but didn't solve the Navier-Stokes smoothness problem.
The Navier-Stokes smoothness problem has previously had significant resources working on it. Computational fluid dynamics is one of the most important tools in modern engineering and is closely related.
You speak of 10,000 agents as though it is somehow extreme, and yet within the past month I've had a single task that used over 100 agents on a mere Anthropic team plan. I think two orders of magnitude more compute to solve one of the greatest unsolved physics problems[1] is nothing.
I don't excuse Brubeck behavior because of this, but that doesn't minimize the achievement here.
[1] Wikipedia quote: In particular, solutions of the Navier–Stokes equations often include turbulence, which remains one of the greatest unsolved problems in physics, despite its immense importance in science and engineering. https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_existenc...
1. I would agree if the rumours were that some mathematician(s) had solved them, but the rumors alleged it was Anthropic. I don't really see what the big deal was. They had a new model that was going along great and wanted to test its mettle.
2. Yes Brubeck's comments were weird at face value. That said, Open AI's proof isn't a duplication of anything. Not only is Tristan's work a sub problem but the methods are different. And what OpenAI didn't want was Levant on the paper OpenAI authored not whatever they were working on (Euler). It's petty sure but it's fair enough. Tristan and Levant didn't have anything to do with the Navier Stokes solution, so it's really their call if they didn't want to collaborate on their own paper with the Anthropic employee.
>OpenAI would have you believe this result shows how powerful their mystery better-than-Astra model is, but the reality here is that this model needed 10,000 agents, $20M of compute,
$20M in approximated API prices doesn't mean they spent $20M worth of compute. The real number would obviously be substantially less.
>and the assistance of a whole team of people at OpenAI
You can't eat your cake and have it. What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one.
>I'd say advantage humans this time....to work on problems that have not been solved yet, and that humans are NOT making nice progress on.
Interesting way to frame progress that didn't move along till an LLM generated proof.
> What sort of guidance do you think is happening in a 10k agent, 320b token, 88 hour run ? AI did this one
If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.
Does this aspect really matter? Not really, other than OpenAI wanting to present this as all the work of their model.
I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien “very little
human input” had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over
their internal chat, it emerged that an entire team had been working on the
problem, that this was one of a number of things that was tried, that work had
started on the unforced problem, that the team first set the model on easier
problems, including Euler, that even the prompt that had been shown to me
had been written by prompting Codex, and that an insane amount of compute
had been used.
I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
>If you read the PDF release by Buckmaster, apparently the initial claim from Brubeck was that there as very little human input involved, then as the call progressed more and more people popped up that has been involved with it.
As it seems and as they tell it, they started the run modestly and diverted more resources towards it as it looked more and more promising. The run didn't start with 10k agents for instance. The point is there isn't anything humans are doing in this timeframe against all this text that would count more than "little human output". It's still a fair assessment I would say.
I put it like that because of Brubeck's own words on the matter. You're acting like we've gotten email receipts here. I'm not really interested in going over a he-said she-said about strangers.
Brubeck has admitted what he said, but claims he immediately retracted it as a "poor choice of words".
Given Buckmaster's telling, this seems beyond "poor choice of words"... It was a veiled threat, that he then doubled down on with his "If you don’t want me to be nice, then I don’t have to be nice." follow-up.
**
I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, “Why would you ruin your career?”
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, “If you don’t want me to be nice, then I don’t
have to be nice.”
**
FWIW there are also other people on Twitter, such as this DeepMind researcher, saying this is a pattern for Brubeck.
Can you specify what leverage you think Brubeck has over an independent professor's career in order to make threats?
I see none, and consequently Brubeck's explanation makes more sense to me. I understand he meant these words, which he supposedly retracted on the spot, in a "why would you ruin your career with this behavior / turning down the opportunity I am offering" way.
I'm not sure the value of speculating, but since you insist...
If someone said to you "you're going to regret this, mo-fo!", do you really need to suppose they had a specific plan in mind, rather than just intent to intimidate?
If you want to suppose that Brubeck had a concrete plan for how he would ruin Buskmaster's career, then my best guess would be that he was threatening for OpenAI to publish without giving any credit to Buckmaster (who has been working on Euler overall for at least 10 years). Obviously this would be absurd, but no more absurd than what Brubeck was also demanding - that Levant was excluded for any credit and could not have his name on any OpenAI writeup (with this being the exact sticking point that Buckmaster was refusing to accept).
The funny thing is that this threat, as is often the case, seems to say more about the insecurities of the person making the threat than the person they are directing it to.
> Interesting way to frame progress that didn't move along till an LLM generated proof.
This part of your argument is totally wrong. The OpenAI approach begins with the B/L work. The belief / knowledge that their approach would pan out is worth a lot - it means essentially “depth-first” search in this direction will be more fruitful than a general search.
Unless you are counting the B/L work as LLM generated. Is that your argument? Even if you do consider it that way, to me racing in for a scoop isn’t a good look.
They've pretty much said their own work was heavily agent driven. Levent is in a particularly bad place here because while he probably had a lot of background in the Jacobian Conjecture problem, he made the solution to that one sound like someone asked the question and he just fed it to Fable during the world cup. Whether that nonchalantness was to just seem hip or was to promote Anthropic, which he has stock in, or was just the truth I don't know though. But it makes this one seem similar, when they might have had really had nearly a year of very valuable feedback to the models.
I was referring to the overall pattern of apparently sniffing around for recent mathematical progress then setting the AI on it to see if the problem is now easy enough to solve (if you have the money).
Terrance Tao has lamented this practice as being unhelpful for mathematics, and likely to lead to humans working in private to avoid this.
Tao has also noted that many of these AI math proofs don't really help mathematics (nor does it seem they are intended to), since for many of them the proof was never the point, it was the math expected to be needed to be developed along the way, which the AI solutions don't provide.
> has lamented this practice as being unhelpful for mathematics
A related point is that the actual solution approach is never revealed. What was the role of humans guiding the agents ? was it fully autonomous ? etc. It is in the incentive of the AI labs to trump the powers of the LLM, but in practice it is humans guiding the agents on the overall approach, This is never admitted. For example, in the announcement on NS there was only an output artifact given but no indication of how it was arrived at, and not even a writeup. This is what disappointed many folks as it was done purely for one-upmanship. As other have noted, the benefit is in the journey or process and not in arriving magically at a destination.
I don't think it's as bad as that sounds; in math people work all the time with conjectures they aren't sure if true, and work out a lot of other interesting math based on whether it is or not. Something like Turing's Oracle machine gives lots of interesting math just assuming one could exist, even if it couldn't. It may be that there are things proved we can never come to a human understanding of, but still keep getting interesting math that relies on it that has aspects we can appreciate and enrich our knowledge from.
Well it looks like they will announce at least one other millenium solution soon. In the same link they say they have "made substantial progress" on another millenium problem. The rumor mill before that statement was Hodge is done and Birch and Swinnerton-Dyer is on its way out.
"I'd say advantage humans this time."
Well LLMs were instrumental in any account of what happened. It's just a question of which company's LLMs did the breakthrough, and most of us outside silicon valley don't care about that part so much. The NYU guy himself said without LLMs the solution is maybe 10 years away.
In the US you can always sue someone in civil court for damages you think they have caused you. For criminal charges you need to have violated the law, but in a civil case it seems it's enough to have suffered monetary or reputational harm which was the other person's fault whether intentional or due to negligence, etc.
IANAL, and I'd be surprised to see any lawsuit come out of this, but you certainly don't need to have "violated the law" to be on the receiving end of a lawsuit.
Every single thing these companies do is dishonest and every word that comes out of the lips of these company execs is a lie, what fantasy land are you living in in which anyone with any amount of power gets punished for their lies?
what benefit do they get from making the statement? they could just say nothing. saying it and having it be untrue opens them to legal issues that are not worth the risk for this nothingburger.
It seems logical since if one used chats in train, one would expect that there would be a delay before their use to get them the form appropriate for batch learning.
The only way the chat could have been used would be for Open AI to baldly violate their policies.
That said, sometimes it take very little information to point someone in a given direction, "I'm working on Navier-Stokes" said by someone with a given specialization might itself be very useful information.
All the interpretability research we have would not indicate that "LLMs have no mind". It seems to me you have a conclusion and are working backwards to justify it. I guess I just don't see where 'they have no mind' would logically follow 'they sometimes cheat'.
reply