Prompting them yes, suggesting potentially fruitful research directions and so on, but the actual research was conducted by hundreds of agents swapping millions of messages and using billions of output tokens over 88 hours. The result being a huge Lean proof: https://github.com/openai/NavierStokesAndEuler. It's not just possible for humans to manually guide such a process in a meaningful way. They can set the direction and attempt to understand the result, but they solution itself must emerge (or not) from the agent swarm.
So yes, the models do seem to be "solving" the problems themselves, but not necessarily in the way we think of mathematical discoveries happening. Academic mathematics has historically been resource constrained: There are a limited number of top-level mathematicians, and they only have so much time and brain power to spend. So when approaching a problem, they are essentially forced to be as efficient as possible, not just searching for a solution, but for one that can be achieved within their cognitive budget. This induces them to develop novel techniques and abstractions, and it is actually those techniques and abstractions that tend to be the valuable part for further research, not the proof itself.
An agentic swarm is like getting a single skilled mathematician, cloning them a hundred times, then locking them in a room with the single objective of solving a problem. No longer constrained by time or brain power, they can approach it differently, using pre-existing techniques to gradually build their way to a solution. This process might not require a single intuitive leap or new discovery, and the solution will not be simple or elegant, but they will probably get there. It is more like a process of intelligently guided search than invention.
The OpenAI team didn't make a Lean proof. They brute forced a counter example. The "other" team was doing what you described but they haven't "finished" their work yet. Also, their Lean proof was for a simpler version of the problem, not the full NS.
Also, OpenAI wanted the actual mathematician taken off the resulting paper. I'm not sure I would describe what OpenAI did as research. What the other team was doing does seem to be more like research but the hardware was still in those cases mostly brute forcing things and then doing something like a genetic algorithm to compose an actual proof based upon the results of a large set of brute force attempts.
Didn't OpenAI make a Lean proof?
https://github.com/openai/NavierStokesAndEuler/tree/main/Nav...
"This repository contains Lean 4 formalizations of the results presented in “Finite time blowup for Navier–Stokes” and “Finite time blowup for the Euler equation” by OpenAI."
The current UK government was also elected by democratic vote, so they have a mandate that postdates the referendum.
You can argue their mandate is invalid due to the first-past-the-post system, but in that case the original decision to hold a referendum, and all subsequent governments that negotiated the terms of the EU exit are also invalid, since they all derived from first-past-the-post elections.
The current UK government have a mandate with regards to running the UK more generally, but leaving the EU was its own very clear message. There was a vote in 1975 to continue EU membership, and again in 2016 [1]. It very much feels likely something that should only be undone on another vote.
I cannot fathom how people enjoy that book. I appreciate the writing (and love other CM books) but Blood Meridian felt like a gratuitous ordeal to me. I tried watching lit professors lecture about it too but I still just don’t get it.
I'm guessing, given they're relying on private financing, the ability to reuse assets from TP1 is the major reason to make it a sequel instead of a new IP.
Personally I liked TP1 a lot for the most part, but I agree that the ending was a horrible misstep. I'm not adverse to metafiction, and this is Gilbert after all, that's kind of his jam, but the sheer nihilism of it all left a horrible taste in the mouth. Ironically, for someone who cares so deeply about game design, it felt like an ending entirely borne of Gilbert and his fellow designers pleasing themselves, instead of thinking about how it would feel for the player.
That said, I'm still glad there's going to be a sequel. Hopefully it has more of what I liked and less of what I didn't.
They are not selling at a loss though. The real issue here is that businesses such as Microsoft do not exist to serve a market or make enough money to survive and pay their employees. They exist to increase the value of their owner's shares. This is predicated on an unsustainable model of continual growth, with the expectation that every market they operate in can produce high margins and endless expansion. But demand is not infinite and persistent high margins are the exception, not the norm.
Maybe it isn't a loss, but as an investor (never directly in them), I consider 3% profit margin a bad sign - at that return I'd prefer a savings account: FDIC insurance means that after accounting for risks the savings account is better. I know stock returns don't directly track profit margin, but that is one input into the complex consideration of stocks.
You are right, and this is why not everything should be a stock. Whatever they will cut to avoid emitting "a bad sign", may involve:
* firing people
* making services worse
* sacrificing their own future
Whether it actually does involve those things is effectively arbitrary, because the consideration of the "bad sign" is also arbitrary. If there is no objective value judgement of their operation there is no objective value judgement in their streamlining either, so all bets are off.
Overall they are not selling at a loss, but parts of what they're selling are being sold at a loss. The market is telling them to slim the f- down and get rid of those money-losing parts.
This is all normal and justifiable. Where is the logic that corporations need to preserve dysfunctional parts of their operations?
Corporations don’t need to, no. But they are systems, and it’s a beginner mistake to assume changing one part is going to affect the whole in some simple, predictable, logical way.
That sounds like an extremely lame excuse to preserve money losing activities.
I'll counter by saying that pruning off failing things is not only good, it's the core of capitalism. Creative destruction, as Schumpeter called it. You get efficiency by hunting down and eliminating inefficiency, redeploying the resources elsewhere.
> pruning off failing things is not only good, it's the core of capitalism
It's quite serious that you see "being good" as something inferior to "the core of capitalism".
Also, the core of capitalism is making money for private individuals, nothing more, nothing less. Whether that's done with or without "failing things", is really beside the point.
The core of capitalism is about where ownership of capital (value producing assets) resides, that is, by private individuals.
What private individuals choose to do with their capital, chase infinite growth and profit or sit on it, is up to them. This is as opposed to say, state ownership of capital.
People confuse the stock market with capitalism. You don't need a stock market for capitalism to function. Publicly traded companies in the United States are legally bound to maximize profit (Dodge vs. Ford Motor Co.)
Error is rampant. I saw a quote saying the difference between a good business and bad one is the good one makes the right decision 60% of time, the bad one 40% of the time.
So errors abound, and have to be subsequently corrected. This correction process is as natural to capitalism as breathing is to you being alive. Without it, things would rapidly grind to a halt. We see this in sclerotic centrally planned economies where errors persist for much longer.
Some of the other examples you could maybe say are just poor industrial design or bad execution of a potentially workable idea, but the Intel MBP thermal debacle was definitely the most egregious example of Apple blindly pursuing form-over-function. They set a goal for ultra-slim laptop forms that the components they were using simply could never achieve. It would have been obvious from prototyping that temperature was a major problem, but whatever decision making process they had in place at the time overrode it. The best you can say is they seemed to have learned from their mistake.
> It's since november 2025, the so called "inflection point", that I'm still wondering for who coding agents become "really good".
The answer is "for lots of people, but not you".
You're doing a vague impression of being fair and even-handed, arguing for non-polarization, but underlying everything you're saying is an obvious attitude of poralizing superiority: That _your_ personal experience with AI is the real truth. That _your_ codebase is more intricate and more challenging than what other people are doing. That everyone else is being led by a "marketing hype train".
So yes, the models do seem to be "solving" the problems themselves, but not necessarily in the way we think of mathematical discoveries happening. Academic mathematics has historically been resource constrained: There are a limited number of top-level mathematicians, and they only have so much time and brain power to spend. So when approaching a problem, they are essentially forced to be as efficient as possible, not just searching for a solution, but for one that can be achieved within their cognitive budget. This induces them to develop novel techniques and abstractions, and it is actually those techniques and abstractions that tend to be the valuable part for further research, not the proof itself.
An agentic swarm is like getting a single skilled mathematician, cloning them a hundred times, then locking them in a room with the single objective of solving a problem. No longer constrained by time or brain power, they can approach it differently, using pre-existing techniques to gradually build their way to a solution. This process might not require a single intuitive leap or new discovery, and the solution will not be simple or elegant, but they will probably get there. It is more like a process of intelligently guided search than invention.
reply