I genuinely have no idea what this comment is about. Why is a to do list a bad example? It's common precisely because it's ideal: it shows a few different basic widgets, and it shows how to update state in a way that keeps them in sync.
Is it because it's clichéd? That doesn't matter so long as it gets the point across.
Is it because you object to having desktop applications that do a simple task like this? I remember when simple single-purpose desktop apps were common e.g. Card File in Windows 3.1. Also this is meant to be an example for both desktop and mobile. Are only SPA web pages allowed to do things like this now? What a depressing world if so.
Or is your comment satire? I honestly can't tell.
Edit: What would you allow as a first example of a new UI toolkit. (Not that Flet is new, it's had plenty of 0.x releases.)
I don't think there is demand for this framework in the first place.
- Cross platform frameworks are starting to become an anti pattern these days
- Hello world examples are losing relevance when I'm not the one writing code
- A To Do list is just the absolute least impressive thing you can showcase
From a marketing perspective, I want to showcase something that other frameworks cannot do and "this only took 10 lines of Python code" is a very weak sales pitch in 2026.
OK, I think your main objection is really that you don't like the idea of a cross platform UI toolkit. I don't share that view, but fair enough. But your first comment, complaining about a to-do list example, is a very confusing way of expressing that.
And you think the front page shouldn't show a basic example because you just vibe code everything anyway. That is a very odd objection - it would be very concerning if a code library didn't include a basic usage example on its front page. Also, guess what, some of us do actually write (or at least look at!) some of the code in the programs we work on. And even AI benefits from a clean API.
And your complaint about a to do list is that it's not "impressive"? It's not meant to be! It's a basic illustration of the API. It's like reading the Rust tutorial and objecting, "ugh, another programming language with for loops".
My main argument is that if you want my attention in a world where an AI agent can build me a native iPhone app in 30min, you need to do better than showing me a to do list app in the hero of your landing page. Your comparison doesn't hold at all. This is not a tutorial page called "first steps", it is the first showcase of what is possible with it.
Anyway, it's fine for us to just disagree here. You just sound like someone that likes to argue for the sake of arguing, and just like to do apps, I don't have time for that.
Because it's literally the least impressive thing you can build and in a world where I don't write the code anyway, a "hello world" example has lost its place (at least for human viewers).
Are there still any reasonable arguments to be mad at OpenAI at this point? Looking at how everything unfolded, this seems to have hit them way harder then they deserved.
They said they started working on millennium problems after hearing a rumor someone had found a solution to one of them. That's just bad taste. It indicates negative motivation rather than positive motivation from the start.
That is not the narrative given in the OpenAI announcement.
They claim to have started working on it after "we heard rumors that two Millennium Prize problems had been resolved", on September 1st, that "we later realized was related to Levent Alpöge" on Sepetember 6th "After the completion of our full project and lean verification"
That does stretch the reader's credulity. I think most people find it far more likely the rumor was related to Anthropic all along.
I find the claims from OpenAI somehow more relatable and reasonable.
- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.
- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.
- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.
- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.
It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.
[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]
Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.
If they were willing to spend 105 million it, perhaps. But that would be surprising. It’s fishy because they spent something like 15 million on the right one.
Again, that is the point. There is rumours that this one thing would be solvable, so they focus on this, spending 15m on one specific thing, instead of 105m on many. How is that fishy?
I understood the rumors (as described by OpenAI) to be that “one of” the prizes was solvable. If they heard NS specifically was solvable and aren’t saying that, they’re intentionally obfuscating that.
"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."
No, sorry I was reading it correctly but they addressed the issue in the post itself. They make it clear they were aiming at all 7 problems looking for the "two" that were rumored to be solved. But they didn't go hog on NS until they made progress. See sibling comments. I misread it the first time to be a claim that they "heard one of 7 human intractable problems are solved and spent 15 million on the right one".
Didn't they say that they launched an effort to evaluate their model on all open Millennium Prize problems, and then narrowed down to Navier-Stokes as the most promising after seeing results on simplified versions of the problems?
Yes, you're right. It seems like they claim to have started broadly and narrowed it down based on some progress. That leads to different questions but does answer my initial "fishy" point.
They probably saw that report years ago of copilot dumping out the fast inverse sqrt function, and assume that's all they can do. From experience most anti-LLM people have either never used them, or used them back in the 3.5-4 era and then never again, though you might have even more experience with those people than I do. :P
I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.
It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.
When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.
Your comment is really strange, why are you defensive towards WarmWash when gemini flash 3.8 high is both 6 times faster and costs less, while having the same intelligence score as claude opus 5 medium?
>Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.
This entire sentence makes no sense given what is being discussed.
I was being pragmatic. These are closed models on closed systems that you cannot hope to host. They are only available as black boxes available over web APIs served by their owners. Within that black box perspective, that we're force to have, the size of the model is, quite literally, just how much memory that server is using.
intelligence/model size is not a useful metric for a black box user.
intelligence/cost and intelligence/speed is a useful metric for a black box user.
Yes, it's cool, but as a black box user, the amount of memory a model is using on a server that I do not own has exactly zero practical use to me.
Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.
Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.
It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.
The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.
I really like that one, but it kinda highlights what I could have far better explained. Their 90% PI is three times in both directions. Between 3T and 24T for GPT-5.5.
That’s a massively wide, inaccurate and at best barely informative range, demonstrating that even the most well thought out method will yield little usable information.
Additionally, I got some private evaluation taking a similar approach towards gauging models in topics I’ve found either over or underfitted by labs. If we just used that to rank models (not get a potential size range but just a rough order) Thinking Machines Inkling would need to be lager than Fable 5.
Stop lying. mattlondon said "gemini-3-8-flash shows an intelligence score of 59" which is undeniably correct. You can't say that number is false. You're literally lying.
All you had to do is go hover your mouse over "Models" in the top bar, hover over Claude Opus 5 and and click on medium: https://imgur.com/mlRCrt1
You have to be an incredibly dishonest person to see a 59 on both pages and say "the initial reported numbers were false and this was simply pointed out. You're changing the subject".
The irony of this article being fully AI generated...
Anyway, it's over for Perplexity. They never had a great a product and the only reason for using them, was when they offered Pro accounts for free. Many people joined. Me included. But with a "meh" product and the general AI business not being very sticky, they lost quite harshly.
I thought they might be able to make money as a search api/index, but this article closed the book.
reply