Hacker Newsnew | past | comments | ask | show | jobs | submit | pietz's commentslogin

Using a to-do list app as the first example to advertise a new app framework is absolutely wild in Q3 2026.

Is it more fashionable to build an AI harness?

We should go back to petstore examples.

A blog in 15 minutes would be great.

That's fine, but a blog doesn't seem technically more demanding or interesting than a to-do list

wasn't the original ruby on rails demonstrating how to create a blog? I think it was a call back to that 2004 era....

I genuinely have no idea what this comment is about. Why is a to do list a bad example? It's common precisely because it's ideal: it shows a few different basic widgets, and it shows how to update state in a way that keeps them in sync.

Is it because it's clichéd? That doesn't matter so long as it gets the point across.

Is it because you object to having desktop applications that do a simple task like this? I remember when simple single-purpose desktop apps were common e.g. Card File in Windows 3.1. Also this is meant to be an example for both desktop and mobile. Are only SPA web pages allowed to do things like this now? What a depressing world if so.

Or is your comment satire? I honestly can't tell.

Edit: What would you allow as a first example of a new UI toolkit. (Not that Flet is new, it's had plenty of 0.x releases.)


I don't think there is demand for this framework in the first place.

- Cross platform frameworks are starting to become an anti pattern these days

- Hello world examples are losing relevance when I'm not the one writing code

- A To Do list is just the absolute least impressive thing you can showcase

From a marketing perspective, I want to showcase something that other frameworks cannot do and "this only took 10 lines of Python code" is a very weak sales pitch in 2026.


OK, I think your main objection is really that you don't like the idea of a cross platform UI toolkit. I don't share that view, but fair enough. But your first comment, complaining about a to-do list example, is a very confusing way of expressing that.

And you think the front page shouldn't show a basic example because you just vibe code everything anyway. That is a very odd objection - it would be very concerning if a code library didn't include a basic usage example on its front page. Also, guess what, some of us do actually write (or at least look at!) some of the code in the programs we work on. And even AI benefits from a clean API.

And your complaint about a to do list is that it's not "impressive"? It's not meant to be! It's a basic illustration of the API. It's like reading the Rust tutorial and objecting, "ugh, another programming language with for loops".


My main argument is that if you want my attention in a world where an AI agent can build me a native iPhone app in 30min, you need to do better than showing me a to do list app in the hero of your landing page. Your comparison doesn't hold at all. This is not a tutorial page called "first steps", it is the first showcase of what is possible with it.

Anyway, it's fine for us to just disagree here. You just sound like someone that likes to argue for the sake of arguing, and just like to do apps, I don't have time for that.


They didn't even show a prompt box with starry icon button!

Why?

Because it's literally the least impressive thing you can build and in a world where I don't write the code anyway, a "hello world" example has lost its place (at least for human viewers).

It's one of the standard benchmarks for web frameworks.

No clue why it's being done here.


Specially one looking so bad.

Thanks for the hard work, but the Tahoe limitation for the GUI app is absolutely ridiculous.

Yeah, I'm holding onto Sequoia until after Golden Gate is GA. I've heard nothing good about Tahoe.

same. "macOS Tahoe 26+" requirement is a bummer.

Are there still any reasonable arguments to be mad at OpenAI at this point? Looking at how everything unfolded, this seems to have hit them way harder then they deserved.

They said they started working on millennium problems after hearing a rumor someone had found a solution to one of them. That's just bad taste. It indicates negative motivation rather than positive motivation from the start.

Not someone. Their single competitor. During a time when they were validating their newest internal model. What would you have done?

Truly, if this is the biggest criticism left, they should be celebrated. While in reality, all of this has a bitter aftertaste.

So weird.


That is not the narrative given in the OpenAI announcement. They claim to have started working on it after "we heard rumors that two Millennium Prize problems had been resolved", on September 1st, that "we later realized was related to Levent Alpöge" on Sepetember 6th "After the completion of our full project and lean verification"

That does stretch the reader's credulity. I think most people find it far more likely the rumor was related to Anthropic all along.


I find the claims from OpenAI somehow more relatable and reasonable.

- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.

- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.

- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.

- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.


It’s fishy though that they heard one of seven problems was about to be solved and threw perhaps 15 million bucks at the right one.

[Edit: they said "two of": "On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. .. we launched an effort ... on all open Millennium Prize problems".]


Why in the world is that fishy?

Isn't that exactly what almost everyone would do given that they wanted to see how capable their model is and the tense competition they have with Anthropic right now? Stealing impressive headlines from your competitor is pure gold.


If they were willing to spend 105 million it, perhaps. But that would be surprising. It’s fishy because they spent something like 15 million on the right one.

Again, that is the point. There is rumours that this one thing would be solvable, so they focus on this, spending 15m on one specific thing, instead of 105m on many. How is that fishy?

I understood the rumors (as described by OpenAI) to be that “one of” the prizes was solvable. If they heard NS specifically was solvable and aren’t saying that, they’re intentionally obfuscating that.

Because that’s the one Anthropic was rumored to have solved.

"On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems."

That doesn't reject my claim. They just didn't name them in this post. It feels like, you're going through great lengths reading something into this.

No, sorry I was reading it correctly but they addressed the issue in the post itself. They make it clear they were aiming at all 7 problems looking for the "two" that were rumored to be solved. But they didn't go hog on NS until they made progress. See sibling comments. I misread it the first time to be a claim that they "heard one of 7 human intractable problems are solved and spent 15 million on the right one".

Didn't they say that they launched an effort to evaluate their model on all open Millennium Prize problems, and then narrowed down to Navier-Stokes as the most promising after seeing results on simplified versions of the problems?

Yes, you're right. It seems like they claim to have started broadly and narrowed it down based on some progress. That leads to different questions but does answer my initial "fishy" point.

Isn't that *exactly* the type of solution you'd expect from AI?

Move 37 comes to mind.


Only if you understand nothing about the difference between LLMs and AlphaGo.

In case someone is asking: THIS is what a launch article should be like. 10/10.

Mission accomplished. That's both cool and fast.


> That's both cool and fast.

and probably a barely modified knock-off of some github project that it trained on


You're so upset that you have to invent an imaginary hypothesis to make yourself feel better.


Yes, very imaginary to think that the code comes from pretrained data and copy pasting whole blocks. It's not like this is exactly how LLMs work.


Correct, that's not how LLMs work [0].

[0] - https://arxiv.org/abs/1706.03762


Are you a frequent user of LLMs? That "copy pasting whole blocks" mental model doesn't hold up to regular usage, in my opinion.


They probably saw that report years ago of copilot dumping out the fast inverse sqrt function, and assume that's all they can do. From experience most anti-LLM people have either never used them, or used them back in the 3.5-4 era and then never again, though you might have even more experience with those people than I do. :P


I am using LLMs daily and I am a Claude max Subscriber. I neither saw any reports about Copilot Dumping sqt methods.

Now that we established that everything you said is wrong, do you have other toxic opinions or was that it?


they're wrong but they are right that this isn't interesting


The bar could not be any lower these days I guess


I know everyone is benchmaxxing but this one feels one step too far. Doesn't DeepSWE have both public and private tasks? I'd love to see the diff here.

It looks more like Google execs losing their mind and pressuring researchers to put DeepSWE directly into the training set.


That's not being debated here. The initial reported numbers were false and this was simply pointed out. You're changing the subject.


Opus 5 medium has the same score as 3.8 flash on artificial analysis intelligence index.

Are you implying Google or Artificial Analysis are reporting false numbers? What's your source?


BTW you're comparing 3.8 flash high to opus 5 medium. 3.8 flash medium scores lower.


Flash models are on the order of 1/10th the size of Opus models, so some flex in the thinking level is fair.


When comparing closed models, the only thing that actually matters to anyone using them is some mix of cost and speed. Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.


Your comment is really strange, why are you defensive towards WarmWash when gemini flash 3.8 high is both 6 times faster and costs less, while having the same intelligence score as claude opus 5 medium?

>Considering how much memory a server is using, when evaluating models that you'll never have access to in order to host yourself, doesn't really make sense.

This entire sentence makes no sense given what is being discussed.

https://artificialanalysis.ai/models/gemini-3-8-flash

https://artificialanalysis.ai/models/claude-opus-5-medium


I was being pragmatic. These are closed models on closed systems that you cannot hope to host. They are only available as black boxes available over web APIs served by their owners. Within that black box perspective, that we're force to have, the size of the model is, quite literally, just how much memory that server is using.

intelligence/model size is not a useful metric for a black box user.

intelligence/cost and intelligence/speed is a useful metric for a black box user.

Yes, it's cool, but as a black box user, the amount of memory a model is using on a server that I do not own has exactly zero practical use to me.

Cheers!


Flash is just a name with no defined or consistent meaning even within labs, let alone between them. Considering both are closed weight, there is no way to truly assess how big the size delta between the two is. Then again, who cares about size, performance and end-to-end speed+cost are what matters along with task adherence, task assessment and so on.

Model size also can not be inferred by tokens/sec for a multitude of reasons, but to showcase two examples, Opus 5 and Sonnet 5, as well as Gemini 3.1 Pro Preview and 3.1 Flash have each very comparable output speeds when using the same deployment as a basis for comparison, despite it being very likely that within their generation, the former are larger than the latter. Feel the need to mention this, as I unfortunately stumble upon so many poorly reasoned, speculative hype post trying to infer model size via utterly unreliable metrics, not based in actual data.

It’s like comments below arguing about the reasoning levels not normalized to some metric (like cost, output token amount or duration) but just the labels or high, max, medium, etc. Those mean almost nothing even when comparing models based on the same pretrain (just compare GPT-5.4 to GPT-5.2), they mean less than nothing comparing different labs releases.


It's not totally a mystery

https://arxiv.org/html/2604.24827v1

The short of it is by using hard facts knowledge that is difficult to compress, and then quizzing models on these facts and calibrating against a bunch of open models, you can kind of feel out the size of closed models.


I really like that one, but it kinda highlights what I could have far better explained. Their 90% PI is three times in both directions. Between 3T and 24T for GPT-5.5.

That’s a massively wide, inaccurate and at best barely informative range, demonstrating that even the most well thought out method will yield little usable information.

Additionally, I got some private evaluation taking a similar approach towards gauging models in topics I’ve found either over or underfitted by labs. If we just used that to rank models (not get a potential size range but just a rough order) Thinking Machines Inkling would need to be lager than Fable 5.


> [...] shows an intelligence score of 59, the same as Opus 5 medium!

Nothing here is false, you are simply confused. You either didn't read what they wrote in its entirety or decided to reinterpret what they did write.


"Beating opus" is the false part, no?


Stop lying. mattlondon said "gemini-3-8-flash shows an intelligence score of 59" which is undeniably correct. You can't say that number is false. You're literally lying.

All you had to do is go hover your mouse over "Models" in the top bar, hover over Claude Opus 5 and and click on medium: https://imgur.com/mlRCrt1

When you do that you arrive on this page: https://artificialanalysis.ai/models/claude-opus-5-medium

The gemini flash page for reference: https://artificialanalysis.ai/models/gemini-3-8-flash

You have to be an incredibly dishonest person to see a 59 on both pages and say "the initial reported numbers were false and this was simply pointed out. You're changing the subject".


The irony of this article being fully AI generated...

Anyway, it's over for Perplexity. They never had a great a product and the only reason for using them, was when they offered Pro accounts for free. Many people joined. Me included. But with a "meh" product and the general AI business not being very sticky, they lost quite harshly.

I thought they might be able to make money as a search api/index, but this article closed the book.


They were way above their main competition at the time Google decided to ignore the entire open web but they were still focused on searching it.

Since then, they decided to change focus into answering questions, and didn't maintain the quality of search results.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: