Hacker Newsnew | past | comments | ask | show | jobs | submit | dwaltrip's commentslogin

The entire point of plain text is that it doesn’t have number formats, date formats, and so on.

This is the eternal tradeoff of specificity vs. generality.


More generally, it's not rich text. The two terms are a pair, where rich text defines additional structure around the characters in order to specify things like that, and plain text has none of it.

There's a way in which it can definitely cgeapen things.

When you are vibecoding or vibe engineering, you're moving so quickly that it's difficult to evaluate what you're actually building and why. This often significantly reduce the quality and usefulness of the result.

It seems very difficult to mitigate, even if you were aware of the phenomenon. I think that the speed alone plays a large role. I'm not sure how one can build quality products without careful reflection and thoughtful iteration throughout the process.


RL is part of the model’s training. It changes the weights.

What distinction are you drawing?


I think the distinction is between "improving the g factor" and "adapting a given level of g to perform certain types of work better".

tl;dr it changes the weights, it does not add new ones.

RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.

Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.

The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.


Hmm interesting idea. I’m pretty confident there is generalization and learning that occurs during RL that does make the model smarter. So I think the distinction doesn’t fully hold up.

Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.

For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.

I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.

This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.

Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.

Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?


Every man and woman and child for themselves! Cooperation and support are for weaklings.

More seriously, it seems like you have an uncommon view about good ways for people to relate to each other. You may recall that humans evolved to live in large, cooperative groups :)


There is a difference between allowing free agency while creating an environment to steer people away from bad things, versus banning said things.

> So as long as these websites aren't defrauding the customers with fake stats, there really is nothing wrong.

This is a moral stance that not everyone agrees with.


Ban this shit. Fuck anyone who works at these companies.

I like the cut of your jib. Do you have a blog?

The exploding rockets were all test launches. SpaceX is pretty good at launching rockets successfully:

> Rockets from the Falcon 9 family have a success rate of 99.57% and have been launched 703 times over 16 years, resulting in 700 full successes, two in-flight failures [...], one pre-flight failure [...], and one partial failure [...]. The active version of the rocket, the Falcon 9 Block 5, has flown 633 times successfully and failed once [...], resulting in a 99.84% success rate.

Source: https://en.wikipedia.org/wiki/List_of_Falcon_9_and_Falcon_He...


Wow. That was amazing... I've watched it like 10 times in the past hour

There's a lot of terrible things that should be said about Musk, but somehow I don't think this film will do a great job of saying those things.

Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: