Hacker Newsnew | past | comments | ask | show | jobs | submit | layer8's commentslogin

FWIW, I’m in the same minority(?) as you. I suspect that others just have lower standards of rigor.

Do you find yourself still trying to use AI tools for the synthesis portion of writing/coding, or have you gone back to the 'manual' ways?

I've pretty much completely dropped it for writing (but still use it for catching issues with clarity or logical flow afterwards). I've yet to get away from it for coding. Perhaps I keep trying with coding because I feel I'm doing something wrong (and because for a a little my performance was tied to usage of it...).


I have no pressure to use LLMs for agentic coding, so I haven’t tried that hard, to be honest. I use LLMs to generate initial code drafts and to perform logical refactors (those which classic mechanical refactoring tools are unsuitable for) that I then touch up or revise manually. Basically, areas where it actually saves time while still fully controlling the design of the code and reasoning through all aspects of the implementation. I don’t see how agentic coding can save time without giving up some level of diligence, coherence, and attention to detail, which I’m not willing to do.

Looks like “ClawGPT” was already taken as a name.

(Indeed: https://clawgpt.com/)


The average consumer doesn’t have a $100+ ChatGPT Pro subscription though.

Isn't that a vote for 'not nefarious' as they are not deploying it widely across their user base? Unclear on the point

My point is that your argument regarding the average customer doesn’t seem to apply. But neither do I believe in a nefarious motive.

What’s an active blogpost? Do you mean an active blog?

Yes. This one: https://bookofjoe2.blogspot.com/

Full disclosure: It's mine (daily since 2004)


> Take the QuickTime record button.

Arguably it’s visually centered, because the menu arrow on its right is changing the visual weight between left and right.

It may be a layout accident, but it might also be deliberate.


I doubt it. That was when starting an audio recording. When starting a video recording, it's off-centre in the other direction, despite the UI being largely the same!

He might have started feeling that way only after trying the VR headset.

> I never understood the code. You think it works a certain way, until you find out that it doesn't. What LLMs make possible is for me to say: find out all the ways this thing works. Analyze the different ways we can run this software, build a fuzzer, build property tests, and run this software in every scenario possible. Log full traces. Log all the outputs. Now, analyze each scenario for bugs. You can't do that by hand.

Testing isn’t the same as understanding the code, or proving (even informally) that it is correct. Having the LLM do all these things above doesn’t lead you or the LLM to understand the code, to logically reason about its behavior over all possible states and inputs.

“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.


You’re technically correct, but the vast majority of software has never been built to the kinds of standards you are describing. LLMs are not displacing that kind of work!

The person they responded to refers specifically to "healthcare, finance, automotive, defense, power plans, aviation, manufacturing", areas where I'd at least hope that we aspire to understand what the code does.

Also software like SQLite:

" SQLite is built using a DO-178B-inspired process. The testing standards for SQLite are among the highest for commercial software.

SQLite is open-source but it is not open-contribution. All the code in SQLite is written by a small team of experts. The project does not accept "pull requests" or patches from anonymous passers-by on the internet. "

https://sqlite.org/hirely.html

https://sqlite.org/testing.html


Having worked in software for healthcare, defense and finance I guarantee you we don't

Do you agree that this isn't a desirable state of affairs?

I have also done those three and we definitely did!

TLA+

TLA+ ensures your model is correct, not that the code is correct

I worked briefly in heath care and the code is so brittle and so poorly understood that almost everyone is afraid to touch anything and instead it's just layers and layers of stuff trying to patch around existing code.

I know this is true. But I don’t think “Some parts of important codebases are black boxes. Therefore it’s fine if all of that code base becomes a far bigger black box” sounds like a good argument.

Also, there was probably some human at some point that had some understanding of what they were trying to do and why. The black boxes generally get programmed around after they long left but at the time they had bugs ironed out over decades. (Yes I know sometimes true slop is done over a short period of time and the programmer leaves. But I’ve generally seen the black box built over decades instead).


In many codebases when you see this happening it is because someone reported a bug and the dev fixed the narrowest instance they could. Typically if(some specific case) do_fix() else do_normal(). Over time this fragments meaning in the source and makes it harder to reason about until it is a giant pile of slop.

The thing the dev should do that often doesn't get done is to take the information about the undesired behavior back up to the design/spec level and rework the definition to account for it. This doesn't normally get done because it takes longer and it is conceptually more difficult than the patch.

This is something that LLMs can accelerate significantly and the result is a shallower gradient of technical debt over the long term.


> back up to the design/spec level and rework the definition to account for it.

> his doesn't normally get done because it takes longer and it is conceptually more difficult than the patch.

This usually isn't done, because design/spec no longer exists.


Sometimes this also happens because the design/spec never existed, but hopefully not in the industries under discussion.

I worked briefly in aviation and the part I've seen was very understandable and easy to extend and modify in understandable way. Some parts were hard, but by necessity. We also had very good tests. But maybe that one software was just a good exception.

Healthcare has featured many a time on thedailywtf, because a lot is (or at least was) written in MUMPS[0]. An example: https://thedailywtf.com/articles/A_Case_of_the_MUMPS

[0] https://en.wikipedia.org/wiki/MUMPS


I worked for a reinsurance company and their main pricing tool is a huge brittle excel file full of spaghetti VBS code. And yet, they manage to underwrite billions.

When the derecho came through Iowa and many businesses were out of power for up to multiple weeks, the very large organization I was working for at the time in the insurance space got to discover just how many of their processes relied on machines sitting under people’s desks. Business critical servers and processing just chilling on a PC under someone’s desk. Also tens of billions of dollars in revenue a year.

yeah, the llm approach is incredibly wasteful wrt pretty much everything. Performance, RAM, Development (Tokens).

But it does give you surprisingly stasble rube-goldberg machines.

And thats basically what 95-99% of enterprises want from their software.

It annoyed me to no end when i began my career, but at this point ive accepted it and can definitely still have fun developing software with llms. As a matter of fact, as my perfectionism approach to software in my earlier years was never really appreciated... So i dont really mind the new MO.

I still occasionally hand write though, esp. at the dayjob where ive got super small token budgets while continuously being told to use more AI. But that's normal, employers usually give off bipolar vibes with multiple stakeholders wanting to advance each of their bonus package KPI of any given quarter


> yeah, the llm approach is incredibly wasteful wrt pretty much everything.

Everything except what matters most: human time.


Oh, it wastes so, so much of that...

well, depends on how rich you are

The vast majority of software is not that important. I don’t really care about easytag (which I use for flac metadata), but I do care about xterm and tmux.

>“Finding out that it doesn't” means that you didn’t properly reason through the code beforehand, checking all your assumptions against what the code and underlying systems are actually guaranteeing. This may be a matter of formal education (proving computer science theorems and algorithmic correctness in university), I don’t know.

We're not writing theorems, dude.

Except in the equally pedantic sense that every program is a proof to a theorem...

We're writing plain enterprise and web software, closer to CRUD than NASA.

If you said that even before LLMs 0.1% of teams "checked all assumptions against what the code and underlying systems are actually guaranteeing" in any kind of formal way, you'd be overestimating it.


I’m not talking about formal verification, but about diligent informal or semi-formal reasoning through the code, so that you can rightfully claim that you understand the code and will be unlikely to be surprised by its behavior. Having learned formal verification does train that form of exhaustive reasoning about properties of the program. This practice also has you structure the code such (and select your dependencies such) that you can reason about all relevant properties. This is perfectly applicable to what you’d call CRUD and enterprise applications (that’s half the projects I earn my living with). Testing and fuzzing are complementary, but not a substitute by any stretch.

GP is correct. Very few people were capable of even the informal analysis you are describing, and fewer did it. I'm not saying it's not valuable... just stating that, empirically, it rarely happened.

And layer8 is saying (two responses upwards by them) that this is a novel benefit of AI: it can do a particularly thorough and repetitive kind of fault analysis that is a real PITA for humans to do (by their nature, versus the nature of computers).

Who’s “we” here? Formal verification isn’t common, sure. But you don’t speak for all programmers. You might work on “plain enterprise and web software”. But there’s still plenty of other software out there that many of us work on. And lots of code being written for internal use (e.g., data analysis code) that needs to be correct.

Of course, even enterprise and web software benefits from a little rigorous thinking. It’s pretty wild that understanding your code and its assumptions and informally proving it works is controversial. But I guess that explains why most software I use has actively gotten worse over the years.


You were really looking for reasons to get offended huh

Your code is only as good as what you can prove. Understanding the code is not the goal, it’s only important insofar as it helps you evolve the codebase predictably and without bugs or regressions, and understanding is not easily measurable or transferable.

Moreover, when your codebase is hundreds of thousands to millions LOC, I question how much you can ever truly understand it at the level you’re saying.


Regarding the last part, the strategy is to not have everything depend on everything, to instead modularize with succinct interfaces, so that you can reason locally. Of course beyond a certain project size, there is no single person who understands every part in detail. But for every part you can have someone who understands it, and can reason about it in terms of the interface contracts with the other parts. It’s also not essential that every detail is still understood at every point in time, as long as it’s sufficiently documented. What is essential is that for every part someone did reason through it with the necessary rigor at some point.

> to instead modularize with succinct interfaces, so that you can reason locally

Okay but how does AI change any of that? You can still do that with AI.

> as long as it’s sufficiently documented.

AI definitely helps with that.

> What is essential is that for every part someone did reason through it with the necessary rigor at some point.

Why is that essential though? What if the person who reasoned about it dies or leaves? Moreover, why is it imperative the reasoning happens at the source code level?


You’re right, you need TWO people who have read and reasoned through the code.

In reality, you need people who understand the design of the code well enough that they can quickly find, read, and understand the relevant parts when something goes wrong.


>> to instead modularize with succinct interfaces, so that you can reason locally

> Okay but how does AI change any of that? You can still do that with AI

With your own code you reasoned about it which contributed to its stability. This meant that you could treat it like a black box. And if the abstraction leaked or was unstable, the code was still fresh enough in your head that you could evolve it and still preserve its invariants etc.

With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.


> With your own code you reasoned about it which contributed to its stability.

Okay but to what extent? People say this but there's no way to measure it really. Did you live through the 90s? People reasoned through all that code and it was very often quite unstable. I'm sure everyone involved with Windows ME reasoned about it quite a lot, probably elements of it locally were very sound, yet in totality it was an unstable mess.

What fixed that situation wasn't that engineers today are reasoning better than engineers in the 90s, but IMO better tooling. Which brings me back to: your codebase is only as good as what it can prove. If there's any question, I just show you the proof rather than appealing to my reasoning being sound.

> With unreviewed AI gen nobody ever understood or reasoned about the code, including the AI.

And? You haven't established reasoning about it is actually necessary and it certainly isn't sufficient.


It made sense when databases and programs were used almost exclusively locally. It still makes sense for local apps (e.g. local-first or local-only smartphone and desktop apps) who typically will automatically do the right thing that way based on the OS regional settings.

It only started causing widespread issues with the rise of cross-region internet SaaS. Database systems, language runtimes, and OS APIs are keeping the default behavior for backwards compatibility.


Yes and no, it was thought to make sense for "end-user-programmers" where it's helpful to be fully locale specific, I'm Swedish and my OS settings makes programs expecting comma (,) signs for decimal separation is something that's actually hit me today when copy-pasting between programs.

So in practice, while it was kinda useful to be locale/region dependant for some users it's probably been more trouble in the long run to be overly helpful.

C# was released after y2k, so they don't have the excuse.

Also, You're missing the biggest sin here however, locale specific time is OK, automatically allowing conversions/comparisons to points in time types without specifying timezones has in principle never caused anything but grief.


It's not even global, is it? It's session scoped.

> Adding a month with + INTERVAL '1 months' is timezone-dependent. […]

Adding months isn’t well-defined anyway, even when using date, for days of month > 28. I think it’s a mistake that systems generically allow such a computation (as opposed to application code implementing domain-specific business rules).


D. Richard Hipp had a great blog about this on sqlite.org quite a while back.

If you have a link that would be appreciated.

Looking... I can't find it with Google, but https://www.sqlite.org/lang_datefunc.html mentions it:

| Because the length of a month or year changes from one month or year to the next, ambiguities can arise when shifting a date by months and/or years. For example, what is the date one year after 2024-02-29? Is it 2025-02-28 or 2025-03-01? Or what is the date that is two months after 2023-12-31? Is it 2024-02-29 or 2024-03-02? There is no consensus on how to resolve this ambiguity, so the "ceiling" and "floor" modifiers (14 and 15) are available to let the programmer decide. If the next modifier after a time shift is "ceiling", then any ambiguity in the date is resolved by choosing the later date. The "floor" modifier resolves ambiguities by resolving to the last day of the previous month. The default behavior is "ceiling".


It’s not just a matter of there being different ways to do it, it’s also that none of the ways obey expected arithmetic laws like associativity. Adding four months and then subtracting four months can yield a different day than the starting point. Or adding three months and then one month can yield a different date than adding four months at once.

Many people like email for a similar reason: Every message is an independent item that can be moved and copied at will, organized into folders, filtered and sorted, moved between accounts and email clients, attached to other messages, saved and loaded as files, and so on.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: