Hacker Newsnew | past | comments | ask | show | jobs | submit | mbrumlow's commentslogin

The problem is not vibe coding. The problem is the people doing this were never good software engineers before vibe coding.

Simply put you still need good practices to produce working good code with AI.

Oddly what I have found best is patterns that many would cringe and demeans unit test and integration test.

Yes I still vibe unit test and integration test, but more than ever I have to put on my PM and QA hat, even more so with web dev.

For me the only part of the process that has changed is I don’t write the code. I still setup environments to run the code, I still manually poke at things I suspect will be fragile. But most importantly I still always run the software before shipping it.


> Yes I still vibe unit test and integration test,

False sense of security. What's stopping the agents from doing an elaborate form of "Assert.IsTrue(true)" in some/most/all of the tests since you are vibing it and not reviewing it?


Exactly what happened in a PR from a few months ago, from the beginning it felt wrong, pytest was moved from the dev group to the main import section. Then imported in our models. To add ignores everywhere.

Agreed. If you don’t want to read the code that’s fine, but you have to read the testing code. Otherwise you have no idea what’s going on.

Yeah, you have to have some level of involvement. It's crazy that we are actually talking of 100% dev involvement to literally 0% involvement.

If you're working with agents that severely misaligned you'll probably have issues way beyond fake tests.

It seems to me that chances of misalignment are a function of accumulated usage time. Eventually, it will happen.

> The problem is not vibe coding.

Except that it is the enabler. Without AI there are no 5K lines of code that no human has ever seen. AI introduces a new category of problem that did not exist before.


Yep agree.

As much as I like to put power in the individual, this has become a huge nightmare for those of us who were kind of decent programming.

This was NOT the case 5 years ago. At all.


Well said. The problem is that vibe coding can paper over bad engineering practices for awhile. And results look good, and when you raise concerns people look at you like you're dumb and trying to slow things down.

But the underlying bad practices _will_ bite the same they would have before AI, you just don't know when.

It's a really frustrating state of the industry to be in.


Models are the worst now that they ever will be. Frontier models do a phenomenal job of employing good engineering practices and code style compared to where they were mere months ago. Unassisted programming is dead. Programming as a human activity will be on a timescale of years (along with a vast swath of other knowledge work).

It's like the casino, crypto gambling, or eating junk food. We're wired for that short-term gains and few of us can escape it. I understand that, just sad that the casino slot machine had to hit our craft.

Couldn't have said it better myself. Issue at hand is if you're working in an environment that enables such people and even encourages that behaviour. Suddenly, their voice is the same where it was never even a whisper before, and you can't do anything about it. It is extremely demotivating and I would actually propose a more radical approach to it - either embrace the shit in full and commit to it as well or really do engage into another company where there's a better filter for hiring. The latter are usually smaller companies because they can't afford no to.

I actually quit bigger companies and SV contracts because the amount of hype was way overblown. Nothing could not be AI, really. Think about how crazy it is that Jensen Huang is saying "AGI is here", seriously. I first thought malice, now I think it's just plain religious Zeal.

We're not a company, it's just me and my pal. I keep 50% of the product and profits and he pays me to do that (because I have way more expertise than him)

curious question where you work, the AI craze is not happening?


curious question where you work, the AI craze is not happening?

of course it is. There's no escape IMO unless, as you've said, go solo or with a small team. I just don't know of any company of a notable size that's not into it full force. It's both disgusting and amazing to look at really. Give it a few years and if it fizzles out I'm curious what will happen. If it turns great, I'm even more curious what will happen. Weird times.

Just "yesterday" we were all careful what and how we do stuff, you have now people who barely understand things slopping together OS and people use it blindly (Omarchy), and then shocked pikachu faces when giant gaping holes are noticed in it.. and people still do not care. Amazing trainwreck of a timeline, hah.

I'm seriously considering pivoting myself to security.


if I got $1 for every time me or anyone I know prompted AI, got the wrong answer, and people claimed "you need good practices / get better at prompting" I'd be a millionaire

The problem is these companies literally say "NO NEED TO LEARN ANYTHING JUST VIBE". Ton of ads on social media and YT claiming that.

Like my mom doesn't even know what a PC, a Macbook or a Linux computer is and we're telling her she can make apps?


This does not look like C

Looks like Rust.

Would use this at a fairly large company if we had windows and Mac support, probably could get budget to pay for the features it’s missing.


That sounds like you should get in contact with Rui Ueyama. If anybody asks, someone on HN said it seemed like a no-brainer.


Absolutely. Rui is awesome. He's always been awesome. I was his director and then vp for a long time (also replaced by awesome people, thankfully). The day he left to make a go of mold and such I was sad for us and super excited for him.


Hm - are the linking requirements for PE and COFF similar enough for that to work?


No. I wrote my integrated linkers with rcc for pe/coff, elf and macho by myself, and macho was the worst to be dealt with. elf was pretty easy, easier than pe/coff.

Mine is also much faster than mold/wild, because I do much less work, and I have all the objects already in memory. Similar to tinycc. Only when someone needs a very unusual unsupported linker feature, I have to fallback to an external linker, which is ~100x slower. The less work it has to do, the faster. I dont appreciate the unix habit to use intermediate files all along, which needs additional readers and writers.


Are you sure? What if I painted the image , but happen to have a photo realistic capabilities ?


Ignorance of the crime isn't excused BTW. You do have a flake chatbot to ask the law about


The only people being ignorant at the ones saying deepfakes are illegal.

It’s more complicated than what is being touted as fact.

In most jurisdictions creation is legal, it’s publishing them that is not. In some jurisdictions threatening to publish them is also illegal. Most jurisdictions that have a creation policy apply to doing deepfakes they sexualize minors, which likely would have been illegal to possess under other laws.

Then we have a consent issue. Creating and publishing them could be entirely legal if the subject consents to such.

So yah. It’s not “deep fakes are illegal, let’s arrest everybody !!!”

To be clear, creation is likely legal, publishing not so much.

Additionally, intent is important with regard to AI hacking.


Same here, 20+ years, no degree. but have a slightly different view about folks with CS degrees.

It really comes down to the individual. Those who are good with CS degrees would likely be good without them. The majority of those who picked CS as a degree because “it pays well” are truly awful. Only a small handful end up enjoying and understand how to write code and make good software.


I agree about the existence of people who really just aimed for the grades (sometimes dubiously) to unlock a career. There were also enthusiastic folks with whom I learned a lot.

While I'm an interested self learner, I significantly benefit from structured learning. I don't think I would have had the breadth of depth of knowledge if I didn't do a CS degree since I didn't think I would have done such a variety of focused projects and homework. I would "shop" for good teachers and wait until they taught to take a course. I would have lots of questions in class. At the end of many of my semesters, I genuinely felt like Neo's "I know kung fu" - knowledge given on a platter. I didn't think I would have left covering undergrad subjects if I didn't, well, feel the need to earn my keep and at least aim for higher level studies and get assistantships for initial financial independence.

That being said, this was all before the YouTube era. Although, YouTube is a double-edged library of distractions and amazing educational material.


Oh man, I have that too. I am always told by my best friend "well, you go it right, but entirely wrong words" -- or something to the effect I got the feeling, notion, or concept right, but the wrong words.

This has affected me a lot though, I know what is going on in songs, but can't ever remember the right words, but still some how end up sining stuff that means the same thing :/


Rust is better. It just is. Go is not bad. But as a long time go advocate, the hurdle for my teams using rust is gone, and thus everything is now rust.


Personally I find Rust a lot harder to read than Go.

If you're going to have an agent write most of your code readability is very important.


I'll qualify this from my POV (which may be different than GP's).

Go's historic maintainability strong suit has been its simplicity and consistency. The syntax is, relatively speaking, lightweight, the language invites complexity through composition, and information density for any unit of code is typically quite low (which isn't necessarily a bad thing).

In my opinion, though, these are all drawbacks, and Rust addresses all of them. It's syntactically and semantically much heavier, leading to its oft-maligned steep learning curve. It has, uniquely among the major languages, I think, a syntax for expressing variable lifetimes (with its own unintuitive semantics). It stuffs lots of abstraction into a hodgepodge of terse semantics and punctuation.

It sucks to read, until you get really used to it. Then it tends to read really quickly, and, at least for me, it's easier to reason about a conceptually-broad piece of logic if I don't have to jump between different locations in a file, a module, or a package to do it.

With Go, I find it more difficult to get into a flow state, and easier for my eyes to glaze over when looking over large diffs.

It's not lost on me that these are purely subjective arguments, though. My preference remains with Rust, and that goes back to before I used LLMs.

I'm also aware that Go is very prescriptive about how you write it; it's explicitly opinionated, and Rust doesn't have that. It means that most Go code bases will look more alike. I consider this an anti-feature; I believe code should be able to conform to the problem space or product and a good team will find the best way to do that.


Yeah, Go is easy to read in the same sense that English limited to its ten hundred most common words is easy to read (https://xkcd.com/1133/). Whether that nature is helpful or harmful to LLMs is an interesting question.


It seems very obviously detrimental to me (in the exact same way it's detrimental to people); e.g. use proper jargon with an LLM and you find it is suddenly an expert. The LLM has no trouble at all perfectly fluently using macros or monads or whatever thing people are afraid of to write simpler, more concise code that directly expresses the business logic in a fully type safe, high performance way that the compiler can introspect for even more information. Go of course lets you do none of those things and can only ever be used for "beginner code" by intentional design.


For whatever reason, maybe not even logical ones, Go repulses me. I don't know why exactly, I like the idea of Go, but the aesthetics rub me wrong.

I think it's the use of pointers and "if err != nil {}" error handling spam. It reads as a highly compromised imitation of Python and C rather than a solid execution of some other idea.

Rust is not the most beautiful language out there but it doesn't trigger any such reaction for me. The ? operator and "match", which I use constantly, more than compensate for some of the sigil noise which I barely need to look at much less write most of the time. So Rust wins on that comparison for me.

The "func name() -> retval" syntax also grew on me. I like the fact that Python type annotations copied that approach, and C-style declarations look ugly to me now. Same with C-style /* */ comments.


Because Go is kind of a Steampunk design. The creators collectively ignored decades of PL developments. What that nets is kind of a C without sharp edges, but can't be used where C can.

Rust is definitely jarring to look at, in the same way that decoding some strange C declaration can be, in ye olde days when you had to float all this context in your mind while doing work. But with modern tooling who cares: "explain this lifetime to me"


For me it's mainly the if err != nil {} stuff and the fact that everything is package scoped (C-style enums and constants).

You can pretty clearly see the limitations if you read, for example, the type of code the Protobuf compiler generates when trying to compile Protobuf/gRPC enums or structs into the way-more-limited Golang type system (this is despite the two being designed to work together). And it could really do with algerbraic data types and other modern programming language features.

Also the type system does have a couple weird behaviors that seem straight out of JavaScript. Like the difference between struct and interface nil for example:

```

var buf *bytes.Buffer = nil

var out io.Writer = buf // now out is nil

if out != nil {

    // This block will execute because out is not nil

    out.Write([]byte("crash")) // This line will crash because out is nil
}

```

Many things about the language almost seem to be designed to simplify the implementation of the compiler rather than to benefit the developer experience.


This is a classic pitfall. Every Go developer knows that using this assignment is discouraged, because interfaces are structures composed of two fields. You are assigning a pointer-type value to an interface. (_type = *bytes.Buffer, data = nil)

`out` is then no longer equal to `nil`.


It's a matter of expressiveness, Rust expresses more.

E.g. make a table that's 3x3 is easier to read (Go), but the equivalent line in Rust would also include material, angles, height, etc. because the type system encodes much more information.

Though I always found Go to be significantly harder to read than Rust. Sure Rust has some crazy syntax at the edges, but Go makes it very hard to know where imports come from (and thus what they do), and the imperative style + lack of clarity about mutability makes code much harder to reason about.


Counterpoint: I find Rust easier to read.


I have an agent read most of the code as well. The agent explains things to me in plain english.

The default sentiment is humans should read code it's more progressive and a leap of faith to start giving that up.

Obviously, I get why you feel humans still reading code is important, but if you look at the progress of AI for the past couple of years, that gap is closing. The trendlines speak of a future where it becomes less and less important.

This was exactly what happened with writing code. Now most people don't write code.


> Now most people don't write code.

I use LLM daily to write code for and "with" me, I also write code without LLM. Most people I come across mix it up. A few do it all by hand, and equally few all by LLM I would say. Is that just in my corner of the world?


The trendlines are moving away from this. It's all happening so fast that not every company is on the same page, but from what I see we are quickly converging on not writing anymore code.

My entire company for example does not write a line of code. We manage agents and that's it. Many, many, many companies and people are already doing this.


I have a few utility go codebases that I simply do not read at all - but it's internal tooling so there's literally no point in reading it when the LLM can modify it in seconds to do new things.


I don't read all of the code produced by my agents any more, but I like to reserve the ability to do so if I run into a particularly confusing bug, or for any code that's security adjacent.


Same. But usually if I need to read code, I end up telling my agent to summarize it for me.


Might be unpopular, but agents write too much code for humans to read in any meaningful time frame. Using agents to generate code to then require humans to slowly consume it defeats a lot of the speed you gain from AI.

I my self and teams members are slowly reading less code and requiring agents to prove things work the way we want in other ways.


Personally, Rust or Typescript both happen to be better than Go for me. TypeScript has better type-safety and tooling for user-facing apps, and Rust has better type-safety and tooling for algorithmic stuff or stuff that needs to run fast.


Ignoring performance for the moment (because most situations are bottlenecked on something else), why is rust better?


I agree, but this doesn't justify anything. Saying rust is better because it "just is" won't convince anyone. I'd like to know why you think it's better.


I guess the difference in compile times doesn't matter enough?


idk. In my experience the build/compile experience has been far worse esp for fast iterating. Even concurrency models did not seem as intuitive as Go's. Im no systems expert - have deployed practical and performant distributed systems though.


Zig seems to have more closely aligned with what Go devs prefer.


Go doesn't have memory safety issues because of its GC, while Zig has UB problems. Zig might have slightly better performance, but I don't think choosing a language without memory safety is a good idea


What is UB?


UB stands for undefined behavior—I've written about it in detail on my homepage[1] [1]https://www.makonea.com/en-US/wiki/undefined-behavior


undefined behaviour


can you elaborate?


Rust compiler is a tyrant. Type system is strict. Borrow checker is relentless. LLMs can't slop too much without being beaten up by the compiler.


This is true. I'd like metrics on this though. It could be that LLMs find go easier so they end up writing better code and rarely hitting static errors like a human would in rust. IT could be through scientific measurements that the benefits of static checking could be negligible for LLMs.

No way to know until someone does the science on this. Until then it's just people saying that more static checking is better. But I do think, anecdotally, python is horrible for LLMs.


Until we can measure slop accurately it's all guesswork


True but if the reviewer doesn’t have an intimate understanding of rust then the fact it can’t “slop” is no different than unreadable slop.

Go is simple, no “magic” marcos or meta programming even with just a little programming in any language it’s not hard to understand what the go code is doing.


That's simply not true. I wrote a piece of software in Rust that is non-trivial, robust, 50k lines of code, and 100% LLM generated, used by four people productively with only one or two minor bugs in the past two man-months


question did you review the code or did you test it was working? these are different things. and if you did review it, could an engineer without deep Rust experience have reviewed it just as effectively?

I have no doubt that you can get a LLM to write working bug free code in any language but that is not the topic of the article or my comment.


Probably at 50x the cost LOL


except Rust is HARD while GO is super easy.


Rust makes you solve many of your problems upfront, which is a nice feedback loop for using with an LLM. Go does much of this too, but I feel Rust is more experessive and takes the frontloading a bit further.


It's not that hard.


I just do all that with Claude CLI :/


Policies like this will result in the death of the branded software.

As we move forward it will be easier than ever to just maintain and keep your fork of software with the changes you want or need. No more approval, bureaucracy, or arguing. Just tell the AI agent want you want changed and you have it.

This will be used for huge things too. Like maybe you want a specific fork of Java that only supports for each iterators, goodby linters, hello compile time error.


You truly envision a future where every program is written in a custom programming language, for a custom operating system, for a single user who will now be in charge of understanding and maintaining it forever? That user being your grandma, your baker, your CEO?


I do.

Software as a list of requirements and that's it. The local LLM appliance everybody has taking in a document specifying hardware, interfaces, and requirements and spitting out software changeable locally via conversation with its users.

In the same way you have a cookbook with recipes to make dinner instead of ordering out.


> Software as a list of requirements and that's it. [...] a document specifying hardware, interfaces, and requirements [...]

For that, you'd want the list of requirements and hardware documentation to be written in a precise, formal language. That's no different than writing them in a programming language (though a declarative one, instead of the more common imperative ones).

I've in the past (way before LLMs existed) thought about automatically generating device drivers from hardware documentation. But besides the need for very precise documentation, hardware never works exactly as documented; a human-written device driver can avoid problematic areas (perhaps even by accident), while a computer-written device driver would end up exploiting every corner case of the documentation.


Requirements ARE written in a precise, formal language. These days most people don't actually have any contact at all with real requirements though as practiced by professionals.

There is quite a bit of distance between the exactness of any human language and an sort of programming languages. When you have a language model loaded with software engineering best practices you do not need the exactness of a programming language to describe desired behavior.

Hardware documentation is indeed often lacking in plenty of ways but in a world where writing your own software through agents is commonplace, the hardware manufacturers (or the community) would make testing that documentation to find the problems an important part of hardware development.


Most people don't bake their own bread even if it's more approachable, and cheaper then vibecoding. Bread is just a recipe one may say. But in modern times have we ever witnessed disappearance of specialization? I don't think so.


That is a cool perspective to consider.


> easier than ever to just maintain and keep your fork of software with the changes you want or need

I don't imagine this is practical or desirable for all situations. Good software is built from being battle tested by many users in many environments. Even with the advancements in AI tools, I don't imagine they'll become omnipotent anytime soon.

> No more approval, bureaucracy, or arguing

For software that can kill people or substantively affect someone’s life in a negative way, the bureaucracy is there for good reason. I don't think anyone should want someone at Phillips to vibe code the control software for an X-Ray machine or an employee at CrowdStrike vibe coding the next update before pushing it out to millions of machines.

We are forced to endure low-quality software because there is little or no accountability. I can only imagine what you propose would make an already poor situation worse.


I have my doubts about this. Brands are a about reputation, and reputation strongly affects responsibility when business decisions are made. If a manager decides to vibecode a solution which eventually fails, said manager will be punished. If instead a mainstream software is bought, and it fails - well, everybody has the same trouble around, right? It's not a personal mistake anymore. I suppose there are niches where branded software may give way, but I don't think it's gonna be a general trend.


This just reflects how absolutely clueless you are about the amount of attention to detail that the OpenJDK folks put into developing the language and ensuring that it works for all its users (which are serious users delivering actual value.) And I say this not even being a Java programmer myself, just an envious C++ dude watching from the sides.


> maintain and keep your fork of software with the changes you want or need

Then you'll have the same problem everyone who forks a piece of software ends up having, sooner or later: as the original evolves, keeping your fork up to date with the upstream changes becomes harder and harder. The bigger and more invasive the changes are, the harder synchronizing with newer releases become.


Someone brings this up on just about every AI-related thread. I think it's nonsense. Nobody wants to maintain a fork of any remotely complex software, not even with AI. And in a corporate setting, nobody wants to use your custom fork; they just want to use the standard software they already know with the quirks they've already learned.


This is far more likely to happen in a corporate setting with competent engineers who want to build/solve now versus relying on others. I've seen it 100% of the time with varying results.


The idea of someone forking Java with vibe code nonsense and expecting them to be able to maintain is laughable.


Except. . . (accept)... SaaS solves the problem of needing one piece of software to communicate between multiple (many) users and/or other pieces of software.

Can't picture a functioning world where every piece of software is custom and requires factorial amount of AI comparisons and reviews to patch the API to communicate. In fact, it's impossible! There's not enough compute to handle a factorial explosion.

I really doubt SaaS going anywhere.


Given how much of the industry is SaaS that would be such a self own! Wow, fantastic, you can build all your software in house, zero dependencies. Wait, your users can do the same? And they don’t need to pay you anymore for any of your cool services because they just asked their agents to recreate your infra from scratch? Interesting, truly the future of humanity


> Wow, fantastic, you can build all your software in house, zero dependencies. Wait, your users can do the same? And they don’t need to pay you anymore for any of your cool services because they just asked their agents to recreate your infra from scratch?

The cherry on top is that OpenAI and Anthropic brainwashed your coworkers and your company's C-suite into uploading the entirety of the "proprietary" codebase onto their servers thousands of times per day over the last three years.


No need to brainwash, they did it willingly and with a smile, tweeting about how they are part of the revolution while telling us all how we will be left behind


Relevant xkcd: https://xkcd.com/605/


I think it's even more pervasive than that. Why bother forking software at all? At the point in which code generation is meaningfully trivialized, software becomes entirely disposable. Anything you want, have a model spin it up. You don't even need libraries, the model can just make everything in-situ, who cares? Why on earth would I ever want to use SQLite if I have access to a sufficiently advanced code generator which can generate me a similarly high quality database system, with the added benefit of conforming to whatever my problem domain is, conforming to whatever branch of database theory I want?

Even SaaS isn't safe. I don't even have to describe your product to my system, I just have to give it a harness with access to the interface and have it replicate it locally. Frankly you can probably already prompt for that.

The only thing holding this future back right now are pricing problems and code generation quality. Both of those barriers are constantly being knocked down. We might never arrive at that future, but it's definitely a higher probability than solving AGI's scaling issues, and would arrive much sooner for technical users.


> Why on earth would I ever want to use SQLite if I have access to a sufficiently advanced code generator which can generate me a similarly high quality database system

Because SQLite has 10k requirements that wouldn't even cross your mind to write down, but 80% of which are useful to you.


The nice thing about natural language is that nesting semantic layers is free and arbitrary, and far more tractable than in a formal grammar. Every natural language is like coherentist ω-order logic. Effectively, I don't have to write the 10k requirements. I only need to provide a sufficient metatheory that can be extrapolatable to those 10k requirements, and that can include embedded theory I did not write myself but am familiar with enough to invoke, as well as refinement criteria ranging from the fuzzy to the explicit with priority weighting parameters to describe the shape in which I want the search space pruned.

This isn't anything new or particularly interesting. It's the entire basis upon which ILP demonstrated generality. A metatheory to synthesize 10 trillion rules isn't even scratching the surface of what you can reasonably do. The key was finding out the tractable semantics for actually computing it in reasonable amount of time, which right now is looking decidedly like informal semantics was the answer the whole time.


Not sure what this is a parody of, but it's hilarious, thank you. Keep it up!


I accept your concession. The humiliation of irrationally committed foundationalists has been a long time coming, so it's good you're trying to get ahead of the curve.


>Why on earth would I ever want to use SQLite if I have access to a sufficiently advanced code generator which can generate me a similarly high quality database system

https://www.youtube.com/watch?v=V_qzqY1bb7I your sufficiently advanced code generator may generate you a high quality database system for some measure of quality, but it will not have SQLite's reliability over the extremely long tail of edge cases proven through its testing and use in real life


Then you've failed the criteria of sufficiently advanced. It's perfectly fine to cast doubt we'll see scaling to this generalization, but you're not casting doubt you're outright rejecting the premise in-confidence. It betrays that you have no idea what you're talking about. May I see your quantification of this long tail? Do you even know how to formalize the mapping from n-bit precision of weights and/or activations to the standard deviation of a transformer's output distribution, such that we could decide whether the long tail of a given behavior is unreachable? Something tells me that no, you don't know how to do that in the slightest. So what drives you to speak with such confidence?

That's before we get into the entire non-linearity of agentic systems introducing massive decidability problems on this in the first place. A little bit of epistemic humility please.


You still spend all your time troubleshooting the reinvented wheels even if AI writes it because you won’t know the edge cases til you hit them. Then you modify the lib, re-release, and update all your apps, but now look where all your time is spent.

The assumption you make is the classic LLM mistake of thinking writing code === building software.

To cite your example SQL has had its tires kicked a lot it’s seen things you can’t even imagine thanks to being used millions of times by millions of people. It’s hard to just replicate all that iteration, learning, mastery, and process. If you reinvent it, users will encounter the dumbest bugs over and over and over. Sure you’ll fix them, but you’re now embarking on this big thing that SQL and others already did.

If you love the problem space definitely do it - go full steam ahead - especially if you’re actually innovating and doing things better, but don’t be fooled into thinking anyone can, or should with every side project.

The new struggle is focus, what not to build, I almost have the purely opposite view of instead of using LLMs for grandiosity, only using the LLM for tedium and making sure it doesn’t do anything too much that I haven’t planned for or want to do. I drive the thing, so every new project is still my time and energy and focus.


> The assumption you make is the classic LLM mistake of thinking writing code === building software.

Not quite. I'd agree with you if you try to do code generation in one-shot, but the agentic loop importantly isn't that. Natural testing at-scale is automated as well.

> but don’t be fooled into thinking anyone can, or should with every side project.

You're right, and I hoped I had conveyed that but I guess I didn't. I'm very bearish on whether or not we're democratizing anything. What I do think is we have a very complicated and interesting tool, and within a reasonable precondition of additional scaling and structural innovation, one that promises a certain kind of person the ability to swing with the weight of an entire institution.

I have been playing around with transformers for 7 years now, and can see and qualify the improvements observationally. Innovations like thinking blocks, the agentic loop, the many architectural refinements flourishing, SAEs, steering vectors, harnesses, etc. Despite what it sounds like, I've yet to incorporate LLMs into any of my professional workflows, they're really not quite up to my standards yet. But then, in my industry I would never in a million years touch anything like SQLite either. Those kinds of things are very taboo for anything but like an internet service and that's more the kind of operations drek that gets shoved off to the IT department. All of this colors my perspective. I was never pulling in dependencies. In the future, it's looking like I'll have infinitely less reason to.


What does the benchmark even mean when people are using AI to make real world things that solve real world problems?

I see people, and my self making amazing things with AI and fixing old projects and having real world impact at the fraction of the cost it would take me to hire people, or hours spent on my own coding.

I have built tools and systems with AI that have allowed me to build windows drivers, android apps, web apps, iOS apps, vm occultation, custom block drivers, custom file systems and more. To the point where entire products have been created.

Not trying to be a doomsday, but yes. It seems as though with the right infrastructure we are at the point where businesses owners can go from idea to product very fast and not need or hire much external talent.


What amazing windows drivers have you sold?


Idk about sold. But it’s loaded on all the windows machines in a fairly big company that solves a real world problem.

It allows us to apply custom ACLs to AI agents and the child process spawned by AI agents. Giving us the ability to control what files an AI agent can read or write to, while still being in the calling users context. It allows us to force all ai derived processes to use a transparent MITM proxy so we can then also apply robust access rules to remote host allow or deny access to specific urls and not others. It also allows us to monitor access to windows Credential Manager with rules ti allow specific singed binaries to access some credentials but not others. It give us complete control of what AI agents on windows can see or not see or access.

Windows native sandboxing is lacking. You have some stuff in WSL that completely are broken once you call a windows native app. Or you have app containers which are too restrictive and result in applying expensive file system ACL to all files the app containers would access, which can take hours when dealing with million of files, and would be required to be applied every time you chains your app container (there are some workarounds, for them but they still have a one time cost a long with a fairly flaky maintenance process). You can get the network part done by running commands as a different user but that would result in the same file system ACL nightmare that app containers has.

Result is we get seatbelt level sandboxing in windows native, and can apply dynamic rules like preventing access to .aws folders regardless of the OS level ACLs, using glob rules like */.aws, so we don’t have to be aware of the exact path ahead of time.

It also has registry tree ACLs and, can prevent process and process trees from gaining administrative access, the list of features goes on and on.


so you decided to vibe code a sandbox because windows sandboxing is lacking, instead of moving to linux where sandboxing is robust?


Some products and companies require windows.

It might please you to know that this driver was vibe coded from a linux vm based sandbox.(that was also vibe coded)


A better question is which products/services have you avoided purchasing.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: