Harnesses are the next frontier. If LLMs are electricity, harnesses are the “electronics.” Right now, it’s like an AC vs DC between Claude and ChatGPT, but once that settles, the harnesses will be the actual value providers.
And Pi is the best harness because of the amazing extension system. You can build extensions that turn Pi into a stock trader, software factory, anything. I tried switching to another harness but none have extension functionality as good as Pi.
Even if there is a new harness or agent project, I tell Pi to dig into the codebase and then make me an extension that brings that functionality into Pi. I did it with Prime Intellect’s and Deepseek’s harnesses and those are built on Pi.
Depends on your product strategy. If you only care about how your model will be used in the context of a harness (perhaps, specifically the harness that you designed), then the incentive is plainly there to optimize the weights within the context of the harness.
What is being absorbed into weights is tool usage. It is incredibly counterproductive to train models on a specific harness, when instead it can be trained to reason about the tools that it has available to itself and how to best use those tools to accomplish it's goal.
Would you rather hire an engineer that can adopt to your org's prefered tooling, or hire an engineer that can only perform well with their own favorite tools? It's the same thing.
I was saying more the custom skills and extensions that make the harness not a commodity. Yes people will use Claude Code, Codex, or Pi but their customizations will make their harness unique and more powerful.
In a sense they are the last frontier imo. At some point a harness will be built that can modify itself to fit the needs of the majority of people's workflows and evolve with them.
Then people will want to share and exchange their evolved harnesses. Ways will be found to modularize certain aspects to enable mixing and matching.
I’m thinking of how in cyberpunk, people are replacing their cybernetic enhancements all the time. You could alternatively bioengineer your own body towards the desired outcomes, but that’s more constrained by the trajectory your body has already taken, whereas the promise of cybernetic parts is that they are more independently replaceable. (Probably an illusion in practice, but I’m talking about the fictional ideal.)
As another analogy, monolithic software tends to quickly become hard to change significantly, whereas a plugin architecture tends to be more flexible and modular, and people can share and combine their various plugins.
Sadly, many people have bought into the cult that LLMs will lead to AGI. I guess if that is your worldview then all this babbling about new frontiers makes more sense.
They probably used an LLM to come up with this bizarre metaphor.
I find it difficult to understand people who are wildly skeptical about LLMs leading to AGI (assuming we can even agree on what that means). Consider:
- They can already reason better than many humans and are still improving all the time
- Harnesses are improving all the time
- We're already exploring things like long term memory, long term goals, and other things that humans have which LLMs traditionally lack
- An AI agent can read and reason about every piece of AI research ever published, including looking for insights that humans may have missed. A team of humans could never do this even if they dedicated their whole lives to it.
- They can design and execute experiments on a mass scale to determine what does and doesn't work
- Large AI labs have more than sufficient resources and motivation to throw at the problem, and are in fact doing this.
So you believe LLMs (despite their inherent deficiencies vs EBMs [0], etc.) can lead to what you'd consider AGI, but you also admit that there is no agreement what AGI actually would be and you further don't provide your own definition? But you are surprised that some (like e.g. Yann LeCun (I am very convinced by his published works beyond his authority in the field, but am willing to admit his could be seen as a biased position)) are skeptical?
If you provide what you'd consider AGI, we may not agree on that definition, but I and other skeptics could at least discuss with you whether A.) that seems reasonably achievable given LLMs inherent limitations and B.) whether any of what you'd listed is actually likely to get us there.
As it stands, neither is possible without knowing what you believe AGI to be, but for what it's worth, coming from someone who both does see LLMs as valuable tools but whose definition for AGI also contains, among other things, reliable self-assessment of factual uncertainty [1] and basic counting and grade school maths [0][2] without tools or eternally scaling training data, I have yet to read any evidence that LLMs can achieve my, rather strict, metric for AGI.
These models are amazing tools, their ability to leverage massive amounts of high quality training data to further sciences truly awe inspiring, but that does not mean intelligence, at least in my definition that requires some internals these models have never been proven to possess. It's nuts that solving Erdos problems can be done by a model which struggles to count or solve a sudoku without external tools, but that's where the technology has been for years now and no paper I have read has shown that LLMs can overcome that to any scalable degree. You can push further with training data, but the limitations remain, albeit less noticeable. Any externalities, be it tools (self-scripted or called by the model), external memory solutions of all shapes and sizes, etc. I personally also feel cannot be required for or lead towards AGI as intelligence may be better leveraged by such externalities, but should never require them, so much of your suggestion I feel shouldn't be considered even if one believes LLMs can yield intelligence. I will admit that I am very extreme here though, this is not a position held by everyone for good reason. At the end I will always point towards the "extraordinary claims require extraordinary proof" of it all and that LLMs, in the face of any doubt, should be viewed akin to how Stockfish can play better than any grandmaster, but that does not mean intelligence, at least in my world.
If your definition for intelligence does not require basic arithmetics or an understanding of ones own knowledge gaps, then maybe LLMs can achieve that, but I'd push back on that truly rising to the AGI moniker. Maybe a more comprehensive or even my definition of AGI is possible whilst keeping the autoregressive nature after all, but there is no evidence supporting that by itself and quite a few things that haven't even begun to be overcome before something of that magnitude could be honestly considered.
It's akin to "let's colonise Mars by 2020 or 2030 or 2040 for sure, then terraform it" proposals. If that were possible, wouldn't we see a lot of these methods applied on earth and in a moon base long before (as in, we'd have had a permanent moon base in the early 2000s)? Same with LLMs, if they can truly yield AGI, we'd see some of the major deficiencies dealt with long before. The fact that we neither are terraforming earth, nor have any permanent off world colonies, nor have solved some of the listed, inherent limitations with LLMs by their design, that's what informs my skepticism that both are reasonably achievable in the timelines some industry "experts" (read hype merchants) propose on the regular. You tend to see some progress, a path toward solving actionable problems long before full implementation, at least in the real world...
Did I even mention AGI? All I’m saying is that we’re hitting a plateau with how good models are while harnesses are untapped potential. And with Pi, you can swap models like electricity companies. Yes for now, the electricity is better with some companies but this will stabilize.
And no I came up with the metaphor all on my own, send me the chat of you getting the LLM to come up with it. Why not argue based on merit instead of strawman and ad hominem attacks?
I’m afraid it’s a terrible metaphor, starting with the fact that LLMs are nothing like electricity, and the relation of harnesses to them is nothing like that of electronics to electricity, save perhaps one is a prerequisite of the other.
Harnesses (and the concept of agents before them) presuppose competence in LLMs which simply doesn’t exist.
“Just as electricity transformed almost everything 100 years ago, today I actually have a hard time thinking of an industry that I don’t think AI will transform in the next several years” - Andrew Ng
I didn’t come with the electricity idea, it was Sam Altman saying it will be like a utility down the line and metered[0]. What would the “electronics” be in your opinion?
If it is metered and like a utility, Sam Altman will not benefit alone. All the models will have plateaued and you can swap for any of them. Then the only differentiator is the harness.
I don’t believe in AGI, but that doesn’t mean I don’t find AI useful. I just understand that the correct harness can take them to the next level.
Guided by humans, code generators which have ingested the worlds’ code and can recognise and generate patterns can be useful tools. I wouldn’t personally qualify it as a wild success as we are early and there are significant downsides.
That doesn’t make them intelligent agents which think independently.
Would you like to see some of my own examples of wild success?
I have a system that entirely reverse engineers old arcade games. Creates semantic symbol mappings that were considered impossible just a couple years ago.
Granted, it took me a couple weeks to build the system.
From impossible to a couple weeks in just a couple years.
Would you like to see it or continue to pretend these things don't exist? Your call.
(It's finding the coolest stuff - the anti-tampering hacks they put into the old machines is fascinating.)
You were making some pretty strong statements about the usefulness of harnesses that seemed to me to be entirely detached from reality. Maybe I misunderstood you? Because I can provide evidence to correct that misconception.
Another of my projects is to incorporate Pixar's ideas from RenderMan into a 3d printer slicer. Displacement shaders, in a 3d printer, have never been done before. Would you like to see that? I had the idea 10 years ago but it was too tedious to implement. I have a working system now in just a couple weeks AGAIN.
Anyone claiming agents aren't profoundly useful is WRONG. If you disagree, please let's discuss it.
In my experience the use of agents and harnesses and other scaffolding around them doesn’t improve the performance of LLMs much, which is adequate for some tasks under supervision but nothing like general intelligence (or electricity for that matter). But good that it works for you.
OH - I think I understand now. You're like LeCun - dismissive of LLMs in all their forms. So of course agents can't lead to AGI, because LLMs just can't.
Which I've always found so strange. Language was always considered the pinnacle of the human mind - right up until we created LLMs. Now it's the world model - the things that animals always had and the thing we used to look down on.
From my perspective, language is still at the top. I'm not a fickle lover of the gaps.
> In my experience the use of agents and harnesses and other scaffolding around them doesn’t improve the performance of LLMs much
I have to say - that seems straight-up crazy to me. Would you mind digging into that discussion?
As just one example - in a harness I can ask an agent to confirm everything it says via a second sub agent - which dramatically improves the output. It almost entirely solves the problem of hallucinations. You don't see that as an improvement?
> Sadly, many people have bought into the cult that LLMs will lead to AGI
You can never tell if the goomba opinion of the forum will agree we have reached AGI (seen that happen on a few threads lately) or will readily call that a ludicrous proposition.
I've never used Pi but I don't see why you can't use stock codex or claude code for the same purpose, what makes Pi special? I've built plenty of custom harnesses on top of claude code and codex using custom skills or simple markdown instructions and subagents. Never had any issues or limitations with that approach.
I do agree that harnesses are going to extend AI capabilities a lot in the next year, but after reading Pi's page I don't see anything that makes it particularly special in terms of functionality, other than being more provider-agnostic.
For one you can ask Pi to create a TUI extension, so along with the agent interface you can add whatever custom TUI you need, such as portfolio stock tickers, alerts, whatever you want.
Many of my harnesses eventually turn into customized UIs around the chat interface.
I was doing something similar months ago with openclaw. I had skills/scripts that replaced my todo list, expense tracker, habits, whatever etc and then would create a minimal web ui. Then a deploy skill that wires it up to my docker/traefik setup. This eventually led to a custom chat dashboard with those wired up as widgets.
Author here. I think our website could be much clearer - but Pi is fundamentally easier to mold than other harnesses. It’s not magic but it strikes the balance well of letting you shape it extensively without letting you break it.
A harness is the bottom layer of a pie that gets fed into the model. In my project, I count 7 more layers on top of it https://replicated.live/blog/wiki
They all affect consistency, coherence, token efficiency. Probably we need some broader term. Like "information architecture", "knowledge architecture"? It's not just shoveling Markdown to nvidias, after all.
Codex and Claude historically had more bloat in their system prompt and tools. Pi is minimal by design so more adaptable. But to be fair Claude Code is moving in the Pi direction with a small system prompt.
I've been getting a little frustrated with having to rearchitect things any time I want to try a new harness. Wrote about my most recent experiments with separating conversation from control loop here and using MCP as the seam here: https://demianbrecht.com/posts/the-harness-within-the-harnes.... This allows me to build a spectrum of agentic to entirely deterministic tools and be able to port them from one harness to another with only a minimal amount of harness-specific config.
This is a plug, but relevant. I recently added a 'build native tools on the fly' functionality to Dirac (https://github.com/dirac-run/dirac) that works like:
1. You can use the '/new-tool' and tell what kind of tool you want (including whether it should be task-scoped, workspace-scoped, or global), the model builds it, the harness runs validation and other tests until the tool is ready
2. The model decides that in such and such task, it would be helpful to have a tool like this, it can build a task-scoped tool.
In either scenario, the tool catalog is rebuilt, and the new tool is instantly available in the next turn.
This type of modification of the harness on the fly to fit the need is the future.
The only thing left after that is the mobile front. I think static app store type software as we know it is a thing of the past. You'll only ever need one self modifying app.
What I can see is a world where we end up with a Chromium-shaped harness, a fully featured standard implementation everyone builds against, because doing every single thing yourself would be crazy.
I disagree with this. Unlike training models (which requires huge compute), harness development is available to anyone with an editor and ideas. That means that solo devs and small startups can still make meaningful progress.
Also, having only a "standard implementation" makes no sense for a harness. A standard implementation would need to try to be as good as possible at all things. But you'd often want a specialised harness designed for exactly your use case.
It is way too early to tell. If we are comparing to web browsers we are in the early 90s with Netscape, IE, Firefox, etc. I don't even think we are at the point where agent's have a metaphorical JavaScript, we are that early.
The thing that makes everyone build against Chromium is because web browsers are very hard and it is well supported by dev tools like Electron and Playwright.
Harnesses are so easy compared to a web browser, I'm curious what in this world you see that would make building your own harness seem crazy, because I don't see it.
I built a software factory and am now building a stock trader using opencandle extension[0] and a custom extension. For inspiration for how to tweak Pi, check out OMP, Prime Intellect, and Deepseek harnesses.
looking at the website. i can't really tell if they have benchmarks and measuremnts on how all that improves capablities over just using regular agent withtout all that
Pi doesn’t have a UI like Claude Desktop. It also doesn’t work with the Claude subscription, only API key and pricing.
So if you do want to use it, use the Codex sub. Once you install it, run Pi and /login and you’ll get login with ChatGPT. From there, Pi can tweak it’s settings if you ask. Check out their extensions (or ask Pi) and that will take you most of the way there.
Yes but Pi has had a minimal system prompt since inception. Skills and Pi extensions let you make a hyper specific harness for specific use cases. For general conversation, harnesses are overkill most times.
There’s evidence of harnesses making a smaller, weaker model perform better than SOTA and some benchmarks ban harnesses because it becomes too easy.
And Pi is the best harness because of the amazing extension system. You can build extensions that turn Pi into a stock trader, software factory, anything. I tried switching to another harness but none have extension functionality as good as Pi.
Even if there is a new harness or agent project, I tell Pi to dig into the codebase and then make me an extension that brings that functionality into Pi. I did it with Prime Intellect’s and Deepseek’s harnesses and those are built on Pi.