You absolutely can, that’s what your harness is for. You don’t need your environment to “reason” about things when deterministic tools exist - You have a really fancy hammer, but that doesn’t make everything a nail.
To offer a possible example: What would the game Zork™ look like with an LLM? Assume we do not want to let players sweet-talk the system into letting them teleport to the end.
The LLM's job would be to channel "I perambulate in the direction of the Arctic circle" into go(north). You saved writing the grammar parser, but you still need to write the game world.
But what if we used the fancy hammer to change the shape of everything to be a nail? And what if we build the handle of the fancy hammer with a fancy hammer? With all of this we could build a very good fancy hammer manufacturing company.
Yet the hammer seller continues to scream everything is a nail and their hammer will replace your entire job eventually. So are you telling me the hammer seller is lying or am I the one using it wrong?
I don’t understand why Mozilla keeps doing these odd projects. They maintain the only real alternative browser stack to Chrome and rather than keep it up to date and triage bugs, so many resources keep flowing into random AI hype crapware that no one asked for.
Well they need a new strategy for finding that. Continually creating half-baked side projects then dropping them after a few years clearly, and unsurprisingly, isn't working.
You may be right, but I'm not sure we actually have strong evidence for that. Mozilla has run like a bigtech firm for so long that it's not really clear whether the open-source portion thereof could survive without.
$500 million in annual revenue and $6 million CEO compensation packages are well in excess of what other similar-scope projects get by on. Mozilla accounts for somewhere between 30-50% of all funding in open-source - if you sum up the Linux foundation, the Wikimedia foundation, the Apache foundation, KDE e.V, the Gnome foundation... it's still less than Mozilla
Honestly, good luck with that. Mozilla spent past 15 years or so creating and killing new projects over and over. Right now I'd be hard-pressed to try anything they create because they would inevitably shut it down next year.
I would like to see this (and any model router, frankly) benchmarked against two things, personally:
1) Claude Code's "advisor mode" (nominally, Sonnet 5.5 with a Fable advisor)
2) Copilot's "HydraFusion" model router/advisor combo.
Specifically I would like to see them compared on architecture planning (both human assisted and hands-off with a draft document) and code review, as these are what I have found the most significant improvement on with multi-model systems.
> our initial hypothesis that an ensemble of models can do better than any single model ever could.
To your hypothesis, anecdotally I find both of these offerings to be far superior to any single model for most tasks of any real complexity, and both to have general frustration/failure cases that single models do not. I would not be surprised that any ensemble approach that utilizes more than one single model meets this hypothesis.
Quite frequently I delegate review and restructuring loops to subagents acting as judges/advisors to tell the primary agent if it met the goal it was instructed to. For some workloads, I will even vet every tool call and user-facing output this way.
I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.
> Also I’m very interested in the unique failure cases you’re referring to! What have you noticed?
Permissions issues would be the most common - models collaborating with eachother on a task seem to try to convince eachother they have either more or less permission to do things than they actually do. Especially when transitioning between creating a plan and executing the plan.
Another is deciding that there is a limit to the “loops” they are allowed to run to iterate on something. In many cases I have set an explicit goal, and come back to an agent stopped and reporting that it has hit the “maximum allowable loops of [insert arbitrary number that changes every time].”
Now that I think about it further, I believe the other examples I have also all fall into the models hallucinating the presence of control instructions, or attempting repeatedly to violate permission boundaries that a single model’s harness instructions would usually guide it away from re-attempting.
> I’m not really sure, I can’t say I trust any kind of synthetic AI benchmarks.
Fair enough but what exactly were you thinking when you said:
> I would like to see this (and any model router, frankly) benchmarked against two things
And thanks for sharing about failure cases! Those do sound like strange harness-level things, honestly we haven't seen failure cases like that crop up in our own usage & testing.
I genuinely don’t understand the point of this product. Anything I could think of for an “always on agent” to do, AI is not yet good enough to do without being supervised.
I think you're operating on a different definition of "work".
The OP is discussing tasks to do, you sound like you are discussing things that need to be done. The OP isn't saying there aren't things to be done, they are saying that no one is going to tell them what needs to be done, so they must discover and formalize it.
I would heavily disagree with this - it depends entirely on the cancer you get. If it's something slow growing that would take a decade plus to kill you, sure - just trying to treat it would be worse than letting it ride.
But my father died of cancer nearing age 80, (most likely mesothelioma, though we didn't do heavy tests as it was discovered late), and it was incredibly painful and miserable for him. He had fast growing tumors in his ribs, forcing bone to spike out and hit his lungs, the tumors that grew in his lungs caused constant pneumonia and kept leaving him unable to get enough oxygen to his brain to communicate for hours at a time.
I feel your pain, I really do, but I think you missed the point. In almost any metric, if you have to die, it is better to die after having lived a full life than it is to die young.
I used to believe this without a shadow of a doubt, but without the choice to go out on your own accord, I don't know if I do any more after what I saw.
I think the original point was that cancer 80+ is better than cancer at 45, so you should still try to avoid habits that make it more likely earlier, even if the chances remain 100% at sufficient time scales
I don't understand what you think this conversation is about. It's not "die quick and painless at 45 or excruciatingly painfully at 80" it's just 45 or 80. Nobody said anything about excruciating pain except you.
Do you think your dad would prefer to die exactly the same way at 45 instead? No? Great, conversation over.
What an absolutely obnoxious way to respond to someone who was willing to share his experience about a subject as personal and painful as losing his father to cancer.
I don't think it's obnoxious at all to point out that forgoing nearly half your lifespan in order to avoid a painful death implies that you don't value your life much
Maybe it’s that I have young kids, but even knowing my death at 80 would be horrible, I’d choose that over a painless death at 45 in a heartbeat. I’d beg and pray for that outcome.
Unfortunately a lot of us will die in comparable ways if we are lucky enough to get that old. My grandmother was healthy & mentally with it until stroking out in her garden one day at the age of 90. That is the best any of us can hope for and the reality will almost certainly be far more painful and prolonged.
It depends on mechanism, but many cancer patients do deal with a great deal of pain, especially near the end of the disease process. I am sorry for your experience. It is not uncommon.
It is unfortunately a complicated situation that often doesn't have a perfect solution. Without trying to cover every possibl nuance, I'll just say that many end-stage patients have pain so refractory that sufficient doses of pain medication are effectively lethal. This is usually offered, but obviously not everyone accepts.
The truly sad cases are those in which the patient no longer has decision-making capacity and it falls on a family member to make such decisions. When you aren't the one experiencing the pain, it's much harder to weigh that decision rationally. This is why so many providers push patients to establish legal documentation of their end-of-life wishes.
Or hear me out here… we as a species have developed measuring methodology that is precise enough at human scale but not actually correct and breaks down at larger scales.
reply