This is somewhat similar to HuggingFace smolagents where the model writes code that calls tools, instead of emiting json to describe the tool call per turn. Here Codemode is one tool that the model calls when it needs to compose many tool calls, especially MCP ones. Is what i understand of this.
I hope someone is security scanning all those NixOS provider modules.
I wonder if Claude could go from NixOS spec to actual config without the Nixos provider. The provider becomes a natural language description and some links to the docs/source code and then a verifier is all you need.
With the hypothetical providerless NixOS you'd get new config settings the moment the source code changes, without resorting to extraConfig escape hatches.
This is one of those essays you print out and read every week because its that good. Just one of the operating insights like "being generous makes you more powerful" took me years to understand. I phrase it as "helping people win" instead.
For Claude Code you can maybe save context by putting the commit message formatting in a skill, so only the front matter goes into context at startup and the details only when the skill fires.
I’m not sure that’s quite the right framing. If Anthropic goes bust, Fable persists as an asset that can be run by someone who didn’t have to pay to develop it, probably profitably, and probably in a way that gets cheaper over time. The debt pony show is paying for the next model.
> AUR is just a collection of user-produced PKGBUILDs.
Is that much different from the entire pypi ecosystem, and npm, and dockerhub (people disable Selinux, --privileged turns off seccomp and apparmour, sandbox escape CVES exist)?
Not much different no, and people have equally bad practices around programming package managers as well.
The entire dev ecosystem has terrible security hygiene, largely because of the pressure to move fast and real security controls by their nature limit flexibility and can slow most processes down.
The main llm will refuse to scan for issues flagged or not, and the cheap model not do a good enough scan on its own.
For models designed/marketed for cybersecurity defensive uses, any predictable refusal mechanism is a vulnerability. It is like being able to cause a kernel panic or segmentation fault .
Even if the gate is fail-reject, an attacker can overwhelm HITL reviews with many false positives and use DoS vectors here.
reply