Hacker Newsnew | past | comments | ask | show | jobs | submit | carlsborg's commentslogin

This is somewhat similar to HuggingFace smolagents where the model writes code that calls tools, instead of emiting json to describe the tool call per turn. Here Codemode is one tool that the model calls when it needs to compose many tool calls, especially MCP ones. Is what i understand of this.

I hope someone is security scanning all those NixOS provider modules.

I wonder if Claude could go from NixOS spec to actual config without the Nixos provider. The provider becomes a natural language description and some links to the docs/source code and then a verifier is all you need.


How is that better than the nixos provider? The trouble is the links to the source code right?

With the hypothetical providerless NixOS you'd get new config settings the moment the source code changes, without resorting to extraConfig escape hatches.

This is one of those essays you print out and read every week because its that good. Just one of the operating insights like "being generous makes you more powerful" took me years to understand. I phrase it as "helping people win" instead.


> "helping people win"

That's a great way to put it!


For Claude Code you can maybe save context by putting the commit message formatting in a skill, so only the front matter goes into context at startup and the details only when the skill fires.


They removed ~80% of Claude Code’s system prompt for Claude 5 models so the harness cost-per-task benchmark could be stale.


Make the most of your heavily subsidised $20 / $200 subscriptions while the credit spreads allow it.


I’m not sure that’s quite the right framing. If Anthropic goes bust, Fable persists as an asset that can be run by someone who didn’t have to pay to develop it, probably profitably, and probably in a way that gets cheaper over time. The debt pony show is paying for the next model.


Good for humanity, and also GDDR/HBM manufacturers.



> AUR is just a collection of user-produced PKGBUILDs.

Is that much different from the entire pypi ecosystem, and npm, and dockerhub (people disable Selinux, --privileged turns off seccomp and apparmour, sandbox escape CVES exist)?


Not much different no, and people have equally bad practices around programming package managers as well.

The entire dev ecosystem has terrible security hygiene, largely because of the pressure to move fast and real security controls by their nature limit flexibility and can slow most processes down.


Pipeline is then: Cheap open source model for flagging potential LLM refusal content -> main LLM check


How will flagging help?

The main llm will refuse to scan for issues flagged or not, and the cheap model not do a good enough scan on its own.

For models designed/marketed for cybersecurity defensive uses, any predictable refusal mechanism is a vulnerability. It is like being able to cause a kernel panic or segmentation fault .

Even if the gate is fail-reject, an attacker can overwhelm HITL reviews with many false positives and use DoS vectors here.


Cheap model replaces trigger words with something innoculous. Of course, this breaks dynamic analysis if malware has unpatched integrity checks


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: