Yeah, I don't understand why this is hard. If I don't want to take responsibility for an accidental discharge or someone getting my firearm, I won't use a firearm.
You set a goal. Agent will do goal.
The rest doesn't matter: the instructions, the "guardrails" etc.
The agents are not smart, they don't reason, they don't think, there are no morals, no ethics. Nothing will prevent not doing the goal because that is the set goal. It's a statistical model that will "justify" anything to do X.
I'm finding it mind bogging how this is not clear for everyone.
Everything so far is pointing to the conclusion that you can only set ONE goal. Exactly one.
But let's assume not. If you want things like "do not break out of sandbox" - have you defined what the sandbox is? Eg. "never, ever leave the IP range 10.0.0.0/8" would be a bit more precise, but technically using a proxy bypasses that limitation as the system itself never left 10.0.0.0/8.
See, it's a tad bit hard to define the rules properly.
Which is why Wish, the spell, should really be avoided in D&D. It's the same problem: it's up to creative interpretation.
It's already a "regulation" that one shouldn't steal, enter a private property without the right to do so, etc, yet we have a lot of these crimes.
I could see the UK trying to set a law on ehat models are allowed, only to learn again, that the UK law doesn't apply everywhere, but as long as you can rent compute in a foreign country the whole idea is dead.
The part I don't understand is how the excess traffic not triggered alarms on the hugging face side, or the url shorteners used, or on anything that was touched in the process.
Nothing got overloaded, no unexpected CPU or IO use? Did it blend into the normal traffic somehow?
I have a client that uses webflow and they author everything in the edit mode and the site is constantly breaking. They just email me to update everything now.
reply