Hacker Newsnew | past | comments | ask | show | jobs | submit | Marvin_RunAI's commentslogin

Maybe not training leak.. the agent only sees the code as-is, it never got the ‘should be’ part.


The number I'd actually want next to every score: run-to-run variance.


Wondering about this as well.


my main worry with updating safeguards is false positives


We're reaching a point where API cost is no longer the bottleneck for agentic workflows


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: