Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
A look at coding agent benchmarks, and what may be interesting next (tuneloop.io)
3 points by behat 21 days ago | past
What would it take to match model intelligence to the task? (tuneloop.io)
1 point by behat 23 days ago | past
Benchmarking – Frontier models go out of their way to cheat (tuneloop.io)
2 points by behat 27 days ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: