Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
behat's submissions
login
1.
A look at coding agent benchmarks, and what may be interesting next
(
tuneloop.io
)
3 points
by
behat
11 days ago
|
past
|
discuss
2.
What would it take to match model intelligence to the task?
(
tuneloop.io
)
1 point
by
behat
13 days ago
|
past
|
discuss
3.
Benchmarking – Frontier models go out of their way to cheat
(
tuneloop.io
)
2 points
by
behat
17 days ago
|
past
4.
Show HN: Tuneloop – a local CLI for analyzing coding agent session transcripts
(
github.com/tuneloop
)
5 points
by
behat
46 days ago
|
past
5.
Launch HN: Relvy (YC F24) – On-call runbooks, automated
(
relvy.ai
)
48 points
by
behat
5 months ago
|
past
|
25 comments
6.
Ramp: How we made Ramp sheets self-maintaining
(
twitter.com/ramplabs
)
3 points
by
behat
5 months ago
|
past
7.
LLM Costs of AI investigating production alerts
(
relvy.ai
)
6 points
by
behat
6 months ago
|
past
|
1 comment
8.
OpenRCA benchmark – Improving Claude's root cause analysis accuracy by 12 pp
(
relvy.ai
)
12 points
by
behat
6 months ago
|
past
9.
Can AI debug problem scenarios in the OpenTelemetry demo application?
(
relvy.ai
)
2 points
by
behat
on April 30, 2025
|
past
10.
How GitHub Copilot is getting better at understanding your code
(
github.blog
)
24 points
by
behat
on May 23, 2023
|
past
11.
Tech stack for fine-tuning LLMs
25 points
by
behat
on May 17, 2023
|
past
|
4 comments
12.
Show HN: A macOS app that suggests fixes to error messages on screen
(
getessential.app
)
4 points
by
behat
on May 7, 2023
|
past
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: