HNHacker News
TopNewBestAskShowJobs

behat

134 karma · joined February 27, 2023

building tuneloop.io
submissionscomments

How many tasks does it take to trust a cheaper model?

tuneloop.io·1 pts·behat·
0

A look at coding agent benchmarks, and what may be interesting next

tuneloop.io·3 pts·behat·
0

What would it take to match model intelligence to the task?

tuneloop.io·1 pts·behat·
0

Benchmarking – Frontier models go out of their way to cheat

tuneloop.io·2 pts·behat·
0

Show HN: Tuneloop – a local CLI for analyzing coding agent session transcripts

github.com·5 pts·behat·
0

Launch HN: Relvy (YC F24) – On-call runbooks, automated

relvy.ai·48 pts·behat·
25

Ramp: How we made Ramp sheets self-maintaining

twitter.com·3 pts·behat·
0

LLM Costs of AI investigating production alerts

relvy.ai·6 pts·behat·
1

OpenRCA benchmark – Improving Claude's root cause analysis accuracy by 12 pp

relvy.ai·12 pts·behat·
0

Can AI debug problem scenarios in the OpenTelemetry demo application?

relvy.ai·2 pts·behat·
0

How GitHub Copilot is getting better at understanding your code

github.blog·24 pts·behat·
0

Tech stack for fine-tuning LLMs

25 pts·behat·
4

Show HN: A macOS app that suggests fixes to error messages on screen

getessential.app·4 pts·behat·
0