HNHacker News
TopNewBestAskShowJobs

jascha_eng

982 karma · joined January 23, 2024

Software Engineer at Timescale

Socials:

- github.com/Askir

- linkedin.com/in/jaschabeste

Opinions are my own

submissionscomments
jascha_eng··on Livenerf: Has Opus 5.5 been nerfed yet?
Yeh it's absurd that people claim this all the time. It's some crazy conspiracy theory and when you ask for examples nothing ever shows up.

It would be economical suicide from anthropic and OpenAI to actually need models intentionally.

But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.

jascha_eng··on Drawgent: Coding agent on a live Excalidraw canvas
I played around a bit with pencil.devs mcp integration before: https://www.pen.dev/

Which was at least fun for a bit although I failed to really see the use after models became better at designing themselves.

jascha_eng··on Drawgent: Coding agent on a live Excalidraw canvas
MCP is the answer but it's only really worth it if you product is reasonably complex e.g. n8n has a community MCP fully for this reason.
jascha_eng··on Grok 4.7
32 on the omniscience index. Not terrible but far from Astra and fable: https://artificialanalysis.ai/evaluations/omniscience
jascha_eng··on I am often wrong
That's what it took to finally support Agents.md I guess. Boris had to blog about having made a mistake.

How about you stop taking things so personal instead and let others steer more.

jascha_eng··on Ask HN: What are you working on? (September 2026)
Self hosted database access management, think Code Reviews for SQL Statements for when your devs need to go to prod: https://github.com/kviklet/kviklet

Recently cleaned up my postgres proxy and built a MySQL wire proxy so that you can use psql/datagrip as a dev but every statement is still logged. With SSO and no password sharing ofc!

jascha_eng··on iPhone Duo
But it makes it really weird to type one handed I think.
jascha_eng··on Artificial Analysis Intelligence Index v4.2
But it is actually a great model it e.g. got the carwash question right from 9 months ago. While openais models all struggled.
jascha_eng··on Artificial Analysis Intelligence Index v4.2
Fair but I think those two behaviours are strongly correlated at least the index does represent my personal experience very well where fable is way better than opus opus is better than sol. And I haven't tried Astra yet but it having 44/43 is very interesting at least.
jascha_eng··on Artificial Analysis Intelligence Index v4.2
I dont think you read my message. Muse and sol are nowhere near fable and Astra on the omniscience index
jascha_eng··on Artificial Analysis Intelligence Index v4.2
Imo the omniscience index they have has the highest correlation to actual usefulness of the models.

https://artificialanalysis.ai/evaluations/omniscience

> measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.

This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. But this index measures how often it is correct while penalizing wrong responses so that a high score means you can trust this model more and when it doesn't know it is more likely to tell you that it really doesn't know rather than making shit up.

Fable also performs a lot better than opus 5 here which correlates very strongly with perceived strength despite the models performing similarly on e.g. DeepSWE

Astra is a big jump from sol and performs the same or slightly better than fable here.

jascha_eng··on What Is a Harness?
The ai hype word for 2026 after agent in 2025 for any LLM powered application.

Well kind of, I wouldn't be surprised to see that some things marketed as agents are actually good old deterministic software.

jascha_eng··on The August 17 outage
Instagram, Facebook and even threads all had much more mundane growth rates and definitely no unexpected jumps like GitHub is experiencing. I'm sure if suddenly the solar system had 10 more earths with each about 10 billion people and they would all start using Instagram tomorrow we would have exactly the same growing pains and outages that GitHub has today.

Luckily for Meta agents are not yet as much into doomscrolling as humans are.

jascha_eng··on Humans missed 1 in 3 threats approving AI agent commands across 40k game runs
1 in 3 is not terrible you just need a few more humans in the loop to reduce the error rate meaningfully. Combined with other classifier models and heuristics you can get good results. Humans can probably also perform better if they don't have to judge every single command but just suspicious ones our attention is limited after all.
jascha_eng··on How to Spot AI Writing
But fables is often worth diving into at least. Opus 5 says this and then rambles on about something completely irrelevant or even wrong
jascha_eng··on Writing by hand is good for your brain
The study he cites is also specifically using digital pens

> Brain electrical activity was recorded in 36 university students as they were handwriting visually presented words using a digital pen and typewriting the words on a keyboard

kinda funny.

It's also not clear from this study that there is actually any benefit to "more brain activity". Of course doing more complex motor tasks requires more brain activity but nobody guarantees that this helps with remembering things.

jascha_eng··on OverpAId – Fire your CEO. Hire the future
tbh AI is great at vague strategic decision making with a low risk bias. I wouldn't mind working for Claude
jascha_eng··on Kimi Work
huh? because im curious what they used? Claude Code and codex take completely different approaches. The core loop of feeding generation and having a bash tool is entirely the same. If you want to start building your own thing I'm sure there is good setups out there to start from rather than going from zero.

Ofc I can just fork gemini-cli if I want my own version but like I don't think that's what OP meant by build your own thing and customize it.

jascha_eng··on Kimi Work
How did you build your own coding agent? What language/framework did you use?
jascha_eng··on Typing Speed Test, but for Developers
It used to be that DevOps was the movement of merging Ops leftwards so that Devs own the Operation of their software.
jascha_eng··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
And that doesn't work with a simple prompt?
jascha_eng··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
Can you give an example? And more curious about what you do with the resulting code afterwards I imagine its gonna be a big chunk then?
jascha_eng··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
Is this useful? I feel like the problem is usually not that the model isn't capable of achieving what I give it, but the way it does it. Especially if originally I didn't 100% know how I would do it myself the model often takes weird paths through the code base, takes shortcuts that end up in weird feature interactions or pulls in a dependency without weighting if it could've been done without that.

I haven't really found a good way to solve this other than:

1. Produce an initial PR fulfilling all the requirements I knew at the start

2. Chat with the model about any weird snippets I notice and talk through alternatives

3. Simplify anything that I think is overengineered or plain unncessary

Sometimes I restart all over with more precise requirements but then it sometimes makes different mistakes/takes different shortcuts.

In practice the earlier I review the better the end result imo, so /goal seems very unproductive to me?

jascha_eng··on Claude Code: Anatomy of a Misfeature
Until the classifier is wrong or also prompt injected. the classifier is just as vulnerable as the model itself is. Yes it is harder to break but trying to make a nondeterministic tool deterministic by adding another nondeterministic one on top just reduces the chance of something going wrong.

Tbf as long as that chance is low enough it doesn't matter in practice, but I have definitely seen the classifier approve things that were questionable, and I've also seen it decline things that were obviously okay.

jascha_eng··on OnePlus halts operations in USA and Europe
But a pixel is quite a bit more expensive no? At that point you can consider an iPhone?
jascha_eng··on OnePlus halts operations in USA and Europe
Sad I have a 6 year old oneplus and was looking for a new phone somewhat soon, would've considered them again for sure. Any alternatives? They always had a reputation for me for being a great no fuss, little bloat and simply fast android phone.
jascha_eng··on The real prices of frontier models
I mean it might lead to better performance on the model side. So the tokenizer is better but more expensive.
jascha_eng··on The real prices of frontier models
Aside of the claudeisms and the obvious AI smell, it overexplains everything and doesn't come to any useful conclusions. It's just not a good post.

The nudge to think about both "tokenization as variable" as well as actual tokens consumed per task is still good.

jascha_eng··on Ask HN: What Are You Working On? (July 2026)
Slowly improving the UX on my SQL review/approval tool: https://github.com/kviklet/kviklet

Also finally closed the first real customer on it recently!

I want to get through a large chunk of the open issues the next few weeks and then spend some time building agentic capabilities for it. I believe a central place to configure database access for your dev team without having to share passwords and with sensible review policies should also help e.g. if claude needs to access production data to validate a premise.

Still have to figure out the right UX though not sure the agent should have the exact same review requirements that a human does. Maybe it needs to be configurable separately

jascha_eng··on GPT-5.6 Sol Ultra will be in Codex
im talking about anthropics pricing
Page 1 of 8Next →