HNHacker News
TopNewBestAskShowJobs

jakozaur

5,792 karma · joined January 3, 2009

Email: jacek (at) migdal.pl
submissionscomments
jakozaur··on September 2026: The world today, as seen by one Polish guy
I'm Polish, but this is such a pessimistic, overdramatic take. The visualisations are impressive, but you're living in Poland's golden age. See this Hacker News story with 1,000+ upvotes: https://news.ycombinator.com/item?id=48062117

The war in Ukraine is a tragedy, but it has proved that Russia is weak. The price shock will speed up the transition to renewables, and Iran is losing its grip on Hormuz. So many of the other issues are solvable.

It's easy to be a pessimist, but there's a bright future ahead of us.

jakozaur··on People Training OpenAI's AI Fired for Using AI to Train the AI
Typically, for AI training, you need human feedback. If AI were enough, then OpenAI would do it themselves. Though we all know how bad AI slop is, running it in a loop can lead to unreadable sentences.

They hired contractors on the condition that they provide human feedback without AI; those people broke the rules, so their contracts ended prematurely.

Not sure why this is news uncovered by an investigative journalist.

jakozaur··on Claude Status – Elevated errors for multiple models
I don't believe we are anywhere close to AGI until I see 99.95% uptime at the frontier AI lab. Now they are undercounting. I hit 50x even when they claim they are up.

We now have alien intelligence that is very different from human intelligence.

jakozaur··on Hackathons in the vibe coding era: our token economics experiment
I would say so. In this exercise, more tokens meant better results. Though DeepSeek tokens were as good as frontier ones.

In this hackathon, how much data you analyzed was correlated with tokens spent.

Analyzing the whole dataset gave the best insights, but cost more tokens. Using a small dataset could save tokens, but some insights were premature.

jakozaur··on OpenJev
Yeah, real Jev got really weird, no benchmarking clause. Their Terms of Use (1(v)) and MCA (2.3(f)) both prohibit users from publishing "benchmarks or performance information about the Services". No major AI has it; we are back to Oracle-style legal.

Though Jev is original, it looks highly replicable.

jakozaur··on RTK reports token savings, but our cost benchmarks disagree
I believe RTK would work well in that use case.

Sometimes creating less verbose variants yourself (a simple script, build.sh, with pointers to logs) can be a quick win.

jakozaur··on PISA 2025 Students' reading and mathematics performance declined across the OECD
The drop is not uniform; see Table I.2.6 (2022 -> 2025). Focusing on math:

Stable, above OECD average: Taiwan, Macao, Korea, Estonia, United Kingdom, Poland, Lithuania, Italy, New Zealand

Stable, around or below OECD average: United States, Slovakia, Greece, Croatia, Qatar, Montenegro, Mongolia, Mexico, Brunei

Dropping, previously above OECD average: Singapore, Japan, Hong Kong, Switzerland, Canada, Netherlands, Ireland, Australia, Belgium, Austria, Denmark, Finland, Germany, Sweden, Slovenia, Latvia, Portugal, Norway, Spain, France, Hungary

Dropping, around or below OECD average: Iceland, Israel, Malta, Serbia, Chile, Bulgaria, Kazakhstan, Argentina, Malaysia, Romania, Peru, Ukraine, North Macedonia

Improving, all below OECD average: Türkiye, Thailand, UAE, Philippines, Jordan, Georgia, Cambodia

jakozaur··on Understanding ChatGPT Work
Somewhat similar to the Claude Cowork pattern; Cowork is one toggle away from Chat.

Also follows the Grok Bot of an advanced computer use. OpenClaw for normies.

jakozaur··on How Dactyl Works
I tried it with a less typical case (the Warsaw, Poland transit app).

I'm impressed that it worked at all and got quite far. Though it is still broken in many ways (random geo locations), the menu bar on top...

The problem is that many web apps are quite simple, while native apps get trickier when it comes to UX. The LLM doesn't work well if it has to interpret animations or cross-correlate map data with locations on a map.

jakozaur··on Oldinsurancemaps.net is now a Charter Project
Two events: Anthropic did end up paying a $1.5B settlement for a case involving the use of pirated books. See Bartz v. Anthropic.

Second, apparently archiving content and providing it to others is risky business. It is often ruled as piracy, see Hachette v. Internet Archive. It went badly for Internet Archive, against common sense.

jakozaur··on AWS Acquires DuckDB
AWS rarely acquired startups, though leveraging someone else's tech is a typical pattern. It did license ParAccel and sold it as Redshift. It also runs Athena, which is based on Trino (a Presto fork).

Now, with Snowflake and Databricks earning big bucks in the intelligence era, time to gain market share by having DuckLabs under its ownership?

jakozaur··on Oldinsurancemaps.net is now a Charter Project
I wish AI leaders would fund more data archival projects, instead of hoarding it themselves. They could build a lot of goodwill, but the current state of the art is destroying old books and denying anyone access to them.

These archival projects have an enormous impact, though they are chronically underfunded. I wish there were a strong economic incentive to support them.

jakozaur··on Claude Code pricing: same tokens, same model, up to 40x the price
Yes, in the linked sources, SemiAnalysis did test for Codex too: https://x.com/SemiAnalysis_/status/2064815044085318040

Similar effect, but even 70 times higher compared to Claude, 40 times.

jakozaur··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
This isn’t limited to large system prompts. Coding-agent harnesses are also becoming more aggressive about using tools, even for trivial requests. In our tests, prompts such as “Hey” or “commit” sometimes triggered 30+ tool calls:

https://quesma.com/blog/the-true-cost-of-saying-hi-to-an-ai-...

Tokenflation seems very real: the number of tokens consumed by simple tasks keeps increasing.

jakozaur··on Claude Code May–July 2026 weekly limits promotion
Fable is also extended by another week till July 19: https://support.claude.com/en/articles/15424964-claude-fable...
jakozaur··on Show HN: Smart model routing directly in Claude, Codex and Cursor
It's rather hard to do at the proxy level with agentic coding, such as Claude Code or similar. These are long-chained sessions of tool use that heavily rely on prompt caching. Changing mid-flight is costly.

It looks like much more context is required to decide on the best model (e.g., summarizing logs might use a cheap model, whereas you likely want Opus/Mythos/GPT 5.6 to debug multithreading logic). In an agentic system, a decision about the model may be embedded in the decision to orchestrate the model.

jakozaur··on Poland is now among the 20 largest economies
Poland's first partly free election was on 4 June 1989, preceded by the roundtable negotiations.

The protests in Czechoslovakia came later, called the Velvet Revolution, from 17 to 28 November 1989. In June 1990, Czechoslovakia held its first democratic elections, a year after Poland.

Poland paved the way for the whole of central and eastern Europe. The Round Table produced the negotiated-exit template that Hungary built on in its own talks that summer, and that Czechoslovakia, East Germany, and the Baltics drew on as their regimes fell within months.

And it did so from the deepest macroeconomic crisis of any of the satellite states: hyperinflation running into the hundreds of percent by late 1989, an unresolved sovereign default from 1981, and chronic shortages.

Since then Poland has converged fastest of any of them. From a low base it has climbed to the upper-middle of central and eastern Europe by GDP per capita PPP, overtaken Hungary, and is now closing on Czechia and Slovenia.

jakozaur··on Poland is now among the 20 largest economies
The story is longer: Poland was the first country to make a remarkable peaceful transition from a bankrupt, failed Soviet satellite state. The shock therapy, plus NATO and EU aspirations, paved the way.

It is a story of a country that made a lot of the right decisions along the way. Managed to keep consistent high growth, not a pony trick or boom/bust mode.

Poland should be a role model for many other countries.

Recommend a book: https://www.amazon.com/Europes-Growth-Champion-Insights-Econ...

And Noah's blog post: https://www.noahpinion.blog/p/the-polandmalaysia-model

jakozaur··on Heat pump sales rise across Europe
A heat pump could win as the best HVAC technology, though a better drilling for ground-sourced ones. Just a shallow drilling (up to 100m) that works in retrofit mode, such as drilling from the basement, would be a great upgrade:

- No outdoor unit that looks awful in many settings

- works well, even in the coldest winter, without a spike in electricity usage, COP 5

- very reliable with long durability

- super quiet, no ambient noise

- 20% more efficient

Currently, drilling is very disruptive in retrofits, but there is progress in compact techniques that might change the equation.

Disclaimer: angel investor in https://www.flexdrill.at/

jakozaur··on $500 GPU outperforms Claude Sonnet on coding benchmarks
Build systems are tested by CompileBench (Quesma's benchmark).

Disclaimer: I'm the founder.

jakozaur··on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them
I can provide trajectories. Though probably we are not going to publish them this time. This would need some extra safeguards.

Email me. The address is in profile.

jakozaur··on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them
As I do eval and training data sets for living, in niche skills, you can find plenty of surprises.

The code is open-source; you can run it yourself using Harbor Framework:

git clone git@github.com:QuesmaOrg/BinaryAudit.git

export OPENROUTER_API_KEY=...

harbor run --path tasks --task-name lighttpd-* --agent terminus-2 --model openrouter/anthropic/claude-opus-4.6 --model openrouter/google/gemini-3-pro-preview --model openrouter/openai/gpt-5.2 --n-attempts 3

Please open PR if you find something interesting, though our domain experts spend fair amount of time looking at trajectories.

jakozaur··on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them
So many models refuse to do that due to alignment and safety concerns. So cross-model comparison doesn't make sense. We do, however, require proof (such as providing a location in binary) that is hard to game. So the model not only has to say there is a backdoor, but also point out the location.

Your approach, however, makes a lot of sense if you are ready to have your own custom or fine-tuned model.

jakozaur··on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them
Oh, nice find... We end up using PyGhidra, but the models waste some cycles because of bad ergonomics. Perhaps your cli would be easier.

Still, Ghidra's most painful limitation was extremely slow time with Go Lang. We had to exclude that example from the benchmark.

jakozaur··on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them
See direct benchmark link: https://quesma.com/benchmarks/binaryaudit/

Open-source GitHub: https://github.com/QuesmaOrg/BinaryAudit

jakozaur··on Ghidra by NSA
Funny thing, AI is not that terrible at using Ghidra. We released a benchmark on that and hopefully models will improve: https://quesma.com/blog/introducing-binaryaudit/
jakozaur··on Show HN: Ghidra MCP Server – 110 tools for AI-assisted reverse engineering
Funny coincidence, I'm working on a benchmark showcasing AI capabilities in binary analysis.

Actually, AI has huge potential for superhuman capabilities in reverse engineering. This is an extremely tedious job with low productivity. Currently reserved, primarily when there is no other option (e.g., malware analysis). AI can make binary analysis go mainstream for proactive audits to secure against supply-chain attacks.

jakozaur··on Pg_tracing: Distributed Tracing for PostgreSQL
Great idea. Currently, people have to rely on client-side spans in OpenTelemetry. However, it would be awesome if we could get spans for slow SQL queries, along with explanations.
jakozaur··on Benchmarking OpenTelemetry: Can AI trace your failed login?
In this benchmark, micro-services are really small, ~300 lines, and sometimes just two of them. More realistic tasks (large codebases, more microservices) would have a lower success rate.
jakozaur··on AI Usage Policy
See x thread for rationale: https://x.com/mitchellh/status/2014433315261124760?s=46&t=FU...

“ Ultimately, I want to see full session transcripts, but we don't have enough tool support for that broadly.”

I have a side project, git-prompt-story to attach Claude Vode session in GitHub git notes. Though it is not that simple to do automatic (e.g. i need to redact credentials).

Page 1 of 27Next →