HNHacker News
TopNewBestAskShowJobs

jampa

1,971 karma · joined January 25, 2016

Hello there!

hn [at] jampa [dot] dev https://jampa.dev

submissionscomments
jampa··on The Test
I think the reality is even simpler: Huang, Altman and Amodei are in the end of the day head salespeople of their respective company… and they use these social platforms to sell their products.

Their job is to convince people that AI has a large upside while minimizing any downside, they dont need to believe in those themselves.

jampa··on Portal by Spotify cut my Claude Code token usage by 90%
> I've never had an issue with Codex or Claude reading massive files

Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).

I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".

jampa··on Gemini 3.8 Flash and 3.8 Flash Cyber
I wasn't trying to be precise originally, I just tried to fit activities into "morning / evening" buckets. I did the whole itinerary with Opus first, but when I gave it to Gemini 3.7 Flash to review, it started correcting it with "this place will close 5PM" or "this place is closed for good".

It was right on every nit, so it was surprising how well the model knows these things. If I ever release this I'll probably need the SERP API or Google Maps SDK (which I've heard is very expensive now), but for a personal trip where I will verify manually, using the LLM is okay for now.

jampa··on Gemini 3.8 Flash and 3.8 Flash Cyber
Eh that one is on me, if I think too much about my HN comment I end up deleting before posting it. I rely on the 1 min `delay` set in the profile page to fix before it goes live, but for some reason this time it was set to 0.
jampa··on Gemini 3.8 Flash and 3.8 Flash Cyber
I asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.
jampa··on Gemini 3.8 Flash and 3.8 Flash Cyber
I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:

- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.

- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.

- Document parsing (extracting the relevant trip info from PDFs).

If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.

jampa··on Ask HN: What is one simple thing LLMs are insanely bad at?
I tried doing something like this Ox Alpha with Opus advisor, having it work layer by layer (specs -> rooms -> room graphs ...), but each deliverable ended up a mess.

The curious thing is when I pointed out the flaws it fixed them quickly, but it's not something it can do without supervision, and supervising it takes more effort than doing the blueprint myself (to be fair, I'm not an architect, so I'm not the best at steering an LLM for this task).

jampa··on Ask HN: What is one simple thing LLMs are insanely bad at?
Serious answer: no model ever gets close to writing an architectural floor plan that makes sense.

They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.

But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.

Floor plans should be the new Pelican Benchmark.

jampa··on Opus 5.0 drives incoherence into the stratosphere
Opus 5 feels like a downgrade from Opus 4.8 overall. It, along with Fable, really has a problem following instructions and staying in scope, and their prose keeps growing, both in explaining what it did and in writing multiline code comments (some comments read like a changelog, e.g. `// sky is blue (changed from red on 2026-01-01 per TCK-234 by @Foo)`).

Every time I ask it to do something, it does 80% of the job, goes off on "side quests" beyond the scope, and then leaves something out of the core ask (and when you tell it to finish, it does the same thing again).

The only advantage of Opus 5 over 4.8 is the better cutoff date for working with 3rd-party tools, though both do a very bad job of "this tool is constantly updated, I should look for the latest version first".

jampa··on AI;DR (AI; Didn't Read)
I used to do that, but the person could be bad at prompting too (which is often why the LLM couldn't give a good answer in the first place).

So the polite version I use now is: "Hey, just to get a bit more context, what was the original problem you were trying to solve?".

That gets them to distill their own problem a bit further.

jampa··on Plug-in solar is coming. Plug-in batteries should follow
The new rule here requires you to pay for part of your imports (up to ~60%), even if you export more. I know people who installed batteries in the inverters because of this. It is more expensive overall (hybrid inverter + LFP batteries) but has other benefits (blackouts are not a problem, most months you don't even use the grid). I imagine that once sodium-ion battery production scales and costs decline, this will become the "default" option for a solar installation.
jampa··on Kimi K3-256k
Anthropic has better SLA in Germany. I’ve heard uptime there can get up to nein nines.
jampa··on SpaceX wants to launch 100k more Starlink satellites for 100x the bandwidth
When COVID hit, I knew a lot of engineers who decided to move to rural areas / small farms because they could leverage Starlink to work remotely.

Last year, when I asked whether they still liked Starlink, all of them said it is amazing, but they had gotten fiber coverage in their area from a local provider, so they don't use it anymore, or just use it as a backup.

I think Starlink was a huge demand signal that there were people willing to pay a premium for faster-than-radio internet. So, unless they manage to be cheaper and faster than fiber, I don't think there is much of an endgame there.

But there are a few places that will need Starlink, like planes, cruise ships, and islands. I'm just not sure if that will justify that $1T valuation.

jampa··on I think I have LLM burnout
I feel the same way about consumer AI tools now. Gemini and ChatGPT have been abysmal lately. They can no longer be relied on to do multi-turn searching and thinking.

Before, they could stay in thinking mode for more than 7 minutes. For example, "find a source for this claim" would search, analyze, and self-adjust the query. Nowadays, even if I push for it, I cannot make these tools work for more than 30 seconds before they give generic answers, even in "Pro" mode.

jampa··on Ask HN: Where is the programming profession going?
From what you said: Not looking at code is bad, not because Claude can slip a few bugs (it can), but because LLMs tend to default to writing more code and features than needed, which isn't a good thing. I see a lot of people making 10+ PRs per day, but most of them are just going back to fix earlier PRs.

Claude always likes to "go big," for example, by choosing tools that can support millions of concurrent users or by adding unnecessary layers of abstraction that create more maintenance pain. I guess that's good for LLM companies, since more tokens are spent fixing the mess it caused.

Every time I enter plan mode for a huge feature, I end up cutting about 30-60% of the task scope before the LLM can actually start the work. I review the final code, and I still find things to cut. As said before "The best code is no code, or code you don’t have to maintain" [0]

0: https://www.simplethread.com/20-things-ive-learned-in-my-20-...

jampa··on Why Is Claude Turning into an a**Hole?
This post needs some examples, because I have never had an interaction with Claude that made me think this way.

LLMs generally have a way to "play a role" (most earlier prompt guides ask you to start with "You are a <role> expert in a <domain>"). So maybe if you interact with it by asking questions, it might assume that it knows more than the operator and adopt that attitude?

jampa··on Claude Fable is relentlessly proactive
I tested it to fix React Native bugs in a project, comparing it with Opus. It fared better on harder bugs, taking less time to find the root cause, but after implementing a fix, it spent a lot of time and effort on validation. This was mostly unnecessary, since most of the bugs were in the JS code, so for most things, hot reloading is enough for E2E validation and to run just the right tests. No need to run a full build and test suite (which takes 10+ minutes); the CI can do this.

I switched back to Opus because of this validation quirk. Overall, Fable spent 20% of the time on coding and 80% on validation.

I think using Fable for planning and Opus for execution could be a "best of both worlds" approach (I need to test this more), but for most cases, it's not necessary, and Opus is enough.

jampa··on Claude Fable is relentlessly proactive
Fable feels like a version of Opus running on a harness that won't let it halt until it's sure the issue is fixed, which makes sense if what you want is a model that's better at benchmarks.

It's a very good model, but it comes at a huge premium: not only do the tokens cost more, but the model itself really wants to spend them all. For example, working with React Native, Fable never just says "okay, I did the thing, that's it." It tries to rebuild the entire app from scratch, run the whole test suite, and watch every log and warning.

This is the first time with LLMs I've felt that upgrading to a model isn't worth it, even if my company lets me use it, because all the building / testing was just destroying my machine and its battery, which keeps me from working on other things.

For now, it feels like Opus with ultracode is a better choice (less pollution of the main context, more parallelism in investigations).

jampa··on What it feels like to work with Mythos
It is hallucinating many flights in my region, some that never existed (so it is not an outdated data problem).

I also see some logic flaws. It overlooks the option of going to a major hub to access faster aircraft, rather than hopping on local hubs.

Also, immigration and customs are cleared at the first airport you arrive at in the country, not at the last one.

In some countries, you need to clear immigration even while going to a third country, so 1 hour is not enough to do it.

jampa··on Wind and solar generated more power than gas globally in April 2026
> there's not enough money to be made via speculation

I mean, there is money to be made. CATL stock (the major producer of EV batteries with 50% market share, with billions of contracts for stationary batteries) rose 48.81% over the last 6 months, for example.

But I agree that news about renewables goes unnoticed. I only see news about renewables because I actively seek out channels and websites that cover it. I wonder if it is because most companies in the industry are Chinese and don't focus on PR in the West as AI companies do.

jampa··on Morningstar values SpaceX at $780B, half its IPO target
The biggest competitor to Starlink is, ironically, traditional fiber.

When COVID hit, I knew a lot of engineers who decided to move to rural areas / small farms, because they could leverage Starlink to work remotely.

Last year, when I asked whether they still liked Starlink, all of them said it was amazing, but they had gotten fiber coverage in their area from a local provider, so they don't use it anymore, or just use it as a backup.

I think Starlink was a huge demand signal that there were people willing to pay a premium for faster-than-radio internet. So, unless they manage to be cheaper and faster than fiber, I don't think there is much of an endgame there.

jampa··on I believe there are entire companies right now under AI psychosis
The last three times I filed detailed bug reports as a client, all I got back were AI replies asking the same questions I’d already answered in the original report and suggesting alternatives I’d explicitly said I’d already tried. No wonder people don’t write bug reports anymore.
jampa··on Two millionth electric car registered as market rebounds from tax changes
This oil crisis was a huge boon for EVs. In Brazil, despite the "hate" most people have against EVs, BYD went from breaking into the top 10 in March to taking the #1 spot in consumer sales for the first time ever.
jampa··on Claude Opus 4.7
Mythos release feels like Silicon Valley "don't take revenue" advice:

https://www.youtube.com/watch?v=BzAdXyPYKQo

""If you show the model, people will ask 'HOW BETTER?' and it will never be enough. The model that was the AGI is suddenly the +5% bench dog. But if you have NO model, you can say you're worried about safety! You're a potential pure play... It's not about how much you research, it's about how much you're WORTH. And who is worth the most? Companies that don't release their models!"

jampa··on Software design is now cheap
I think the main point is that "reinventing the wheel" has become cheap, not software design itself.

For example, when a designer sends me the SVG icons he created, I no longer need to push back against just using a library. Instead, I can just give these icons to Claude Code and ask it to "Make like react-icons," and an hour later, my issue is solved with minimal input from me. The LLM can use all available data, since the problem is not new.

But many software problems challenge LLMs, especially with features lacking public training data, and creating solutions for these issues is certainly not cheap.

jampa··on YouTube's $60B revenue revealed amid paid subscriber push
Not sure if WhatsApp paid off, though. There are reports of up to $1 billion in annual revenue with the Business API, so this is far less than what they paid. I think Meta's strategy was to create a Western version of WeChat, which has a very high ARPU, but for some reason, they never invested in it properly... They added Stories, a "Venmo" feature, and then gave up.
jampa··on Top downloaded skill in ClawHub contains malware
Thanks for the write-up! Yes, this clearly shows it is malware. In VirusTotal, it also indicates in "Behavior" that it targets apps like "Mail". They put a lot of effort into obfuscating the binary as well.

I believe what you wrote here has ten times more impact in convincing people. I would consider adding it to the blog as well (with obfuscated URLs so Google doesn't hurt the SEO).

Thanks for providing context!

jampa··on Top downloaded skill in ClawHub contains malware
This article is so frustrating to read: not only is it entirely AI-generated, but it also has no details: "I'm not linking", "I'm not pasting".

And I don't doubt there is malware in Clawhub, but the 8/64 in VirusTotal hardly proves that. "The verdict was not ambiguous. It's malware." I had scripts I wrote flagged more than that!

I know 1Password is a "famous" company, but this article alone isn't trustworthy at all.

jampa··on Claude Code daily benchmarks for degradation tracking
I am using API mode, and it's clear that there are times when the Claude model just gives up. And it is very noticeable because the model just does the most dumb things possible.

"You have a bug in line 23." "Oh yes, this solution is bugged, let me delete the whole feature." That one-line fix I could make even with ChatGPT 3.5 can't just happen. Workflows that I use and are very reproducible start to flake and then fail.

After a certain number of tokens per day, it becomes unusable. I like Claude, but I don't understand why they would do this.

jampa··on Things I've learned in my 10 years as an engineering manager
There are already some great responses, but I want to add that one effective way to coach senior employees is to give them responsibilities one level above their current role and then provide feedback.

For engineers aiming to move into management or staff engineering, you can assign them a project at the level they aspire to reach and give feedback once they complete it. For example, for an engineer aiming to be an EM, I expect them to lead not only meetings but also all communications related to this project, while I act as their director. Afterwards, I provide feedback.

It doesn't have to be that extensive right away. You can start small, like asking them to lead a roadmap meeting, and then increase responsibilities as they improve. Essentially, create a safe environment for them to grow.

Page 1 of 6Next →