HNHacker News
TopNewBestAskShowJobs

mixermachine

322 karma · joined July 6, 2017

submissionscomments
mixermachine··on Pi 1.0
Copied from myself above:

That goes against the philosophy of pi. Nearly Everything is a Plugin. Need to set the reasoning effort? Install pi-reasoning. Need subagents? Install pi-subagents. Need permission gating? Install one of the permission extensions (otherwise the agent can do everything).

A feature like cache warming on Anthrophic feels very strange in this environment. NOT saying that it is not useful. It just should be a plugin.

mixermachine··on Pi 1.0
Pi with Caveman Plugin :D
mixermachine··on Pi 1.0
That goes against the philosophy of pi. Nearly Everything is a Plugin. Need to set the reasoning effort? Install pi-reasoning. Need subagents? Install pi-subagents. Need permission gating? Install one of the permission extensions (otherwise the agent can do everything).

A feature like cache warming on Anthrophic feels very strange in this environment. NOT saying that it is not useful. It just should be a plugin.

mixermachine··on What California is learning from solar panels built over irrigation canals
I wouldn't touch fields or canals. The US has incredible amounts of parking lots. They are also mostly built in very developed parts of the land with a lot of users for electric power in close proximity. And a parking lot is already "sealed ground". Might be less problematic to get a permit there?

Additional plus: cars don't heat up in the sun. And at some point you might even charge EVs there.

mixermachine··on Nvidia announces native GPU programming in Rust
After reading through the threat here it seems more like a cultural thing. The US has quite a lot of filters for profanity. I remember from my youth that in 2009 Eminem was a guest in a Germany TV show and very happy to swear as much as possible without being censored. https://www.youtube.com/shorts/2OC-yKZ5Yag
mixermachine··on Nvidia announces native GPU programming in Rust
Am I, with around 30, in this generation? Putting * in words seems like self censorship to me. Still, might have a cultural component. German here.
mixermachine··on Breaking the 1.58-bit Barrier for Ternary LLMs
I also no longer trust benchmarks on this one. When the context gets a bit longer and the problem harder low quant models often produce worse output for me. Sometimes they even loop.

Interestingly different formats also often behave differently. GGUF unsloth is so far the best for me.

mixermachine··on Let's make quality the norm again
Project Farm on YouTube has never disappointed me. The guy buys everything himself and finds creative & reproducible ways to test things.
mixermachine··on DeepSeek v4.1 Flash
The engram stuff is great because RAM is often still cheaper (or at least expandable). My company does currently look into buying some hardware as we handle confidential data and code.

Qwen 3.8 Flash is viable on two Nvidia 6000 96GB with a wood quant because you can put the 50GB Engram into RAM and the hit should be below 10% performance. At least that is what I have seen so far. Correct me if I'm wrong.

mixermachine··on I'm becoming AI-blind
I need to support integrators for a mobile app SDK. Bank stuff.

We have one integrator which consistently uses a truly bad AI. My brain skips after two sentences already. There are always +9 questions what should just be 2 max. Redundant info is requested (e.g. please provide a change history of this function and if it will be deprecated) and everything is just unbelievably verbose.

90% of questions we get from them are already answered by the web documentation (which even has a search function) or are just non sense.

I'm truly considering to build an MCP just for them...

mixermachine··on Qwen 3.8 27B
Did you already test TranslateGemma? I use this model for my Android Studio Translation Plugin (https://plugins.jetbrains.com/plugin/30265-localizepipe) and so far it produces great results for its size.

If there are other models (of similar size) out there, that are better at this, please let me know.

mixermachine··on AMD acquires Taalas to boost inference performance by etching models in silicon
Scaling a model on a chip is quite hard. ChatJimmy is based on Llama 3.1 8 billion. Kimi K3 has 2.8 trillion parameters. That are 350x more parameters.

I would expect that Gemma 4 E2B (approx 5.1 billion parameters) or maybe even Gemma 4 26 billion A4B at some point is running on a chip.

mixermachine··on Teardown: A Generic 7-Port USB 3.0 Hub That Wasn't
Channels like project farm https://youtube.com/@projectfarm or other reviewers that are not sponsored are truly my main source of information in this age.

Some direct reviews between 2 and 4 stars are also sometimes useful. Always discard the 5 star ones...

mixermachine··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
If there are no changes to the system message, yes this is possible and also likely done by Anthropic. When there are additional local MCPs, reusability will be lower.

I think that Anthropic will bill you in any case :D

mixermachine··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
right, but the cache retention time is very short for Anthropics LLMs. 5 minutes or 1 hour (with additional costs). So you have to prompt basically non stop to not get a cache eviction.

Anthropic even changed this silently: https://www.reddit.com/r/ClaudeAI/comments/1sk3m12/followup_...

mixermachine··on Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k
The OpenCode CLI does not work as well for me as the PI CLI. I'm a subscriber of OpenCode Go (the sub, good value for me really) but I had not great experiences with OpenCode CLI. It multiple times with different models deadlocked itself into listing endlessly to non ending processes (Android Debugging Bridge, COM serial log, ...). There was also a problem where the OpenCode CLI would crash after sometime with a Bun error.

I switched to the PI CLI and have no problems with hanging processes anymore. OpenCode Go allows for API access so I'm keeping this sub.

mixermachine··on A complete ClickHouse OLAP engine, compiled to WebAssembly
Works on Brave (Chromium) with Android 16
mixermachine··on One million passports leaked online
It is quite interesting how this is handled world wide. For me PII is very sensitive and I advice people to be very cautious. Every business in the EU (were I live) also has to be very careful with such data by law. Fines are now at a level were they can hurt the business significantly.

During vacation in an Asian country on the other side all of this was basically a no brainer for smaller to medium businesses. I once rented a scooter there and the business owner had all her documents organised in WhatsApp chats. Including now my passport plus drivers licence... The people in general in that country were also very relaxed when it came to giving out their contact details to random businesses.

I don't want to throw shade on them, thus no country name. Incredible friendly and welcoming people there.

mixermachine··on Was my $48K GPU server worth it?
Right, I did swap that. Still, you have to pay that 4k then every year and give out the code. I also assume that prices will go up as no AI company (but NVIDIA -> selling shovels) is currently making any money.

For some projects the giving out the code part might be ok (i use Codex there too) but for the core app at the company I'm working at there is currently a strict no-AI policy. A local GPU solves this.

mixermachine··on Was my $48K GPU server worth it?
With parallelism of 16 you can still get around 25 to 30 tokens per user when all 16 channels are running. Not everyone will use the model at the same time but it certainly will be tight, especially for agentic coding. For pure chat applications this should be quite fine.
mixermachine··on Was my $48K GPU server worth it?
The 5h quota of Codex Pro on GPT 5.4 Medium lasts me for around an hour and a half, maybe 2 hours. And this is already the "savy" setup. Enable GPT 5.5 High fast and you will be beached in 30 minutes with active development.

For continues all day work you definitely need a higher tier sub level.

I'm actually looking into deploying a GPU at my company because we can not give out our code. Qwen 3.6 looks good

mixermachine··on Removing the modem and GPS from my 2024 RAV4 hybrid
That is crazy. 5 years and they are already shutting down the servers? They should be forced to open up the API when they shut it down. Running a replica yourself should be pretty doable.
mixermachine··on The 'Hidden' Costs of Great Abstractions
RAM + GPU are getting more expensive but mostly for applications that require a lot of it like AI. The hardware cost for regular applications has not vastly increased (especially when factoring in inflation). Spending 2x development time on a problem often is not worth it (or only with large deployments).

UI development is an even more special case here. The customer buys the machine which runs the code, not the company. So sadly "good enough" is the standard.

One example for me here is the "switch product option" button on Amazon listings (e.g. switch green to blue color, smaller to larger model). On my phone this sometimes takes >5 seconds to properly load. Horribly optimised.

mixermachine··on Stop trying to engineer your way out of listening to people
In my experience in software architecture, drawing a diagram often saves you >60 minutes of discussion and potentially multiple meetings. This works even with a badly drawn but truthful one.

Use an Ai agent + Mermaid.js for a quick scribble if you are in a remote meeting. Use white boards or pen + paper in a local meeting.

Diagrams are so much clearer then words, especially if the concept or logic in question is not trivial.

mixermachine··on Are the costs of AI agents also rising exponentially? (2025)
For programmers maybe. I do this too. But think about all the regular users out there. Your dad and your mum, maybe even your grandparents. This is a huge marked too and for that we can use these special chips at scale.
mixermachine··on Are the costs of AI agents also rising exponentially? (2025)
The cutting edge, max size models will likely stay in the GPU space for a long time. But these models are not needed for most general requests. With a fine tuned 30B quantisized model you can serve a large portion of requests with around 32GB of RAM. Free users will likely only get these kinds of models.

At some point we will get these models in hardware and the cost per token will be minimal.

mixermachine··on Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference
What spec of Framework Desktop do you run this on?
mixermachine··on Pro Max 5x quota exhausted in 1.5 hours despite moderate usage
I'm using the Codex Business subscription (about 30€) already for multiple months. Even there they cut back on the quota. A few months back it was hard for me to reach the limit. Now it is easier.

Still, in comparison with Claude Code, the quota of Codex is a much better deal. However, they should not make it worse...

mixermachine··on Sweden goes back to basics, swapping screens for books in the classroom
Fully agree. I went to school in Germany and many of our textbooks were free there. Sometimes you would get a textbook that is already >= 10 years and out of shape but who cares? Especially the basic knowledge does not change often. Buying all these textbooks new every year feels like a scam to me as they are then only used for one year by the pupil.

Btw when you damaged a book beyond repair, you needed to pay the full price. Only the exercise books needed to be bought freshly as they were "used up" fully after the year. Still, they were often seen as optional.

mixermachine··on Solar is winning the energy race
This seems quite strange to claim. Basically every city in the developed world already has power plants on the outside and a lot of wires to get the electricity in
Page 1 of 6Next →