HNHacker News
TopNewBestAskShowJobs

ricardobeat

22,735 karma · joined August 16, 2010

working on the boomkat javascript engine: https://github.com/ricardobeat/boomkat

contact: hn at ricar.do

submissionscomments
ricardobeat··on AI Is Throwing a Roadside Picnic
Why wouldn’t the AI be able to answer questions about the proofs or organize a talk? Those are pretty much the things it excels at, exploring and summarizing previous knowledge. The problem is the volume of questions and data a human needs to step through to arrive at the same understanding.
ricardobeat··on WSL3 Performance is about 5-60% faster than WSL2 depending on the workload
herdr might be the culprit there, it can significantly slow down tui apps with its terminal capture. This happens on mac/linux, WSL will surely make it worse.
ricardobeat··on Germany transforms former coal mines into Europe's largest lake landscape
Off-topic, but what a horrible mobile experience on this site. The screen is >50% covered in ads, and both close buttons are a trap that let clicks through and open a new tab.

Those arrows also initially align perfectly on top of the featured photo, making it look like a gallery, but actually navigates to another article (more ad views, yay).

https://postimg.cc/ftkRMXwb

ricardobeat··on A font recreated from photographs of classic Commodore 64 keycaps
What a confused comment.

"There is no RETURN, SHIFT...", "Too many F keys". They are confusing a typeface with trying to replicate the keyboard keys.

The arrows are not the cursor keys, but based on the ESC key.

They are right about the @ character though, it looks quite wonky.

ricardobeat··on ESP32-C3 Adblock
Let’s not trivialize “human rights”, those are much more fundamental issues and this kind of discourse does not help achieve anything.
ricardobeat··on ESP32-C3 Adblock
I use arduino-cli exclusively, lighter than ESP-IDF and works completely standalone from the IDE.
ricardobeat··on Strands Decider 2B: a small, open-source, decision model
Most LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.
ricardobeat··on Strands Decider 2B: a small, open-source, decision model
I wish these would stop using JevBench. It focuses way too much on text classification tasks, and some of the models perform very poorly on tasks that need actual intelligence.
ricardobeat··on Mistral Large 4
Ever seen a movie chase scene where cars are going right to left?
ricardobeat··on Mistral Large 4
just “a random word” gives you Zephyr in Gemini, and “Lantern” in Claude and ChatGPT.
ricardobeat··on Mistral Large 4
These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.
ricardobeat··on Mistral Large 4
They don't publish numbers, but Anthropic has a single DC with 200k+ GPUs for inference, GPT-6 Astra is said to have trained on 100k+ GPUs.
ricardobeat··on Using Blu-ray M-Disk as backup of last resort
You could say so from a consumer perspective, but it is not what the MTBF you quoted measures. Read errors are not considered a hardware failure. Not an opinion.

That's a lot of work to avoid accepting your previous comment was incorrect.

ricardobeat··on We are going to kill "unalive"
Probably to make the point that suicide is usually caused by something, not merely a self-inflicted choice or an act of violence.
ricardobeat··on Using Blu-ray M-Disk as backup of last resort
MTBF is for hardware failure, not data loss.

This recent study I found [1] shows almost 4% HDD loss after only 2.5 years for consumer hardware, enterprise products being a bit better.

[1] https://research.cs.wisc.edu/adsl/Publications/latent-sigmet...

ricardobeat··on Worth Building
I wouldn't read too much into it, as that sentence is also clearly AI juice.
ricardobeat··on Web Search API
This practice of having the one provider should be eliminated. Companies self-inflict lock-in to large platform providers, prevent their own teams from using better technology options and stifle innovation. It's crazy that even with a pile of SOC/ISO/PCI/HIPAA/NIS certificates, procurement is still a months-long process, it should be much easier to do business.
ricardobeat··on Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
What are you running it on? I'm getting mixed results using the IQ3_XXS quant which supposedly matches baseline, it feels significantly degraded.
ricardobeat··on Show HN: Local pretrained classifiers, GPU not needed
This is like creating a markov chain bot and naming it ChetGPT.
ricardobeat··on Getting the most out of Opus 5.5 in Claude and Claude Code
Plan mode has become pointless since Opus 5 came out, they know when to switch between planning and execution now. But that iteration/discussion is still necessary unless you're building completely blind - the model cannot read your mind.

I've had it running 8h+ of non-stop optimizations, chasing a performance target, rewriting systems or building a series of prototypes for research. All it needs is a clear goal.

ricardobeat··on Cloudflare OHTTP gateway
So.. they continue having access to private identifiers, while you willingly give it up to "protect privacy"? Piping all of that data into a massive central database instead of your nginx logs? How is this supposed to be better?
ricardobeat··on Clef: Open-weight decision models, and new RL fine-tuning platform
Because with a tiny model you're skipping all the intelligence and world knowledge that makes it useful without fine tuning. `typed-decisions` is almost entirely text classification tasks.
ricardobeat··on Clef: Open-weight decision models, and new RL fine-tuning platform
In my experiments decider-4B performs better than Kev with significantly lower latency. It's remarkably good for it's size, shame it wasn't included in the benchmarks. Laya on the other hand shouldn't even be featured - despite being 'the original' decision model, it can only do simple text classification and is nowhere near usable performance for anything else.
ricardobeat··on Pi 1.0
All of which would require shipping a compiler along with the app.
ricardobeat··on Pi 1.0
I enjoyed Pi for a few weeks, but ultimately moved to other harnesses - the plugin ecosystem became a sea of slop, large vibecoded projects that don't work at all, and yet have thousands of stars. At some point I gave up trying to get subagents working.
ricardobeat··on Livenerf: Has Opus 5.5 been nerfed yet?
This is the case since Opus 5, the latest models (from all providers) favor using shell tools instead of the View/Edit tools available in the harness, and “accept edits” doesn’t let those calls through. Auto mode is the best option.
ricardobeat··on Livenerf: Has Opus 5.5 been nerfed yet?
“Nerfing is a myth” - “I made an imaginary chart to show you why”
ricardobeat··on Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
Please don’t post AI generated replies.
ricardobeat··on Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound
This will only work for simple text classification tasks, which is the least interesting possible use of Jev.
ricardobeat··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
You might be overusing subagents. Especially with a chatty model like DS, you’ll be wasting millions of tokens on re-discovering the project and facts instead of actual reasoning.
Page 1 of 34Next →