HNHacker News
TopNewBestAskShowJobs

Zababa

5,804 karma · joined October 29, 2020

submissionscomments
Zababa··on I'm 38 and I Can't Support Myself Anymore
This assumes production of food is the bottleneck rather than distribution. If distribution is actually the bottleneck, overstocked food is not a measure of excess in society.
Zababa··on Apple Will 'Watch Everything Burn' When the AI Bubble Bursts
>At their very core, Large Language Models' costs run contrary to basically every model of selling software.

>Consumers and enterprises alike have been trained to pay a monthly fee for a service, and while these services might have limits or strictures, basically nobody buying software expects to have a metered service, let alone one that's both metered and with hard to measure costs.

Has Ed Zitron not heard about the cloud? Unpredictable AWS bills?

Zababa··on The new rules of context engineering for Claude 5 generation models
Interesting how everyone's favorite language seems to be even better in LLM era, almost like passion, skill level and having LLMs matters more than the language.
Zababa··on The new rules of context engineering for Claude 5 generation models
Maybe this one will even be successful!
Zababa··on ARC-AGI Leaderboard
>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.

They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.

Zababa··on Be skeptical of OpenAI's rogue hacker agent story
>The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.

DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper.

You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.

Zababa··on Does creatine make you smarter?
I think the way you interpret the null result is downstream of considering creatine as "a supplement". You can make the null result say anything by changing your prior about creatine or supplements. That's the issue with priors.

You also seem to reject the possibility that some things help just a bit. The author has another article in the same vein about "things that can help maybe a bit but the evidence we have doesn't really help detecting small effects" https://dynomight.net/vitamin-d/

Zababa··on Shinjuku Station in 3D
Train station: :D :D :D <3

Train station, Japan: :D :D :D <3

Transit infrastructure is really cool

Zababa··on Annoying and alarming things about OpenCode
The mix of humanizing the LLM, calling it "clanker" and being very aggressive towards it is really weird. I don't think it's a good habit to take, it feels like it could bleed into how you interact with people. Many interactions are through text interfaces these days.
Zababa··on Moonshine: Lets you stream games from your PC to any device running Moonlight
There's also an alternative which is that token prices are already high, downtimes do exist (semi frequent on Claude) or are managed by serving degraded versions/quants (speculation that has never really be proven afaik) or reducing thinking time ("juice" values from open AI), and that free tiers are not used much at all or are a loss leader.
Zababa··on What's the deal with all the random weekly quota resets for agents lately?
>Resets of the weekly quota for all users must be ludicrously expensive for these companies

Why can't people see the alternative hypothesis, inference has huuuuuge margins?

Zababa··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
Interesting, good find! Yeah I may be wrong and this may be an error in the leaderboard. Weirdly it shows no reasoning cost and no reasoning tokens used, but for example here https://huggingface.co/datasets/arcprize/arc_agi_v1_public_e... the answer is super short but it says "4945" completion tokens.
Zababa··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
I'm not referring to this paper, I'm referring to this leaderboard: https://arcprize.org/leaderboard. Set it to "arc agi 1", "base LLM" and you'll see deepseek at 57%. Submitted 2025-12-01, $0.120 per task. The paper you linked was later than that, and also says "We do not report an official ARC Prize leaderboard score".

So this paper doubled the price to get the same exact result at base Deepseek 3.2 at launch, and wasn't even tested on the verified set.

Zababa··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.
Zababa··on Mozilla: The state of open source AI
Haven't tried Kimi K3 for now but there was a huge difference between GPT 5.6/Fable and GLM 5.2/Kimi K2.7 that were previous frontier open models.
Zababa··on The state of open source AI
The thing about not much difference between models and the harness making them deterministic and useful is wrong. Also models have different strengths and weaknesses and some are better at almost everything by a large margin compared to others.

As for your speculation, I think it's hinging on some companies releasing models for free or no big differences between models. In a world with hyperscalers and companies training models you can quickly recreate Anthropic or OpenAI by having an hyperscaler ally with a model training company, train a good/a better model, and not release it.

Zababa··on Pebble Mega Update – July 2026
This is true, but there's a big difference between saying "15-20 hours battery life, which is 11000 5-seconds activations, which last you a few years with 10 5-seconds activation a day" and "years of battery (btw in small text the real number is given). Especially since they mention that this project is hackable/you can do other things with it, knowing in advance you have something like ~100k button presses means some projects feel perfectly and some others won't really work.
Zababa··on Pebble Mega Update – July 2026
I don't really care about the environmental consciousness, my issue is that presenting a product with a battery that lasts for years when it actually lasts 15 to 20 hours makes me feel like I'm being lied to.
Zababa··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
"don't reinvent the wheel" isn't a law of physics and I think is mostly said by people that never designed anything with wheels.
Zababa··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
I think I'd typify it as "ARC-AGI doesn't matter" more than "harness matters". Or maybe "harness matters for some very specific tasks".
Zababa··on Pebble Mega Update – July 2026
"Battery that lasts for years" being actually 12-15 hours of recording is a huge turn off honestly.

>How long does the battery last?

>Roughly 12 to 15 hours of recording. On average, I use it 10-20 times per day to record 3-6 second thoughts. That's up to 2 years of usage.

They then say:

>Wait, it's single use?

>Yes. We know this sounds a bit odd, but in this particular circumstance we believe it's the best solution to the given set of constraints. Other smart rings like Oura cost $250+ and need to be charged every few days. We didn't want to build a device like that. Before the battery runs out, the Pebble app notifies and asks if you'd like to order another ring.

My oura has lasted ~3 years, I recharge it twice a week usually, and I think it has spent way more than 15-20 hours turned on.

Zababa··on Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
Considering progress in the rest of computing stuff (RAM, CPUs, storage) is kind of "linear"/exponential it sure looks like it's an engineering gap and we're on the right track. GPT 3 was 175B parameters and is today crushed by models that are 32B parameters, that's a lot of progress in 6 years.
Zababa··on Sleep regularity is a stronger predictor of mortality risk than sleep duration (2023)
Yeah, in an ideal world where I'm the ideal me I wouldn't use my phone in my bed, but I haven't found a way to stop doing that which I can stick with, so I try to limit the damage.

Part of what I wanted to say is, there is conventional wisdom, then there is how you actually put that wisdom in practice in a way you stick with. I've struggled a lot with the implementation, but sometimes by throwing lots of stuff at the wall I find something that brings me halfway there. It's not the "golden way" but it leaves me in a better place than before, with a bit better sleep, a bit more self knowledge, and a small victory.

Zababa··on Sleep regularity is a stronger predictor of mortality risk than sleep duration (2023)
Hard to answer precisely without knowing what conventional wisdom didn't stick.

The common levers I know and that worked at least a bit for me:

- start by having a fixed waking time, and get sunlight or bright light quickly after waking up. Normally relatively fixed sleep time is supposed to follow. For me waking up is the easy part, transforming that into getting up and going outside is harder. Another option here is a strong (like, really strong) lamp on a timer, or letting the morning light in your bedroom (this one is usually not recommended I think, most people seem to be blackout curtains style, but for me it gave me a nice 6am waking time with good sleep last summer).

- melatonin. Two main ways: using it as a kind of hypnotic, so ~30 minutes before sleep, experimenting with 0.3mg to ~2mg doses ; then using it as a circadian regulator, this is a good resource https://lorienpsych.com/2020/12/20/melatonin/, search for "TO TREAT" in the page.

- app timers, for me it was mostly no twitter and no youtube, or a very low time for each.

- light, ie reduc light before sleeping. Not just blue light and not just screens, if I'm on my phone in bed I'll reduce the luminosity a lot, same with computer, same with e-reader. I also try to avoid using too much the lights in my room. More light tend to make me feel more "wired" and less ready to sleep.

- "meditation" to cut rumination, by which I mean "lay down in my bed, gently try to find sensations in the body and to stay focused on them, by gently I mean it's a very low stakes game where the goal is to find sensations in the body and give them attention, but losing focus for a while is not a big deal".

- shower in the evening, as I don't like feeling dirty when I am in my bed, but also not just before bed as sometimes I don't really want to go take a shower and this delays my bedtime

- clean bedsheets, bedroom, stuff in/on your bed

- AC in the summer, I wouldn't be able to sleep properly without it

- sleeping mask. It helps going to sleep, but it falls of my head every night so it doesn't prevent waking up with light too.

- making getting good sleep the priority of the evening. This is easy/possible for me due to my circumstances (ie low responsibilities in the evening). The way I do it is that unless something is actually important, what I'm trying to accomplish in the evening is prepare myself for sleep and get good sleep. This can look like not starting a movie at 11pm, not booting up games, not eating a super heavy meal, not drinking too much water after 6pm to avoid waking up to pee, if I have things I want to do try to do them early so they're done earlier, move some stuff I want to do every day like spaced repetition in the morning.

Zababa··on How I use HTMX with Go
I've tried to like Go with HTMX, but the big issue was always Go templates. I feel like if there was something like JSX/TSX but for Go, it would be a way better dev experience, but right now it's mostly a pain. Templ tries to go in that direction, but a year or so ago editor integration and tooling weren't great.
Zababa··on Bonsai 27B: A 27B-Class model that runs on a phone
There is value in splitting things but there is also a cost. You have to train the specialized model, for that you have to know your use case, you have to hope the use case is going to be stable over time, you then have to see if you can remove english -> slovakian or coding from a model without affecting the useful parts.
Zababa··on A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese
That makes sense. It seems a bit too slow to be really useful, as in, I feel like you'd want either to replay the audio as much as you want, or you can kind of listen passively/actively to something simple and slow.
Zababa··on A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese
Copy-pasting "ずんだもん" into Google gives you everything you want with the sidebar info, copy pasting "zundamon" into Google gives the same Wikipedia link on the sidebar. "Popular TTS character called" is enough to imply that what follows is the name, that you can then search.
Zababa··on A voxel Tokyo in real Japan time – ride the Yamanote line and study Japanese
This looks good and I like the idea but I don't get what the "practice" is here?
Zababa··on Cargo-nextest: 3x faster than cargo test, per-test isolation, first-class CI
Fungible/non fungible is a good alternative, and maybe the technically correct word. But I think in that case it doesn't apply and the change the author did is better.
Page 1 of 34Next →