HNHacker News
TopNewBestAskShowJobs

gopalv

4,060 karma · joined August 20, 2012

@php.net / @apache.org / @isotopes.ai

https://aidnn.ai

submissionscomments
gopalv··on RIP, vector database
> This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

This is a direct parallel to how Postgres and Mysql built indexes.

Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).

But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

[1] - https://www.uber.com/us/en/blog/postgres-to-mysql-migration/

gopalv··on Gemini 4 Argon
> taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.

This is good, but they're the slow mover due to this exact thing.

Google is getting punished for not letting the models enter an echo chamber and go faster than humanly possible.

gopalv··on Claude Opus 5.5
The whole thing reminds me of the Apple feature flag story[1] from a generation ago.

[1] - https://news.ycombinator.com/item?id=6372466

gopalv··on Cloudflare Quick Tunnels
> it’s missing the bigger integration offerings Tailscale has/does still.

Until I had tailscale serve generating valid certs, I had a good reason to use Cloudflare tunnels.

But in general I don't want to put everything on the internet side of things.

Mostly, I don't want something open, but more like a "share with" for people who are in the same office (virtually over tailnet, not physically on the same LAN).

This still works great for a demo instead of a product pitch, to send an link out to see something.

I'd still use a real host over a laptop for those.

gopalv··on Show HN: Toast, a by default in-terminal IDE
The "Emacs doesn't have a file-tree" is when I stopped short and wondered what this person's background is in IDEs.

My emacs days are 1999 during my "functional or get out" phase, but my .emacs had close to 4000 lines in it, including xterm mouse support, a clone of midnight commander + sftp to a wehbost and deep integration with ctags/cscope/c++filt. There was even a line by line debugger with custom .gdbinit commands - look at the php project's gdbinit[1] if you want an idea of how to walk your own super special data structures.

You can't put Emacs in an IDE comparison and say anything about features without at least one person saying "well, actually".

I eventually switched to vim but that had to do with needing to ssh into 100k+ machines on support rotation & never having emacs on any of them.

[1] - https://github.com/php/php-src/blob/master/.gdbinit

gopalv··on Tyranny of Optionality
> doesn't make you a martyr because you get anxious

Every person's biggest problem is their biggest problem and it eats up their life in the biggest way possible.

Maslow's hierarchy is not a "smaller and smaller" problems list, it is a definition of what your biggest problem is & the solutions just scarcer as you climb up the ladder.

It's close to 20 years since I wrote about the same "problem" [1] in my life and the intervening years have brought much more grief, toil and resilience. Death of several loved ones causes you to consider what a virtue sounds like an eulogy and what yours will sound like in the end.

As an aside, I've felt like literature & art was backwards in the way it was introduced to me - where the first came "man vs nature", then man vs fellow man & then a "man vs self" battle. After fighting self and others, most of us will succumb to nature instead.

So, whenever someone vents or complains, just remember that this is not a contest or at least rage bait, but to be seen as a genuine struggle within yourself that you have put to words.

There's an episode of "Person of Interest" where the AI singularity struggles to play chess as each move cuts the possible moves and the first move is paralyzingly hard for it. Some people are like that, but regret is a form of hope without feathers. It comes from a strong belief that if you could wind the clock back and start over, it would be all different.

I nearly fell into the trap of Nihilism there ("Consider the Total Perspective Vortex")

[1] - https://notmysock.org/blog/me/a-buffet-intellectual

gopalv··on LibreOffice breaks download records after declaring it has no AI features
> Weekly downloads number doesn't.

The ChatGPT app downloads a copy of LibreOffice when you ask it to make pptx files.

So I've downloaded LibreOffice 3 times this week without really paying attention to it.

gopalv··on Project HydraFusion: Frontier quality via multi-model orchestration
> One model drafts a result, an independent read-only critic from a different model family reviews it

Multiple model vendors is key here, the cascade pattern doesn't need it, but the critique pattern does.

Last Nov, my team wrote a paper ("Team of Rivals") on the difference between using an OpenAI model to Critique an Anthropic model's output vs running a self-review agent loop on the same vendor.

The ablations [1] proved that neither company alone was better than using both.

The paper was a general response to "What does your company do that Anthropic can't?" but more so a demonstration of how to make something 90%+ good with models which eval at 60% or so (& Gas Town post unblocked our "this is a trade secret" argument about the paper).

[1] - https://github.com/t3rmin4t0r/critique-evals

gopalv··on We could save petabytes of cache storage with Zstandard and Pingora
> I don't see how that's possible if the cache no longer stores the uncompressed data.

Zstd has a seekable format for frames, similar to pigz --independent works.

[1] - https://github.com/facebook/zstd/blob/dev/contrib/seekable_f...

gopalv··on Memory Ordering in CPUs
The good part is that at least there's some explicit C++ std::memory_order and std::sync::atomic::Ordering lets me pick where it really matters.

For example, I built a skew handling model which needed low overhead cross-thread counters, where Ampere and Graviton was different from the M1 mac in benchmark - even down to the same assembly on different systems (cmov specifically).

gopalv··on You Probably Don't Get Why Stripe Bought OpenRouter
Assuming AI is a significant budget item for companies going forward, then this makes complete sense.

Someone's spending money and they can take a small % to move it from your account to an inference vendor's account. Except for this case you have to also ship tokens across.

If you see Ramp's counter point with launching router.com today, this move makes sense.

Because the next problem is cap budgets and block runaway spending.

gopalv··on Rethinking Database Programming
CTEs are how you compose SQL.

I don't quite like how the same CTE lives in 60 different places in my codebase, but at least the WITH clause changed things for me.

Also really liked Snowflake's result_scan for composing chains, mostly because I don't rerun expensive parts again and again. You can use ->> as a shortcut, but I don't think it uses results caching internally to skip waiting for them to all re-run & actually optimizes the whole thing.

gopalv··on California's new tire efficiency rules could save drivers $1B a year
> energy efficient tires have much longer wet braking distance

This is a labelling law, not a "You can't buy the bumpy tires" law.

This improves the quality of my choice here, because I already look for the little snowflake on tires when buying them for the 3 weekends I drive in the snow.

I'd be revolting too if they banned all season tires over this, because I am not swapping winter tires for thanksgiving weekend, Xmas and ski week.

Also efficient tires don't mean driving becomes boring, a GT86 is on skinnies and that makes it more fun at 35mph in a tight corner than a BMW with pilot sports.

I think they're only calling out cost, because the skinny tires cost more money overall, but there is a net payoff period on bills.

gopalv··on Models Are Getting Dumber on Purpose
> I want to click together a model that is laser-focused on what I am doing

This is roughly what multi-agent systems are built for.

This is possible with models too, but "making one on the fly" is much easier with agent coordination rather than model weights, since they all speak the same language.

There is an IBM Mainframe vs Google Distributed system division here. Like Seymour Cray said - two oxen or 1024 chickens.

Chickens are harder to harness, so a lot of my work is in sled-dog territory for agent harnesses & command structures.

gopalv··on People who grew up with high economic connectedness earn more
Your success eV is opportunities x success-rate.

Out of the gate, hard work improves success rates, but not the count of opportunities.

Connections improve opportunities, it doesn't do shit for success-rate.

Luck does for both sides of the mix.

As a startup founder in my 40s, the entire VP/SVP level in silicon valley is people one or two hops away for an intro, though admittedly a lot of the first hop there is through the other founders.

The original paper about importance of weak ties is under appreciated, because people you work with tend to go on to better & bigger things - especially in the beginning of your career, if you are a nice person.

Veritasium put out a very nice video[1] about knowing which sort of game we're playing. It is a power law game.

There are only two takeaways I had, "show up" and "get lucky". There is no winning without showing up.

[1] - https://www.youtube.com/watch?v=HBluLfX2F_k

gopalv··on 'Pervert glasses': Backlash against Meta's smart glasses grows
> notice the real world without the idea of augmenting it.

The core problem of solitude is curiosity and sonder.

When I smell fennel on a hike, I want to know what it is - yeah, this is dog fennel.

When I see a strange symbol on a taxi car, I want to know why it is 3 arrows on bow and what that is.

Seeing a strange dog and wondering what kind of dog that is (a whippet?). A glimpse of a black bird with red shoulder patches.

Hearing someone's ringtone (not these days anymore) and wonder why it is familiar ("ah, it is spelled fa9la, who'd have guessed").

There is a twinge of "This is why we can't have nice things" about the whole article for someone who can't get enough of the world you can see through your eyes.

gopalv··on How We Pushed CDC into Postgres
This was basically Vertica's party trick for quite a long time to have a WOS and ROS formats for the same row and anti-caching between those two.

You could've built a similar system with dezebium and delta lake for quite some time but it would fail compactions, if you run it fast enough. I've seen Oracle GoldenGate 12c do this trick in 2014 or so, using Mysql as the cheap replica. But they are all fragile to schema updates in some direction.

The closest batteries-included equivalent to this is the Aurora -> Redshift bridge[1].

[1] - https://aws.amazon.com/rds/aurora/zero-etl/

gopalv··on Discovery Loop
> a magnet for many more talented ML scientists to leave google.

Also to leave Meta, Amazon, Microsoft and everywhere else.

There would be more people who wouldn't join Google, but would love to do this instead.

gopalv··on Degrees of Wealth
> The Netherlands has made the intentional decision to reduce wealth inequality and to achieve a fairer society

Netherlands was the poster child of the "Dutch Disease"[1].

They in some ways could've turned into Qatar (and Norway into Aramco) without the right people twisting a few dials at the right time.

[1] - https://en.wikipedia.org/wiki/Dutch_disease

gopalv··on Not hiring junior engineers won't solve the problem you think you have
> What jobs are Juniors doing that AI can now do?

I used to use the new hires to hold expertise so that I can spend a weekend at the beach without getting paged. That was the true selfish incentive in spending negative productivity where it takes me more time to teach than to do.

AI absorbs knowledge that the company can keep when I leave.

This used to be a capacity and bandwidth issue, as well as a timing issue with spare time.

I couldn't spend 9 PM to 11 PM on an odd Thu doing this with a human.

gopalv··on Explanation of INT8 ConvRot (FP8 is no longer needed)
The other CloudFlare post on the front page has an interesting passage in it

> It is compute-bound, and INT4 weights have to be expanded back out before the model can multiply with them, so that extra step makes prefill slower rather than faster, GLM sustains about 10,160 tokens per second of prefill in FP8 versus 8,660 in INT4. As with the KV cache, the disaggregated design turns this into a choice rather than a compromise: we run INT4 for decode, where it wins, and FP8 for prefill, where it wins.

So each of these improvements are useful even if they have a narrow area of applicability, since the systems can be hybridized for performance.

[1] - https://blog.cloudflare.com/smaller-faster-safer-models/

gopalv··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
> It always felt as though we already figured out how to break up large files and parse them efficiently with very little memory.

The A in 26B-A4B is the active weights.

The problem is that this is a per-token load/unload at best, not for the whole prompt.

The division happened until one of these can fit in a single GPU and they stopped scaling it down any more, because you can wire up 8 of them to do their share of the work.

gopalv··on Are AI Labs Pelicanmaxxing?
> Catching a lab cheating specifically on my one dumb benchmark would be really funny.

Similar thing happened when TPC came up with SQL benchmarks.

If you're not good at TPC, your engineering team is no good.

If you're good at TPC, then (as a customer) we will actually include you in a bake-off benchmark for our specific problem.

Winning on it is the price of admittance into the game, especially in a crowded market.

But how narrowly you benchmarket matters, you can't just hard-code that specific scenario & not fix anything adjacent while you're at it.

For example when it comes to GPUs, the "Quack3" (sic) benchmark on ATI cards comes to mind.

gopalv··on How to pack ternary numbers in 8-bit bytes
> You can beat the efficiency of 5 trits in 8 bits

The single trit packing is a commonly optimized DBNULL structure for booleans.

Bits/trit approaches the 1.5 asymptote, because that is the fundamental packing limit.

The trick is to use it when you have a trit to start with, like when you have a set which is a tiny bit over a power of two.

There are places where you end up with odd numbers in set sizes, for example when storing a poker hand.

Read Cactus Kev's trick[1] which I think needs a 27 bit section & optimizing it was where I first ran into trit packing.

[1] - http://suffe.cool/poker/evaluator.html

gopalv··on Agent swarms and the new model economics
> does it matter whether we arrived there by random keystrokes?

Douglas Adams said this beautifully in the Campaign for Real Time & the poet Lallalfa.

And I just found out h2g2.com is no longer online, so no reference.

gopalv··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
Taalas HC1 is the closest thing to this.

Last I saw they posted Deepseek R1 numbers in Feb of this year.

The challenge is rolling out a new one every 7-8 weeks as the weights change & cheap enough for a hyper scaler to afford to buy one and save enough on power over the next 8 weeks as a payoff.

gopalv··on Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
> Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe?

This is starting to look at a lot like Intel vs Arm from the last era.

The Fable & Mythos are starting to look like a giant Xeon, while the smaller lighter models are starting to look like a lot of tiny ARM chips which sip on power instead.

The risk is the same as what Intel had. There is a group who are pushing them to go bigger and with a resource no limit approach, who have a lot of dollars to push you that way.

Follow them and they lead you to a pile of money, but then you risk something like Apple Silicon happening.

Something which got better because of efficiency & continuous improvement, not neutered due to it.

gopalv··on Old and new apps, via modern coding agents
I mostly use HTML, but it is much more flexible than what you would assume if you leverage some standard formats instead of building everything from scratch.

Mermaid, Graphviz and friends but in HTML pages.

Sometimes it is turning other things into perfetto.dev format for multi-machine tracking (like turn a build process into the same format as Chrome traces).

If you need more flexibility, you end up reaching for p5.js and three js (rather, tell the model to use it).

Once you're touching distance from WebGL, the equivalent of something you make can start looking like something from ciechanow.ski over a single weekend.

gopalv··on Weightlifting beats running for blood sugar control, researchers find (2025)
Outside of the (in mice) factor, the study compares optional exercise with mandatory exercise factors.

> To eat, the mice had to lift the lid while wearing a small shoulder collar, causing a squat-like movement that engaged the muscle contractions people use during resistance exercise.

vs

> For the endurance group, mice were given open access to a running wheel, an established model of aerobic exercise

The study is comparing the exercise that came in right before eating, which is effective at sugar control over the exercise done at any time as desired.

Speaking as a runner, I ignore the diet bump which makes me put on extra fat when I am training up for the SF (+2.5 kg over June & July is normal).

Mostly because I eat more the night before and mostly light carbs.

In fact, I'd bet my resting metabolism is actually slower when I'm training and the resting heart rates drop to 45 bpm & sleep takes up fewer calories too.

The muscle mass increase from lifting probably never cuts your metabolism needs when you are recovering or resting.

Cardiovascular fitness doesn't really cause weight loss when you're resting. So you'll be comparing something which reduces the calorie spend for the all the time you're not running vs something slightly bumps the spend when you are not lifting.

gopalv··on The future of Flipper Zero development
> I don't think I would use it enough to justify the investment

This is not a rational purchase - most of the rule breaking done with the zero is for fun or convenience, rather than being truly illegal.

It used to be more fun before the hotels started handing out NFC unlocks with your phone.

Still, being able to send each other a key for a hotel room on Signal is a nice trick if you are traveling with a sufficiently tech savvy group of people.

Page 1 of 25Next →