HNHacker News
TopNewBestAskShowJobs

halflings

2,566 karma · joined September 4, 2013

Software Engineer specialized in applied Machine Learning, and in particular ranking and search quality.

Get in touch: https://kachkach.com/ ahmed@kachkach.com

submissionscomments
halflings··on Men are ditching TV for YouTube as AI usage and social media fatigue grow
> Youtube charges $10 per month and doesn't produce a single video

It is different from Netflix (that pays upfront for production costs), but there's of course a revenue share + the bulk of the revenue for creators is actually from sponsorships (which YT doesn't take a share of).

halflings··on Generating one token at a time is a blessing in disguise
LLMs generate their output one token at a time. The first thought when you learn this is that this is a huge performance bottleneck, as we are used to highly parallelized systems.

However, a large part of what makes LLMs feel so magical comes from this bottleneck.

halflings··on The Codex App
The main thing I noticed in the video is that they have heavily sped up all the code generation sections... seems to be on 5x speed or more. (because people got used to how fast and good Sonnet, and especially Gemini 3.0 Flash, are)
halflings··on The Codex App
Deploying from Antigravity is as easy as say connecting the Firebase MCP [1] and asking it "deploy my app to firebase".

[1] https://firebase.google.com/docs/ai-assistance/mcp-server

halflings··on Python 3.15’s interpreter for Windows x86-64 should hopefully be 15% faster
+1, reading through the post, the PR updating the documentation... thanks for being transparent, but also don't be so hard on yourself!

That was a very niche error, that you promptly corrected, no need to be so apologetic about it! And thanks for all the hard work making Python faster!

halflings··on Getting a Gemini API key is an exercise in frustration
"The models perform differently when called via the API vs in the Gemini UI."

This shouldn't be surprised, e.g. the model != the product. The same way GPT4o behaves differently than the ChatGPT product when using GPT4o.

halflings··on US vs. Google amicus curiae brief of Y Combinator in support of plaintiffs [pdf]
I would also add that search has already moved elsewhere.

Less and less people are using search engines to shop, ex:Amazon makes >$57B a year from search ads, but also look at Temu and Shein which are mostly glorified product search platforms.

No one is searching for "funny videos" when you can just open Instagram and Tiktok.

The only real unique thing that search engines can do is queries that are not directly commercial (e.g. education, information seeking, etc.) and competition is insanely intense (w/ ChatGPT, Perplexity, etc) there.

halflings··on Gemma 3 QAT Models: Bringing AI to Consumer GPUs
That's what the chart says yes. 14.1GB VRAM usage for the 27B model.
halflings··on Genie 2: A large-scale foundation world model
> I cannot be the first person to think about such possibilities

Differentiable Rendering [1] is the closest thing to what you are describing. And yes, people have been working on this for the same reason you outline, it is more data/compute efficient and hence should generalize better.

[1] https://blog.qarnot.com/article/an-overview-of-differentiabl...

But also: > While cool, this also seems utterly wasteful. Video games offer known "analytical" solutions for the interactions that the model provides as a "statistical approximation", so to say.

A bit of the same debate as people calling LLMs a "blurry JPEG of the web" and hence useless.

Yes this is a statistical approximation to an analytical problem... but that's a very reductive framing to what is going on. To find the symbolic/analytical solution here would require to constrain the problem greatly: not all things on the screen have a differentiable representation, for example complex simulations might involve some kind of custom internal loop/simulation.

You waste compute to get a solution that can just be trained on billions of unlabeled (synthetic) examples, and then generalize to previously unseen prompts/environments.

halflings··on Open source AI is the path forward
Training code is only useful to people in academia, and the closest thing to "code you can modify" are open weights.

People are framing this as if it was an open-source hierarchy, with "actual" open-source requiring all training code to be shared. This is not obvious to me, as I'm not asking people that share open-source libraries to also share the tools they used to develop them. I'm also not asking them to share all the design documents/architecture discussion behind this software. It's sufficient that I can take the end result and reshape it in any way I desire.

This is coming from an LLM practitioner that finetunes models for a living; and this constant debate about open-source vs open-weights seems like a huge distraction vs the impact open-sourcing something like Llama has... this is truly a Linux-like moment. (at a much smaller scale of course, for now at least)

halflings··on Suicide is on the rise for young Americans, with no clear answers
> The world is teetering on the edge of world war

The world probably has never been as peaceful as in the last 50 years or so.

Same goes for access to drinkable water, food, decent shelter, gender equality, freedoms, technology, etc.

But I suppose your comment is a good illustration of the problem at hands (that so many people deeply believe that things are fucked)

halflings··on Groq CEO: 'We No Longer Sell Hardware'
Thanks for putting this together! Will give it a watch now
halflings··on Groq CEO: 'We No Longer Sell Hardware'
The # of chips is not the most important metric.

Most important, even ignoring latency, is throughput (tokens) per $$$. And according to their own benchmark [1] (famous last words :)), they're quite cost efficient.

[1] https://www.semianalysis.com/p/groq-inference-tokenomics-spe...

halflings··on Groq CEO: 'We No Longer Sell Hardware'
No HBM because they use tons of fast SRAM instead. Isn't that the main driver for performance here?

(the way I understood it => it's still cost effective at scale due to throughput increase this brings)

halflings··on Google's First Tensor Processing Unit: Architecture
Agree re:hallucinations/safety issues, that was likely one of the main blockers.

And here's the sad part: they had this back in 2019... see this paper released in Jan 2020: https://blog.research.google/2020/01/towards-conversational-...

halflings··on Google's First Tensor Processing Unit: Architecture
This (innovator's dilemma / too afraid of disrupting your own ads business model) is the most common explanation folks are giving for this, but seems to be some sort of post-rationalization of why such a large company full of competent researchers/engineers would drop the ball this hard.

My read (having seen some of this on the inside), is that it was a mix of being too worried about safety issues (OMG, the chatbot occasionally says something offensive!) and being too complacent (too comfortable with incremental changes in Search, no appetite for launching an entirely new type of product / doing something really out there). There are many ways to monetize a chatbot, OpenAI for example is raking billions in subscription fees.

halflings··on Pyenv – lets you easily switch between multiple versions of Python
uv has been really awesome as a replacement for pip: https://github.com/astral-sh/uv

So fast it finally made virtual environments usable for me. But it's not (yet) a full replacement for conda, e.g. it won't install things outside of Python packages

halflings··on Show HN: Matrix Multiplication with Half the Multiplications
This looks pretty cool! What's the catch? e.g. why isn't this already implemented in accelerators, is it really just a forgotten algorithm, or this has some implications on the cost of building the accelerator or else?
halflings··on Among the A.I. doomsayers
> Your analogy is the same as early Intel engineers completely unaware that those chips would bring on the ramifications of social media

Exactly! As they should be. (for both Intel engineers developing chips, and physicists developing nuclear research)

There were a billion more potential dangers from those technologies that never materialized, and never will.

I'm glad we didn't stop them in their track because a poll of 10 leaders in the field thought they were too dangerous and progress should stop. (note that no one is against regulating dangerous uses of AI, e.g. autonomous weapons, chemical warfare; the problem is regulating AI research and development in the first place)

halflings··on Among the A.I. doomsayers
This is the definition of strawman. "Advocate for killing all humans" sounds like someone advocating for a genocide, but instead it's just the same transhumanist thinking (which Yudkowsky also believes in, FYI)
halflings··on Among the A.I. doomsayers
Your comment (+ username) reads like what I would have written once upon a time when I was fully in the EA bubble.

Truly no offense meant, as I was deeply into the EA movement myself, and still consider myself one (in the original "donate money effectively" sense), but the movement has now morphed into a death cult obsessed by things like:

* OMG we're all going to die any time now (repeated every year since circa 2018)

* What is your pDoom? What are your timelines? (aka: what is your totally made up number that makes you feel like you're doing something rational/scientific)

I'm deep in the weeds w/ LLMs, e.g. I probably finetune an average of 1 model a day, and working with bleeding edge models... and AI safety just sounds so silly. Wanting to take drastic measures today to prevent an upcoming apocalypse makes as much sense as taking the same drastic measures when gradient descent was invented.

halflings··on Mamba Explained: The State Space Model Taking On Transformers
The importance (e.g. attention) needs to be dynamic, e.g. one token will be important to some other tokens but not others.

tf-idf and similar heuristics are what we were using before attention came along, e.g. tf-idf weighted bag-of-words representation of word2vec embeddings. That approaches fails in so many cases.

halflings··on Our next-generation model: Gemini 1.5
Yep that's pretty much it! That's what they call needle in a haystack. See: https://github.com/gkamradt/LLMTest_NeedleInAHaystack
halflings··on Mozilla 2023 annual report: CEO pay skyrockets, Firefox market share nosedives
Not if you account for the increase in ad revenue from the company paying them.
halflings··on Airbnb is deploying AI to block New Year's Eve bookings that could be parties
> The technology looks at hundreds of signals that could indicate a booking is higher risk for this type of incident, like the duration of the trip the guest is trying to book, how far the listing is from their location, the type of listing they’re booking, and if the reservation is being made at the last-minute, among many more.

This is most likely a (simple) ML model trained on previous reports of such bookings.

Not newsworthy, but probably not a bunch of if-statements.

halflings··on PeerTube v6
> Nowhere it's touted as a rival to youtube in the popularity sense, just like you wouldn't call WordPress a rival to Twitter.

... the first image on the article linked here shows a monster called "Videorapter" with YouTube, Vimeo and Twitch logos, and calls for donations to "help push back Videoraptor"?

halflings··on Google's advanced music generation model and two new AI experiments
> Why not output MIDI instead, and let the artist manipulate that?

This is not true at all, it all started with generating MIDI, and even the very link I shared is of a system that generates MIDI.

halflings··on Fast Llama 2 on CPUs with Sparse Fine-Tuning and DeepSparse
The page fully explains what they mean by this, showing results on benchmarks etc.
halflings··on The curious case of the abominable shower
Anyone from Scandinavia at least would be horrified by this story. It makes the US presidency look like a 3rd world dictatorship.

If it's OK for a president to behave in this way, then you better believe this goes all the way down (and every person is abusing public funds to various extents).

halflings··on Google's advanced music generation model and two new AI experiments
Growing pains. Remember how this all started (Google Magenta, circa 2019): https://www.youtube.com/watch?v=ZRnbbtqxBEc
Page 1 of 21Next →