HNHacker News
TopNewBestAskShowJobs

gandalfgeek

3,101 karma · joined November 4, 2008

meet.hn/city/33.6856969,-117.825981/Irvine
submissionscomments
gandalfgeek··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
Wow. Jeff and Sanjay both departing. Truly end of a golden era.

There is an entire cohort of work-optional very senior engineers for whom one of the last reasons for hanging on was "at least Jeff and Sanjay are around".

gandalfgeek··on AI financial advice is surprisingly good, especially if you ask right questions
"2. AI misses important nuances. Better prompts could help."

My experience has been very different: I give it a ton of personal context (positions, portfolio, account balances etc). I find it's advice to be exceptional, even on advanced topics (tax planning, asset location, long-term planning and scenario testing).

None of the professionals I've engaged or consider engaging (2-3 orders of magnitude more expensive than annual cost of Pro/Max subscriptions) come close.

In fact, it (both Opus 4.8 and GPT-5.5) found a tax overpayment issue my tax guy missed. I basically read out what Codex told me to the pro on the phone to get him to understand and acknowledge the issue. Paid for the annual subscription right there.

gandalfgeek··on Kimi Work
wow -- "Let's take something off your plate" just copied word-for-word from the Claude Cowork UI.

And it is above the fold in their hero image!

gandalfgeek··on Schema Harness Achieves ~99% on Arc‑AGI‑3 Public
Big jump for sure, but definitely comes with a giant grain of salt lacking open-sourcing the harness itself and measuring performance on the held-out set.
gandalfgeek··on Mag 7 starting to underperform [pdf]
6 months ago when Mag7 was overperforming everyone was worried about it being too high a fraction of the S&P500.
gandalfgeek··on Google employees internally share memes about how its AI sucks
(ex-Googler, spent 18 yrs there)

Memegen is a key part of the culture. Its default mode is over-the-top mocking, of course, with a grain of truth. Nobody and nothing is spared. C-level execs, products, the perf process.

So this by itself is not quite the scoop 404 media thinks it is. You could take the front page of memegen on any given day and construct twenty scandalous headlines of it.

gandalfgeek··on Nvidia just paid $20B for a company that missed its revenue target by 75%
> About a year ago, Groq announced a $1.5 billion infrastructure investment deal with Saudi Arabia. They also secured a $750 million Series D funding round.... Then in maybe one of the best rug pulls of all time, in July they quietly changed their revenue projections to $500 million. A 75% cut in four months. I’ve never seen anything like that since the 2008 financial crisis.

Not following the core argument here. Author seems to be comparing valuation in funding rounds to revenue projections. Revenue projection was revised downward, valuation was not.

Good point about not running the proprietary models, but that doesn't preclude strategic fit with Nvidia.

gandalfgeek··on The Undermining of the CDC
> Government agencies in general are largely insulated from politics.

This was obviously false during the pandemic when these “health” agencies did what the White House wanted, from the actual “science” to the messaging.

gandalfgeek··on Warp Code: the fastest way from prompt to production
> Initialize projects with their own WARP.md files (compatible with Agents.MD, Claude.MD and cursor rules).

Can we please standardize this and just have one markdown file that all the agents can use?

gandalfgeek··on MIT Study Finds AI Use Reprograms the Brain, Leading to Cognitive Decline
The coverage of this has been so bad that the authors have had to put up an FAQ[1] on their website, where the first question is the following:

Is it safe to say that LLMs are, in essence, making us "dumber"? No! Please do not use the words like “stupid”, “dumb”, “brain rot”, "harm", "damage", "brain damage", "passivity", "trimming" , "collapse" and so on. It does a huge disservice to this work, as we did not use this vocabulary in the paper, especially if you are a journalist reporting on it.

[1]: https://www.media.mit.edu/projects/your-brain-on-chatgpt/ove...

gandalfgeek··on MIT Study Finds AI Use Reprograms the Brain, Leading to Cognitive Decline
The title of the study is provocatively framed and the actual findings don't live up to it. I made a short video explaining it-- https://www.youtube.com/watch?v=hLDCi0VwyiQ
gandalfgeek··on Wasting Inferences with Aider
Very cool. Even cooler to see it upload itself!!
gandalfgeek··on Wasting Inferences with Aider
There is no fundamental blocker to agents doing all those things. Mostly a matter of constructing the right tools and grounding, which can be fair amount of up-front work. Arming LLMs with the right tools and documentation got us this far. There’s no reason to believe that path is exhausted.
gandalfgeek··on Heavy chatbot usage is correlated with loneliness and reduced socialization
Wet streets cause rain.
gandalfgeek··on Stay Gold, America
Most charities on that list are on sharp sides of deeply polarizing culture war issues and it is not at all clear that their causes align with the “American Dream”.

If the last election was any indication then more than half the country explicitly rejected many of them.

gandalfgeek··on AI Predictions for 2025, from Gary Marcus
I don't understand why he has such an axe to grind. Is there some historical baggage here?

Of course there are plenty of problems with the current state of AI and LLMs, but to have such a preconceived pessimistic outlook that can't even acknowledge their massive and quick adoption and usefulness in multiple domains seems not intellectually honest.

gandalfgeek··on Fogus: Things and Stuff of 2024
Love Fogus, but zero mention of how LLMs impacted programming in 2024?
gandalfgeek··on How Google spent 15 years creating a culture of concealment
This is a BS story.

Pretty much every public company, at least every bigtech company, follows the same conventions -- don't say incriminating things in chat, trainings for "communicate with care" (definitely don't say "we will kill the competition!!" in email or chat), automatic retention policy etc etc.

No need to single out Google.

gandalfgeek··on Scientific American's departing editor and the politicization of science
There was one slogan that was repeated during COVID that perfectly encapsulates the degeneration and capture of science: "Follow the science".

That's not how science works. Religions are "followed". Science is based on questioning and skepticism and falsifiability.

gandalfgeek··on Alan Kay on Messaging (1998)
Thanks for the pointer!

"Call by meaning" sounds exactly like LLMs with tool-calling. The LLM is the component that has "common-sense understanding" of which tool to invoke when, based purely on natural language understanding of each tool's description and signature.

gandalfgeek··on The Friendship that made Google huge (2018)
(former Googler)

It was really special to see how this pair basically laid out the foundations of large-scale distributed computing. Protobufs, huge parts of the search stack, GFS, MapReduce, BigTable... the list goes on.

They are the only two people at Google at level 11 (senior fellow) on a scale that goes from 3 (fresh grad) to 10 (fellow).

gandalfgeek··on DJI ban passes the House and moves on to the Senate
Keep seeing DJI drones at local police dept open houses. They even have a "drone unit" that specializes in SAR, hazardous recon type scenarios. Given extensive existing use throughout US local law enforcement, fire depts etc, not sure if this will actually happen.

Or maybe they all mass-migrate to Anduril solutions?

gandalfgeek··on What we've learned from a year of building with LLMs
Agree that your use-case is different. The papers above are dealing mostly with adding a domain-specific textual corpus, still answering questions in prose.

"Teaching" the LLM an entirely new language (like a DSL) might actually need fine-tuning, but you can probably build a pretty decent first-cut of your system with n-shot prompts, then fine-tune to get the accuracy higher.

gandalfgeek··on What we've learned from a year of building with LLMs
This was kind of conventional wisdom ("fine tune only when absolutely necessary for your domain", "fine-tuning hurts factuality"), but some recent research (some of which they cite) has actually quantitatively shown that RAG is much preferable to FT for adding domain-specific knowledge to an LLM:

- "Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?" https://arxiv.org/abs//2405.05904

- "Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs" https://arxiv.org/abs/2312.05934

gandalfgeek··on Ask HN: Is RAG the Future of LLMs?
#1 motivation for RAG: you want to use the LLM to provide answers about a specific domain. You want to not depend on the LLM's "world knowledge" (what was in its training data), either because your domain knowledge is in a private corpus, or because your domain's knowledge has shifted since the LLM was trained.

The latest connotation of RAG includes mixing in real-time data from tools or RPC calls. E.g. getting data specific to the user issuing the query (their orders, history etc) and adding that to the context.

So will very large context windows (1M tokens!) "kill RAG"?

- at the simple end of the app complexity spectrum: when you're spinning up a prototype or your "corpus" is not very large, yes-- you can skip the complexity of RAG and just dump everything into the window.

- but there are always more complex use-cases that will want to shape the answer by limiting what they put into the context window.

- cost-- filling up a significant fraction of a 1M window is expensive, both in terms of money and latency. So at scale, you'll want to filter out and RAG relevant info rather than indiscriminately dump everything into the window.

gandalfgeek··on Groq CEO: 'We No Longer Sell Hardware'
> No HBM because they use tons of fast SRAM instead. Isn't that the main driver for performance here?

No doubt fast SRAM helps, but from a computation pov imho its that they've statically planned computation and eliminated all locks.

Short explainer here: https://www.youtube.com/watch?v=H77tV1KcWIE (Based on their paper).

gandalfgeek··on Groq CEO: 'We No Longer Sell Hardware'
They're calling the lie on needing bleeding edge hardware for performance.

5 yr old silicon (14 nm!!) and no hbm.

Their secret sauce seems to be an ahead-of-time compiler that statically lays out entire computation, enabling zero contention at runtime. Basically, they stamp out all non-determinism.

https://wow.groq.com/isca-2022-paper

gandalfgeek··on Covert Racism in LLMs
The cited paper seems to be really bending over backwards to find some trace of bias. Unwinnable game for LLMs.

E.g. cited work claims "LLMs assign significantly less prestigious jobs to speakers of African American English... compared to Standardized American English". You don't say! Formal/business language has higher association with prestigious jobs than informal/street/urban language. How is that even classified as "bias"?

gandalfgeek··on 401(k) Will Be Gone Within a Decade
Money I earn is taxed.

I invest what remains, gains are taxed.

I buy a house with what remains after that, I owe property tax.

I buy things to live with what remains, I pay sales tax.

To keep up the pretense of caring the state occasionally throws a bone like "tax advantaged" accounts. It either defers taxes or takes after tax money. Oxymoron to start with.

The state is a parasite that only knows how to sink it's fangs deeper into you.

Afuera!

gandalfgeek··on Jeff Dean: Trends in Machine Learning [video]
Appreciate the feedback. This was a weekend project.

I wanted to explore summarization without paraphrasing, so that the output was a cut up version of the input video. Agreed that it terms of conceptual clarity often a textual summary that is synthesized comes out ahead.

Page 1 of 5Next →