HNHacker News
TopNewBestAskShowJobs

recitedropper

313 karma · joined November 18, 2023

submissionscomments
recitedropper··on Gemini 3
Sure, but the extent to which you bend the truth to get those impressive numbers is absolutely gotcha-able.

Showing a new screen by default to everyone who is using your main product flow and then claiming that everyone who is seeing it is a priori a "user" is absurd. And that is the only way they can get to 2 billion a month, by my estimation.

They could put a new yellow rectangle at the top of all google search results and claim that the product launch has reached 2 billion monthly users and is one of the fastest-growing products of all time. Clearly absurd, and the same math as what they are saying here. I'm claiming my hottake gotcha :)

recitedropper··on Gemini 3
There could definitely be a chance. I was just responding to what in your comment sounded like a question.

That said, I think there is a good reason to be skeptical that it is a good chance. The consistent trend of finding higher complexity than expected in biological intelligences (like in C. Elegans), combined with the fact that the physical nature of digital architectures versus biological architectures are very different, is a good reason to bet on it being really complex to emulate with our current computing systems.

Obviously there is a way to do it physically--biological systems are physical after all--but we just don't understand enough to have the grounds to say it is "likely" doable digitally. Stuff like the Universal Approximation Theorem implies that in theory it may be possible, but that doesn't say anything about whether it is feasible. Same thing with Turing completeness too. All that these theorems say is our digital hardware can emulate anything that is a step-by-step process (computation), but not how challenging it is to emulate it or even that it is realistic to do so. It could turn out that something like human mind emulation is possible but it would take longer than the age of the universe to do it. Far simpler problems turn out to have similar issues (like calculating the optimal Go move without heuristics).

This is all to say that there could be plenty of smart ideas out there that break our current understandings in all sorts of ways. Which way the cards will land isn't really predictable, so all we can do is point to things that suggest skepticism, in one direction or another.

recitedropper··on Google Antigravity
Yes I recognize that, for various reasons, people will fail to document even when it is a profesional expectation.

I guess in this case we are comparing an idealized human to an idealized AI, given AI has equally its own failings in non-idealized scenarios (like hallucination).

recitedropper··on Google Antigravity
OK yes, you are right that we might be talking about employing AI toolings in different modes, and that the paper I am referring to is absolutely about agentic tooling executing code changes on your behalf.

That said, the first comment of the person I replied to contained: "You can ask agents to identify and remove cruft", which is pretty explicitly speaking to agent mode. He was also responding to a comment that was talking about how humans spend "hours talking about architectural decisions", which as an action mapped to AI would be more plan mode than ask mode.

Overall I definitely agree that using LLM tools to just tell you things about the structure of a codebase are a great way to use them, and that they are generally better at those one-off tasks than things that involve substantial multi-step communications in the ways humans often do.

I appreciate being the weeds here haha--hopefully we all got a little better talking abou the nuances of these things :)

recitedropper··on Google Antigravity
For seasoned maintainers of open source repos, there is explicit evidence it does slow them down, even when they think it sped them up: https://arxiv.org/abs/2507.09089

Cue: "the tools are so much better now", "the people in the study didn't know how to use Cursor", etc. Regardless if one takes issue with this study, there are enough others of its kind to suggest skepticism regarding how much these tools really create speed benefits when employed at scale. The maintenance cliff is always nigh...

There are definitely ways in which LLMs, and agentic coding tools scaffolded in top, help with aspects of development. But to say anyone who claims otherwise is either being disingenuous or doesn't know what they are doing, is not an informed take.

recitedropper··on Google Antigravity
Alright, I'm glad to hear you've had a successful and rich professional career. We definitely agree that engineers generally fail to document when they have competing priorities, and that LLMs can be of use to help offload some of that work successfully.

Your initial comment made it sound like you were commenting on a genuine apples-for-apples comparisons between humans and LLMs, in a controlled setting. That's the place for empiricism, and I think dismissing studies examining such situations is a mistake.

A good warning flag for why that is a mistake is the recent article that showed engineers estimated LLMs sped them up by like 24%, but when measured they were actually slower by 17%. One should always examine whether or not the specifics of the study really applies to them--there is no "end all be all" in empiricism--but when in doubt the scientific method is our primary tool for determining what is actually going on.

But we can just vibe it lol. Fwiw, the parent comment's claims line up more with my experience than yours. Leave an agent running for "hours" (as specified in the comment) coming up with architectural choices, ask it to document all of it, and then come back and see it is a massive mess. I have yet to have a colleague do that, without reaching out and saying "help I'm out of my depth".

recitedropper··on Google Antigravity
"major architectural decisions don't get documented anywhere" "commit messages give no "why""

This is so far outside of common industry practices that I don't think your sentiment generalizes. Or perhaps your expectation of what should go in a single commit message is different from the rest of us...

LLMs, especially those with reasoning chains, are notoriously bad at explaining their thought process. This isn't vibes, it is empiricism: https://arxiv.org/abs/2305.04388

If you are genuinely working somewhere where the people around you are worse than LLMs at explaining and documenting their thought process, I would looking elsewhere. Can't imagine that is good for one's own development (or sanity).

recitedropper··on Gemini 3
I'm sure each of the frontier labs have some secret methods, especially in training the models and the engineering of optimizing inference. That said, I don't think them saying they'd keep a big breakthrough secret would be evidence in this case of a "secret sauce" on ARC-AGI-2.

If they had found something fundamentally new, I doubt they would've snuck it into Gemini 3. Probably would cook on it longer and release something truly mindblowing. Or, you know, just take over the world with their new omniscient ASI :)

recitedropper··on Gemini 3
Astroturfing used as evidence of domination. Public forums truly have come full circle.
recitedropper··on Gemini 3
I think in this case, tokenization and percpetion are somewhat analogous. I think it is probably the case our current tokenization schemes are really simplistic compared to what nature is working with. If you allow the analogy.
recitedropper··on Gemini 3
Who wants to bet they benchmaxxed ARC-AGI-2? Nothing in their release implies they found some sort of "secret sauce" that justifies the jump.

Maybe they are keeping that itself secret, but more likely they probably just have had humans generate an enormous number of examples, and then synthetically build on that.

No benchmark is safe, when this much money is on the line.

recitedropper··on Google Brings Gemini 3 AI Model to Search and AI Mode
That's a good point, although given I'd never seen this rule I question if it is commonly known enough that it is actually the reason I'm being downvoted.

Do you not think what has happened today is suspicious? The Gemini 3 posts are, to my eye, out of hand..

recitedropper··on Gemini 3
Not sure if this is agreeing or disagreeing with there being astroturfing.

But I'd reckon that the negative sentiments at the top, combined with that there are over eight Gemini 3 posts on the front page recently, is good evidence of manipulation. This actually might be the most posted about model release this year, and if people were that excited we wouldn't have negative sentiment abound.

recitedropper··on Gemini 3
This is the million dollar question. I'm not qualified to answer it, and I don't really think anyone out there has the answer yet.

My armchair take would be that watt usage probably isn't a good proxy for computational complexity in biological systems. A good piece of evidence for this is from the C. elegans research that has found that the configuration of ions within a neuron--not just the electrical charge on the membrane--record computationally-relevant information about a stimulus. There are probably many more hacks like this that allow the brain to handle enormous complexity without it showing up in our measurements of its power consumption.

recitedropper··on Google Brings Gemini 3 AI Model to Search and AI Mode
The amount of capital that rests on releases like these is insane. The incentive is just too high to not manipulate places like HN, which have a surprising amount of sway with tech industry sentiment.

Edit: Check out how my claim against astroturfing did, one of the first comments posted on the primary release blog: https://news.ycombinator.com/item?id=45967999#45968295. Could always be over-eager Google employees, or maybe the tech community really is this excited for the Gemini 3 release. Seems fishy to me, though...

recitedropper··on Gemini 3
Perception seems to be one of the main constraints on LLMs that not much progress has been made on. Perhaps not surprising, given perception is something evolution has worked on since the inception of life itself. Likely much, much more expensive computationally than it receives credit for.
recitedropper··on 5 Things to Try with Gemini 3 Pro in Gemini CLI
Nice, without this thread I would never have known Gemini 3 released today.

Going to download Gemini CLI right now™ and see how it performs™ against Cursor, Claude Code, Aider, OpenCode, Droid, Warp, Devin, and ForgeCode.

recitedropper··on Gemini 3
I'm pretty sure they mention in their various TOSes that they don't train on user data in places like Gmail.

That said, LLMs are the most data-greedy technology of all time, and it wouldn't surprise me that companies building them feel so much pressure to top each other they "sidestep" their own TOSes. There are plenty of signals they are already changing their terms to train when previously they said they wouldn't--see Anthropic's update in August regarding Claude Code.

If anyone ever starts caring about privacy again, this might be a way to bring down the crazy AI capex / tech valuations. It is probably possible, if you are a sufficiently funded and motivated actor, to tease out evidence of training data that shouldn't be there based on a vendor's TOS. There is already evidence some IP owners (like NYT) have done this for copyright claims, but you could get a lot more pitchforks out if it turns out Jane Doe's HIPAA-protected information in an email was trained on.

recitedropper··on Gemini 3
They claim AI overviews as having "2 billion users" in the sentences prior. They are clearly trying as hard as possible to show the "best" numbers.
recitedropper··on Gemini 3
Perhaps I shouldn't have implied an expectation of lots of explicit mentions of "AGI". It is more the general sentiments being expressed, and the extent to which critical takes seem to be quickly buried.

I'm totally open to being wrong though. Maybe the tech community is just that excited about Gemini 3's release.

recitedropper··on Gemini 3
I definitely believe it--I'm not a total AI hater. The jump on the screen usage benchmark is really exciting in that it might substantially help computer-use agentic workflows.

That said, I think there is too much a pattern with recent model releases around what appears to me to be astroturfing to get to HN front page. Of course that doesn't preclude many organic comments that are excited too!

A bit of both always happens. But given how important these model releases are to justify the capex and levels of investment, I think it is pretty clear the various "front pages" of our internet are manipulated. The incentive is just too strong not to.

recitedropper··on Gemini 3
"Since then, it’s been incredible to see how much people love it. AI Overviews now have 2 billion users every month."

Cringe. To get to 2 billion a month they must be counting anyone who sees an AI overview as a user. They should just go ahead and claim the "most quickly adopted product in history" as well.

recitedropper··on Gemini 3
Peek the other threads.
recitedropper··on Gemini 3
I'm primarily reacting to the other threads, like the one that leaked the system card early. And, perhaps unfairly, Twitter as well.
recitedropper··on Gemini 3
Inevitable... certainly more so than AGI :)
recitedropper··on The Dragon Hatchling: The missing link between the transformer and brain models
For the most part I think we agree. There is a lot of uncertainty around the mechanics of consciousness, a lot of reasons to doubt the existence of those mechanics in current AI, and a lot of failed endeavors to use biological mimicry to improve AI state of the art.

I don't think that precludes remaining concerned with the continued push to make current models more humanlike in nature. My initial comment was spurred by the fact that this paper is literally presenting itself as solving the missing link between transformer architectures and the human brain.

Here's to hoping this all goes toward a better world.

recitedropper··on The Dragon Hatchling: The missing link between the transformer and brain models
I hope your generous interpretation is right... I can't really tell what's going on with Anthropic's theater either. They definitely seem like they are vigilant of bad outcomes, going as far as to publish their own economic index trying to monitor how AI is affecting labor markets.

That said, the cynic in me thinks they give lip service to these things while pushing fully ahead into the unknown on the presumption of glory and a possibility of abundance. A bunch of the leadership are EAs who subscribe to a kind of superintelligence eschatology that goes as far as to give a shot at their own immortality. Given that, I think they act on the assumption that AGI is a necessity, and they'd rather take the risks on everyone's behalf than just not create the technology in the first place.

Them recently flirting with money from the gulf states is a pretty concerning signal pointing to them being more concerned with their own goals rather than ethics.

recitedropper··on The Dragon Hatchling: The missing link between the transformer and brain models
I am empathetic to arguments against consciounsess being computational. Definitely strange to imagine an algorithm played out on trillions of abacuses being conscious.

That said, I don't think it is a sufficient appeal to entirely discount the possibility that the right process implemented on silicon could in fact be conscious in the same way we are. I'm open to whether or not it is possible--I don't have a vested interest in the space--but silica seems to be a medium that can possible hold the level of complexity for something like consciousness to emerge.

So this is to say that I agree with you that consciousness likely requires substrate-specific embodiment, but I'm open to silica being a possible substrate. I certainly don't think it can be discounted at this point in time, and I'd suggest that we don't risk a digital holocaust on the bet that it can't.

recitedropper··on The Dragon Hatchling: The missing link between the transformer and brain models
In this thought experiment, I am considering artificial life genuine. I would agree that there could be productive outlets for our selfish impulses if there was something that mimicked their targets without consciousness to experience the externalities of such impulses.

That said, I think probably the best path would just be to build and foster technologies that help our species mature, so if one day we do get the ability to spin-up conscious beings artificially, it can be done in a manner that adds more beauty rather than despair to our universe.

recitedropper··on The Dragon Hatchling: The missing link between the transformer and brain models
For the record, I'm agnostic to whether or not consciousnses is possible upon silica. I think it is pretty safe to say though that it likely is an emergent property of specifically-configured complex systems, and humanlike intelligence on silica is certainly something that might qualify.

I don't think appealing to whether or not inanimate objects may be conscious is sufficient to discount that we are toying with a different beast in machine learning. And, if we were to discover that inanimate objects are in-fact conscious, that would be an even greater reason to reconfigure our society and world around compassion.

I agree that LLMs are a great breakthrough, and I think there are many reasons to doubt consciousness there. But I would suggest we rest on our laurels for a bit, and see what we can get out of LLMs, rather than push to create something that is closer to mimicking humans because it might be more useful. From the evil perspective of pure utility, slaves are quite useful as well.

← PreviousPage 3 of 4Next →