HNHacker News
TopNewBestAskShowJobs

somebodythere

1,000 karma · joined July 23, 2018

submissionscomments
somebodythere··on The turbulent AI era is here
Did you read the article?
somebodythere··on Nvidia agrees to acquire Hugging Face for $13B
Not with the help of the models!
somebodythere··on AI 2040: Plan A
This prediction can't be scored until the 2028 election cycle.

You may think it's very unlikely the prediction will have turned out to be correct by the 2028 election cycle, but that is not the same thing as the prediction being scorable as false today.

somebodythere··on AI 2040: Plan A
The thing about exponentials is if you admit 60%, it's pretty easy to admit 95%.
somebodythere··on Replies to comments on my "LLMs are eroding my career" post
That is not true. You can tell you are on the latter part of the S-Curve you are on, if the rate of change of capabilities has decreased compared to before. That is not what we are seeing right now. The rate of change is increasing, or is at best, stable.
somebodythere··on Sauna effect on heart rate
There is some evidence suggesting that "blue zones" are largely about pension fraud. https://fortune.com/europe/2024/12/14/are-blue-zones-myth-ex...
somebodythere··on Anthropic takes legal action against OpenCode
Using your API key in third-party harnesses has always been allowed. They just don't like using the subsidized subscription plan outside of first-party harnesses. So this seems to be out of spite
somebodythere··on We tasked Opus 4.6 using agent teams to build a C Compiler
even a squirrel that needs guidance from a human grandmaster, is heavily inspired by existing games, and who can use Piece Mover library is incredible. 5 years ago the squirrel was just a squirrel. then it was able to make legal moves. now it can play a whole game from start to finish, with help. that is incredible
somebodythere··on AI2: Open Coding Agents
No. You can point e.g. Opencode/Cline/Roo Code/Kilo Code at your inference endpoint. But CC has high install base and users are used to it, so it makes sense to target it.
somebodythere··on The future of software development is software developers
Why would I ask the model to reverse the string 'glorbix,' especially in the context of software engineering?
somebodythere··on Mozilla appoints new CEO Anthony Enzor-Demeo
You personally wouldn't use live captions and dubbing, so there's no point building it for the millions of people who need it as an accessibility feature?
somebodythere··on Anthropic taps IPO lawyers as it races OpenAI to go public
Rufus is a Claude Haiku, yes.
somebodythere··on Code Wiki: Accelerating your code understanding
I've seen a few of this type of thing pop up in search results ("DeepWiki" by Cognition.) I'm not a fan. It is just LLM contentslop, basically. Actual wikis written by humans are made of actual insight from developers and consumers. "We intend you use it in X way", "If you encounter Y issue, do Z." etc. Look at arch wiki. Peak wiki-style documentation, LLMs could never recreate. Well, maybe with a future iteration of the technology they can be useful. But for now, you do not gain much by essentially restating code, API interfaces, and tests in prose. They take up space from legitimate documentation and developer instruction in search results.
somebodythere··on Anthropic acquires Bun
I think this wound up being close enough to true, it's just that it actually says less than what people assumed at the time.

It's basically the Jevons paradox for code. The price of lines of code (in human engineer-hours) has decreased a lot, so there is a bunch of code that is now economically justifiable which wouldn't have been written before. For example, I can prompt several ad-hoc benchmarking scripts in 1-2 minutes to troubleshoot an issue which might have taken 10-20 minutes each by myself, allowing me to investigate many performance angles. Not everything gets committed to source control.

Put another way, at least in my workflow and at my workplace, the volume of code has increased, and most of that increase comes from new code that would not have been written if not for AI, and a smaller portion is code that I would have written before AI but now let the AI write so I can focus on harder tasks. Of course, it's uneven penetration, AI helps more with tasks that are well-described in the training set (webapps, data science, Linux admin...) compared to e.g. issues arising from quirky internal architecture, Rust, etc.

somebodythere··on Many countries that said no to ChatControl in 2024 are now undecided
Roughly, this is the Electronic Frontier Foundation (and comparable lobbying orgs in other countries.) However, an org like this doesn't have much power to compel individuals to give them $1.
somebodythere··on Learning Is Slower Than You Think
LLM argumentative essays tend to have this "gish-gallop" energy; say a bunch of tenuously related and vaguely supported things, leave the reader wondering if it was the author who failed to connect the dots, or them
somebodythere··on The natural diamond industry is getting rocked. Thank the lab-grown variety
Maybe it's my engineer-brain talking, but "lab-grown" actually biases me towards the diamonds. Feels precise and futuristic.
somebodythere··on There are no new ideas in AI only new datasets
I don't know if it matters. Even if the best we can do is get really good at interpolating between solutions to cognitive tasks on the data manifold, the only economically useful human labor left asymptotes toward frontier work; work that only a single-digit percentage of people can actually perform.
somebodythere··on Claude 4
My guess is that they did RLVR post-training for SWE tasks, and a smaller model can undergo more RL steps for the same amount of computation.
somebodythere··on The unreasonable effectiveness of an LLM agent loop with tool use
I see what you are getting at. My point is that if you train and agent and verifier/governor together based on rewards from e.g. RLVR, the system (agent + governor) is what will reward hack. OpenAI demonstrated this in their "Learning to Reason with CoT" blog post, where they showed that using a model to detect and punish strings associated with reward hacking in the CoT just led the model to reward hack in ways that were harder to detect. Stacking higher and higher order verifiers maybe buys you time, but also increases false negative rates + reward hacking is a stable attractor for the system.
somebodythere··on The unreasonable effectiveness of an LLM agent loop with tool use
Because if the agent and governor are trained together, the shared reward function will corrupt the governor.
somebodythere··on AI 2027
I took your original post to mean that AI researchers' and AI safety researchers' expectation of AGI arrival has been slipping towards the future as AI advances fail to materialize! It's just, AI advances have been materializing, consistently and rapidly, and expert timelines have been shortening commensurately.

You may argue that the trendline of these expectations is moving in the wrong direction and should get longer with time, but that's not immediately falsifiable and you have not provided arguments to that effect.

somebodythere··on AI 2027
AGI timelines have been steadily decreasing over time: https://www.metaculus.com/questions/5121/date-of-artificial-... (switch to all-time chart)
somebodythere··on AI 2027
Did you see the supplemental material that explains how they arrived at their timelines/capabilities forecasts? https://ai-2027.com/research
somebodythere··on The sins of the 90s: Questioning a puzzling claim about mass surveillance
Federal interests can easily tell the local prosecutor "hey, don't prosecute this, it risks setting bad precedent".
somebodythere··on The sins of the 90s: Questioning a puzzling claim about mass surveillance
Seems not tinfoil and rather plausibly a pragmatic decision by the prosecution.
somebodythere··on Bitcoin puzzle #66 was solved: 6.6 BTC (~$400k) withdrawn
The market is liquid enough to absorb a sale for $400K.
somebodythere··on GitButler is now fair source
Sure. Presumably also the developers at FooLabs would like to continue having a job developing Foo, and the broader software community would like to continue benefiting from additional features and improvements to Foo, which probably wouldn't happen if developing Foo was economically unviable.
somebodythere··on Ly: Display Manager with Console UI
To start an X session after logging in
somebodythere··on Creativity has left the chat: The price of debiasing language models
Instruction is not the only way to interact with an LLM. In tuning LLMs to the assistant persona, they become much less useful for a lot of tasks, like naming things or generating prose.
Page 1 of 14Next →