765 karma · joined July 30, 2015
my hn username [at] jyl [dot] liCan we call that Overton window defenestration?
Meta does lose revenue when Apple and Google introduce privacy features and decrease ad tracking efficacy. They also hate that unlike the web, where their cookies and tracking pixels provided near 100% coverage of your behaviour, on your phone you might be doing something that they can’t log and analyse.
Yes, LLMs can be SOTA for NLP, but you’re going to have to use them to write software or workflows that are more deterministic.
What you call a model acquiring agency I call plain old software with productivity workflows designed by humans, with deliberate goals. We must separate “model” and an execution environment using a model. [Model] ≠ [A glorified shell script doing API calls in a control flow based on heuristics]. Agents are not AI, they are plain old software. The weights are the model, and that very much remains a static artifact (and pre-post training models haven’t improved much over the last few years).
What you call self improvement is a duck tape hack to imitate persistence and save on inference. Every time you do an API call, anything that needs to be processed is sent to the model. Narrowing that context down saves money. Finding clever ways to do that improves apparent performance and value. The cleverness is still human.
These are all useful innovations on top of LLMs, which remain models that generate text and symbols based on static weights, which in turn represent training data and the provider’s preferences.
The idea that “AI models” have acquired “agency” as of 2025 and are working on “self-improvement” in 2026 is closer to delusion than exaggeration.
To be useful they’re gonna have to be capable and powerful. If they’re good enough to be useful, they will be dangerous.
Advanced household robots are closer to fully autonomous cars and airplanes than roombas. The code is safety critical. Imagine something capable of flooding or burning down a building with hundreds of people if left to the vibe coders?
Privilege enables you to rent competence, historically by paying other people. The slop companies will now sell you a simulacrum of competence by the token.
The fact that competence can (could?) only be acquired through sustained effort over a long period of time is (was?) levelling the field.
Selling simulated competence perpetuates privilege, instead of dismantling it like you seem to claim.
There’s a similar thing going on with emails. Dozens of services ”decide” that you need to update your email address, because ”they can’t reach you”. Many of them even stop sending you emails you explicitly subscribed to, perhaps to maintain an archive, ”because you don’t seem to open them”.
No, dear Linkedin and others, you’re reaching me just fine, and it’s none of your business whether, when and where I open them. Maybe I just read my emails offline and strip your tracking links (and avoid clicking on links in emails in general).
Inexplicably LinkedIn’s UX for changing the old email address, the one they cannot reach you at (!), to a new email address, starts with confirming your current email address (THE ONE THEY CANNOT REACH YOU AT). Brilliant.
You are able to one-shot reverse parallel park into a much narrower gap without hitting the bumpers of the cars around you or getting onto the pavement.
My comments are more in the context of OLAP queries and other non-normalised data often queried via SQL.
I train non-LLM transformer models on (older and rarer) datasets, and automating the ingestion of sprawling datasets with hundreds of columns, often in a variety of local languages and different naming conventions adopted over decades, with quite a few duplicated columns…. The LLMs perform badly, it’s nigh impossible to test (for me as a user in prod) and it’s nearly impossible for the LLM companies to test (in training) to RLVR and RLHF this.
The Erdős problems have turned out to be largely brute force or finding older results.
The Feb 2026 GPT-5.2 theoretical physics paper was a result of “dialogue between physicists and LLMs”, called “grad student level” by experts in the field, used a “custom harnessed” “internal OpenAI” model with “20 hours of reasoning”. Quotes from OpenAI blog.
The Matthew Schwartz physics paper with Claude this March involved “51,248 messages across 270 sessions, producing over 110 draft versions and consuming 36 million tokens”, and the actual contribution was Schwartz finding an error in Claude’s solution.
I would have personally gone for 75%, 85% and 95%, which are all still best case scenario answers.
Had I taken on chatbot advice on electronics or chemistry I’d have died every couple of weeks (doing some hands-on real world R&D in my basement as a distraction from software).
Plane of words: broadly correct. Everything is flattened to tokens and token sequences, and the training data is dominated by text tokens.
Reasoning: CoT tokens are mostly just tokens, more appropriately called intermediate tokens, and are largely disconnected from the end result. Including them improves the end result (user satisfaction), but does not imply reasoning. See for example Turpin 2023, Mirzadeh 2024, Pournemat 2025, Palod 2025.
Synthesising evidence: You can achieve SOTA summaries with LLMs, but this involves, for example, using a harness to generate dozens of summaries with different models, separately using some kind of vector embedding model to compare results to the original, and selecting the best match. This is not how most people are using LLMs for summaries. While this is being slowly RLVR’d in post-training, a one-shot naive summary underperforms more complex methods significantly.
Video thumbnails are a different beast altogether. And you might want to double check your assumptions about security considerations. If any of your ffmpeg, opencv, pyscenedetect code is running on your server, it might well be exploitable.