Does this mean that a single incorrect word or twitch will completely derail the task you’re trying to performance? Or will you, like any other intelligent being, recognize it and compensate?
926 karma · joined April 21, 2017
username (at) penrose.com
Does this mean that a single incorrect word or twitch will completely derail the task you’re trying to performance? Or will you, like any other intelligent being, recognize it and compensate?
Modern chain-of-thought models with RL post training on verifiable tasks + realistic environments + rubrics are worlds apart from models trained on a simple next token prediction objective.
More money goes into the rubrics and RL environments than individual training runs themselves.
(Yes, at inference-time LLMs still output words one at a time, much like human speakers. But don't confuse the mechanism with the training objective.)
Recent advances in mathematical/physics research have all been with coding agents making their own "tools" by writing programs: https://openai.com/index/new-result-theoretical-physics/
Example: hypertargeted ads for F-35 engine upgrades in the DC metro - https://x.com/JosephPolitano/status/1683476652276236295
Is that clear enough?
> That the engine's website puts a different thing that whatever the search algorithm thought was the best thing at the top because an advertiser paid for it to be so, is strictly worse.
This is not clearly worse than the result being selected by the whims of some arbitrary Google engineer, or being easily gamed by SEO blogspam bots. At least the advertiser stands to lose something if they bid incorrectly.
>How long had ChatGPT been out before OpenAI's first ad for it?
Just because some products were able to grow organically doesn't imply that paid marketing never benefits startups. This is a false equivalence.
I also find it funny that the vast majority of your example products (everything except Huel or Aeropress?) make a lot of money from advertising. Maybe consider why they still exist.
Yes, it does - it’s called advertising. In the US, the average promotional spend per physician exceeds $20k/yr. As a result, a lot more patients are able to quickly benefit from new medications like Dupixent or Ozempic as a result of wider awareness.
Suppose we banned Google ads and you are searching for a plumber. You are now entirely at the whims of whoever designs the ranking algorithm on Google/the Yellow Pages, who has nothing at stake here. Meanwhile, advertisers have to bid for your attention - making them at least somewhat aligned with your buying intent.
The same applies for doctors searching for state of the art diabetes treatments. It’s hard to say that relying on a fuzzy notion of “legitimacy” (or entrenched status-quo cliques) is a more fair system.
Word of mouth benefits incumbents. Advertising at least enables newcomers to temporarily burn money to gain mindshare, while “slow diffusion” will lock society into a “nobody ever got fired for buying IBM” state forever.
These exist? Companies are making billions of dollars selling persistent environments to the labs. Huge amounts of inference dollars are going into coding agents which live in persistent environments with internal dynamics. LLMs definitely can live in a world, and what this world is and whether it's persistent lie outside the LLM.