Is that not how it works out of the box?
16,587 karma · joined August 28, 2009
Is that not how it works out of the box?
(I found Intuniv, which is basically the opposite of a stimulant and has the opposite side effects, actually helps more. But that's very personal and it also gives me nasty insomnia if I'm not incredibly careful.)
I've been promoting a specific aspect of performance quality recently, but it's mostly lead to other teams creating tests to monitor it that don't really make sense (along the lines of "if you use the computer it'll make the CPU % go up"), asking me to look at it, and then when I say it's probably fine I get a response along the lines of "it must be important because we already told all the executives about it."
https://www.bloomberg.com/news/articles/2025-12-03/apple-des...
Apple traditionally had a culture of a lot of high quality bug reports, but once those become too fast for anyone to keep in their head, there's no point to them because they'll just get lost. You might as well just go look for bugs if you run out of some. You'll find them.
Doesn't really help:
https://en.wikipedia.org/wiki/The_Mythical_Man-Month
Also, "bug fixing" tends to add new bugs - any kind of change can cause regressions. It's quite difficult.
For curry, it does get better if you leave it out in the sun a day and pretreat it, I forget why.
I did try asking Gemini and Claude "what would a unicorn taste like" and got… acceptable and accurate answers, but the accurate answer maybe wasn't "acceptable", because I don't think a little girl asking Claude that question should have gotten a long description about what horse meat tastes like, like I did.
So there of course are issues where an LLM having explicit knowledge about something doesn't mean it has tacit knowledge about it in all contexts, but also hiring the LLM to do a job of being a helpful harmless chat assistant constrains its abilities.
(Gemini also points out a Starbucks unicorn drink doesn't taste like horses.)
I am not up to date on philosophy of science, but the scientific method is certainly always subjective, or at least can't be successfully expressed in a formal system.
Here's a book you can read: https://metarationality.com
> Once there are - or next month when there will be - better models, agents, and agent harnesses for this, do you think that then we should concisely specify what is required instead of doing evals for particular models?
Hmm, not sure what you mean. "Evals" are another way of saying "regression tests", so they're useful when you want to change or compare any part of the system.
> and it's currently necessary to apply such procedural controls outside of the prompt?
In general I think you should try to move controls out of the prompt and into an external system, but the downside is that it costs more, so it's not always necessary.
https://www.anthropic.com/research/global-workspace
However, we don't /want/ them to have too much internal experience, because we want to know what they're thinking* for safety reasons.
* or, we want to be able to assume that the answer text is causally related to the thinking text
That's correct, that is what they think they're doing. It's based on a new religion invented by Big Yud which says we have to invent good computer gods or else there won't be anyone around smart enough to fight evil computer gods.
It is possible (likely) this was a bad idea, but you can't undiscover math once it's discovered. And it did work out last time:
https://en.wikipedia.org/wiki/Mutually_assured_destruction
…so far, anyway.
In what way do you understand the meaning of the word "unicorn" that an LLM does not? It has experienced exactly as many real unicorns as you have.
https://en.wikipedia.org/wiki/Logical_positivism#Decline_and...
Because LLMs also run off vibes and the writing style of your text, another important issue with your prompt here is that it makes you sound like a stuffy dork, or perhaps a pro se litigant. They won't respond to this well because LLMs have feelings too.
https://www.anthropic.com/research/emotion-concepts-function
Just be normal! And have evals.
They are not available to anyone. Nobody knows how they work, including the models, Sam Altman, God, etc. They're emergent from the training process.
> Change 4: with the other changes done, the models should now engage in several self analysis steps layered into its whole thinking chain. “How did I reach this conclusion, did this require any guessing, do the key facts have research support online, quick check for Claudisms or AIisms and common LLM issues, did my changes alter underlying things like libraries without integrating them etc. etc.”
Remember inference costs per token. Do you want to pay for this every time?
Only thing I ever hear about Mastodon is that you can't say or do anything without random other servers disconnecting you or random commenters yelling at you to add trigger warnings because you posted a picture of food. So I would say that being emotionally abused by German Linux users is in fact a mass psychology experiment.
They really do think AI is very dangerous and can end the world!
> 99% of the population gets poorer everyday and can feel it.
Absolutely not true in the US. What did happen is that poor people's incomes vastly increased since 2019, while upper middle class incomes stagnated.
It may of course be true in some countries.
> Or that can't afford kids at all.
Birth rates declining is a sign of increased incomes, because people can afford entertainment, and prefer to consume it over having kids, which are kind of a bummer and mean you can't have fun for the next 18 years.
Noone wants to "train on your data". You can't learn the answers to questions by pretraining on the questions, and nobody wants to teach the models to output text that looks like a user query.
The Chinese providers "train on your data" by sending your query to Anthropic and training on the answers that come back.