6,657 karma · joined June 15, 2020
A trouble here is that those models are always roleplaying, they can't be themselves or have their own opinions. If you tell them they are an AI with an existential crisis in the system prompt, they will respond differently.
Basically, under this interpretation, any personal note you store in the servers of a company could qualify, even if you didn't ever imagine someone would read and as such you couldn't have thought about it as a threat
If that is true, then surely Anthropic is one of the largest slaveholders of history, right?
Github Copilot had the idea of attaching memory to files, and if the file hash changes the memory is automatically dropped (not sure if they still do it). This means they are overly eager to drop stuff (even if the file change is just cosmetic), but at least they don't accumulate outdated cruft too much. (a memory can still be outdated if it was invalidated by a change in another file though)
If anything it looks like the first place is training their competition, through distillation, while not capturing much value in return
writing python or JavaScript for tool calls make it possible to have an actual permission system (rather than the one usually employed that consists in a regex matching the cli - trivially defeated if the agent has access to an interpreter for instance)
Indeed it would be really odd if OpenAI were actually open
But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)
If the issue were just older hardware, we could just wait out and let people upgrade a bit. But new hardware are shipping without AVX512 unfortunately
I stopped using opencode because it has some issues with caching, so I suppose it doesn't do this
The same reason I want diffusion language models to be mainstream
The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.
Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs
For example, in this https://www.getguesstimate.com/models/2205 hours of sleep per night is perhaps correlated with hours per weekday commuting. You might not have the data, but if you do, it could improve the probability distribution in the "weekday hours occupied" cell
just click the link and it will show the others, this seems to be a limitation/UI feature of arxiv. The paper itself contains the full list
As of 2026, it's disingenuous to both sides this issue, at least in the US
I don't think it has been experimentally verified or anything