19,137 karma · joined December 24, 2010
http://nt4tn.net/
Yes they do, they have their KV caches-- it's a pure function of the input tokens, sure but that doesn't prevent it from containing latent 'insight'. LLMs can and do pre-form the tokens they're expecting to output multiple steps in the future.
I wouldn't argue that the 'aha' means anything, but the structural argument that it can't that I think you're making isn't sound.
Asking it to identify problems such as spelling, grammar, or general clarity is fine.
If you tell it to "polish" many commercial LLMs will absolutely slopify your text with characteristic AI catch phrases and non-sequiturs. Beware.
It should be the other way around-- you give the AI instructions and it does stuff. Any time the AI is giving you instructions and you do stuff, that should be a big red flag.
It also has a lot of resolution and not a lot of noise. Better would be multi-turn benchmarks with tools but getting good precision and accuracy for that is hard and computationally expensive.
> --temp 0.2
Looping is a common symptom of changing the sampler settings from what it was RL trained with.
A downside is that you can't just download a lot of that knowledge, vs with the weights the copyright infringement has been outsourced to the lab. Nor can you just search for the info because the internet as a whole is increasingly aggressive at blocking anything that looks like an AI agent.
I'd love to see more retrieval powered local AI-- I think it's an area that open source development could excel. ... but there are advantages of having the knowledge in the weights!
Perhaps what needs happen is for someone to make an "ultrapedia", an AI restatement of a huge library of reference works-- created expressly for the purpose of being a locally stored corpus for AI agents.
For some usages that speedup more than makes up for it being inferior to Qwen intelligence wise.
So essentially the headline sells this as work to keep your data private, but really it's work to keep the AI-- which was trained on your code and your writing-- private.
A lot of the total cost of AI is fixing its "truth shaped errors", particularly in the presence of models that are very "gaslighty" when corrected.
GLM-5.2 is really the only model I've spent much time using that I didn't fatigue from being regularly lied to by the model, but that might be partially luck.
Not really, but a lot of what isn't used isn't very good.
More important is synthetic data. Use a teacher model with RAG with a huge reference library to write synthetic transcripts of idealized behavior for the model. Use models to judge and correct these transcripts. Train on the good ones. Use bad traces to train the model to correct its own errors (e.g. don't train it to produce a bad transcript but if it finds itself in the middle of one train it to self correct).
Similarly, for tasks that can be closed loop evaluated -- e.g. running computer software and programming, unlimited amounts of novel training data can be generated... including for highly original tasks: e.g. run publications in any domain through a model prompted to look for programming problems suggested by the material. Then write/judge/improve transcripts of solving those novel problems.
I expect in the future smaller models won't be directly trained on any internet data at all-- but entirely on simulations of idealized expected behavior from the model under construction. Raw internet data in that case would show up in prompts, but never in the target output (except of course for prompts that are asking it to copy the input).
I had some tedious debate on HN once where I asked if anyone had any pointers to good parallel SMT solvers, only to fall victim to someone dedicated to dying on the hill of "parallelization can never make this kind of search faster" due to (often inapplicable) complexity theory fixation.
For example, the minimum set cover problem shows up in cases like "What minimal set of test vectors covers all the conditions in my code?". There is an obvious greedy algorithm: "Start with nothing, pick the vector that covers the most yet-uncovered cases, repeat until all are covered".
There is an approximation result that says no polynomial time algorithm can do more than a small factor better than this greedy algorithm.
But this is a _worst case_ result, and absolutely useless for any problem you will encounter in practice.
It's trivial to come up with ways of improving the greedy algorithm: First off the simple greedy algorithm will often produce output which has completely redundant elements that can just be removed, because some collection of later added items that were necessary to cover some rare cases completely cover some earlier added item. Adding a simple postprocess to remove redundant elements immediately improves the greedy solution, particularly when the frequency of elements follows something power-law ish.
You can measure the frequency of each element and weigh uncovered elements by how rare they are (E.g. using entropy). This avoids the primary cause of the above duplicate selections.
You can use lookahead e.g. pick the pair of elements that together improve the score the most but then only commit to one.
You can use rarity weighed random starts, complete using whatever search you have, then retry multiple times.
You can compute new solutions using only the results of prior attempts. etc. etc.
In my experience basically any improvement over the greedy algorithm works on real problems, even before getting to a proper ILP solver. The greedy algorithm is just pathetic and will result in solutions much worse than you get from simple elaborations.
But over and over again you can find people being told to use the greedy algorithm because no polynomial time algorithm is better -- even in instances that are small and where actually enumerating all solutions might be tractable and justified.
But they still use system dram, so this approach allows looking into those parts of your own computer from which you're normally blocked. At least on some hardware...
https://www.policygenius.com/auto-insurance/what-is-broad-fo...
But I believe the intersection of states with highly private registration and where broad form is available is the empty set.
> I am not convinced your insurance company would data sharing with flock or its subsidiaries if they only have your vin.
Go pull your lexis nexis report.
If you imagine it being written for alt.sex.stories.moderated but somehow failing to be erotic by virtue of being too explicit, and being a few orders of magnitude better writing for that venue... then you wouldn't be too far off. [Started making a silly example of how some sex story could fail like that and then realized I was litterally describing a scene from prime intellect].
It's also arguably the origin of a lot of the mentally ill ai-safety hysteria. Arguably an enjoyable romp for those who understand that it's fiction, but it seems a lot of people cannot.
I liked the film noir voice-over.