868 karma · joined July 17, 2016
Apple Intelligence was an outlier in how they historically operate. I mean they literally just put out a folding phone after waiting what, a decade, for others to perfect it first.
That's what I was hoping for with Taalas, ultimately ending up with a microsd-ish card that's hot swappable intelligence.
Okay let me just install a tool to do that because the OS doesn't usually have it built in in 2026.
Oh, it's riddled with ads. Oh it's paid for. Etc.
It's the same thing with the recent opening of knowledge thanks to AI.
Before: search for recipe with search riddled with ads, find website riddled with ads cookies notification requests forced sign up and a million other dark patterns, then an entire fucking life story before the recipe.
Now: "find me a good recipe for pasta that uses the stuff I have left over from the party last weekend, I don't mind buying a few more bits".
And there's only so much attention to go around.
I count syllables, not rain—
Whose noticing?
Generative Pretrained Transformer 5.
So more tokens/variability and slow or fewer tokens and fast.
There seems to be a threshold tho, like taalas is super fast but that model is so dumb, being dumb faster doesn't work, seems to be some minimum requirements.
The above is a serious threat though. What do you do, in a democracy, with people who are dangerously uninformed?
I'd always thought we'd eventually hotload loras or MoE experts.
It would certainly be useful on the robotics/VLA side of things as well; more limited mobile hardware, download and load/unload new skills as needed.
Tbf I also don't really care what facts my models have baked in (for llms at least). I care most that the model understands general logic and then general knowledge of some level is secondary. Reason being is that everything is RAG'd in anyway.
Models spitting out well established facts is cute but I don't really ever want to rely on say "electronics knowledge" that exists in a tenuous and vague form in the model weights.
Humans write books (and datasheets) for a reason. Books are RAG.
I've just been doing research and experiments for work related stuff.
Typically we've used plain embeddings for a lot of high contrast documents aka discrete facts.
However I've been working with a >1000 page document of complex procedures with incredibly low contrast where embedding falls flat.
There's top down/graph searching, bottom up/embedded; alts like colbert, reranking, reasoning, search agents and now (though seemingly quite new) specific search agent models.
Ultimately I found that a reasoning enabled search agent doing a hybrid of bottom up (with reranking) followed by top down, gave the absolute best results. Paired with Luna for cheaper and faster tokens it benchmarks pretty well even for vague references to procedures.
I would imagine that search specific models just coming out are even better and I'll have to evaluate using these but for now the above works well for us.
Having an agent get vector search results to use as anchors and then being able to explore the sections and subsections above that, then eventually digesting as much as is relevant (big context, cheap tokens) is amazing.
Speaking to the entire company gets me nervous, but strangely not as much as other things, I think because I have to be delivering a presentation there's less time to spin in a loop worrying.
My GW4 classic has been dogshit slow for years now, they have absolutely fucked it up with no concern for older devices in the updates. And all fucken reddit can say is the usual tribally inspired "Well it's old and has a slower processor" yeah no fucken shit.
I'm not buying a brand new smart watch every 3 years, fuck that.
How exactly? Samsung watches only require the pin to be enabled if you want payments enabled. You only have to enter the pin if you have taken the watch off. As long as it can detect that it remains on your wrist then payments continue to be enabled no pin required...
See also ds/3ds flash used in game carts. They need to be powered on every so often to avoid bitrot. Their memory controllers use a form of ecc that causes weak bits to be rewritten. But if it's left too long then ecc cannot correct. This is over 15-30 year timescale fwiw.
But all chips will die in like 50-100 years or so because physics gets em with electromigration.
Fears are further reinforced if, when you trip over your words and get all anxious, your audience laughs at you rather than with you.
Prevents AI being used to one shot the writing and helps them practice public speaking as well (even for people like me who hate public speaking).