But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"
But you're right in sense: moderate intelligence at superhuman rates (and presuming moderate energy usage) is very compelling compared to an intelligence that takes 1000 years to return "42"
So imagine taking a year-old SOTA model and running it at 100 tokens per second on an edge device. That's enough to feed a screen's worth of content through it and power decent multilingual message suggestions on IM.
Imagine running it at 1000 tps. That's enough to reparse that screen mid-keystroke, and give you semantic autocomplete in text. Or fully general "the phone has a good idea of what you're attempting to do" context at all times.
There's many, many new classes of features that will open up if decent enough models can be run on edge devices at 100+ "intelligence per second".
I'm poking (lightly) at the locallama game, have a fancy MBP5 with gobs of ram (so I can demo/trial locally) and have been semi-waffling between whether to chase a mini or studio for local "always on" type stuff.
My outcome was "CapEx v. OpEx", and dropping another $5k for an aluminum cube buys a lot of OpEx (eg: just trickle-drip HF/OpenAI credits to a raspberry pi or VPS orchestrator rather than trying to do the inference locally), ie: $5/mo inference for 1000 months.
HOWEVER, there's definitely a role for that 1-10 TPS type "ambient inference" that I wouldn't mind sustaining on any sort of always-on / local / private compute. My main email address is still on ...@yahoo.com and their spam filtering has gone to absolute shit.
Being able to have the always-on mini (local, trusted, no private data leaves my control) poke at the IMAP/email and thresh it into SPAM/HAM/Personal/Political ... random spot check, I'm getting ~5 emails per hour, and that's completely tractable for staged low-med-high processing. (Subject + rules.py? Subject + Body + LocalSlowTPS? Subject + Body + LocalDeepTPS? Subject + Body + RemoteLLM?)
Even if you did 10000tps of "jev" that's an incredible value... not quite "Literal AI Packet Router", but as you're dancing aroud saying... "Speed is a Weapon"
Look into "OODA Loop" => """The OODA loop is a four-step decision-making model—Observe, Orient, Decide, Act—created by U.S. Air Force Colonel John Boyd to help leaders make fast and accurate choices in chaotic situations. // The main goal is to cycle through the loop faster than an opponent or changing environment. By operating inside another person's loop, you create confusion and outpace their ability to respond. While initially designed for aerial combat, it is now widely used in business, sports, and crisis management."""