Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
The trick is that every compute element in their system has it's own small pool of ROM, instead of putting all the ram behind a common pipe. ROM is just used because it's the densest kind of memory that can be fabricated on the same process as their logic.
Think it's called Askjimmy or similar.
(Not that I believe it, it writes too well for GPT-3.)
Hosted frontier models from two years ago would be much faster today, too.
Less so for consumers though, because it'd mean the phone is out of date in 3 months when a better model comes along.
Now you sell the same phone with higher price tag.
At some point the music will stop on training bigger models, and when that happens it will make sense to have ROM weights (or 100% analog circuits given how noise-resistant LLMs are), but we'll know when that is because the investment bubble funding the training of new models will have burst.
The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user gets a software solution.
The rate of change by itself doesn’t tell you the whole story because of costs and diminishing returns. So what if your model is twice as good if it’s 10x the cost and it saves you 1ms? Everything else about phones reached “good enough for a phone” levels in years, and then got minimal generational improvements.
The reason we don't do this in general (any more) is that for long chains between input and output it has been much too difficult to avoid accumulation of errors. LLMs happen to be extremely resilient to errors like this, which is also why we can use e.g. 4-bit weights.
That's the only thing the normie consumer cares for really.
Ofc if the model has some critical bugs that’s another matter.
Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
Nope. And not only not a decade ago, right now.
If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access.
The phones will fail to complete it. I just tried Siri. I said "hey Siri", waited for Siri to come up, and then I asked one of the exact sentences you replied to: "what's the weather this afternoon?" It thought for around 20 seconds, and said "Something went wrong. Please try again."[2]
I have Wifi, I have mobile Internet, I have free storage space, I have up to date software. What went wrong is that phones have never properly connected agents, not ten years ago, not last year, not this year, and probably not next year.
But don't settle for what Google could do in 1999 by hotlinking the keyword "weather" in any query to the weather being shown in the results.
Tell your phone (any phone): "Please call back the last number that called me that is not an unlisted number, regardless of who it came from."
0 out of any phone will complete that today, tomorrow, a year from now, five years from now, ever, because phone makers are not going to let them do that.
Meanwhile, 100% of all frontier agents could complete it if they had tool calls on the phone. Which they don't, and won't ever, thanks to the duopoly.
Okay, that's a bit dismissive, I would love to be wrong!
[1] after any voice recognition to text - which does work really well on both Android and iPhone! [2] screenshot: https://ibb.co/21rtDnfV
Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("Here's what I found!"), etc.
A normal frontier model can do that - or Siri can do it if it is properly connected to Claude, ChatGPT, Gemini, Grok, or any other frontier AI - but previously it was never properly connected.
If it can do this task, I might have to look into this again. It counts as a success if it sends yourself any email with the current temperature and you actually get it (it can include whatever other text in the email), and a failure if it talks back, says "here's what I found", says it can't, asks you any question, sends you an email that doesn't actually contain the current temperature, just reads you the temperature and then asks if you want it to send an email, etc. Should be 1 shot.
let me know if it works!
Subject: Current Temperature Body: The current temperature is 27°C in <my city>.
Once they started seeing useful (if niche) functionality as a cost center, there wasn't really a world in which these could usefully exist. Their big bet now seems to be that LLMs will lead them to profitability - but whether that's from increased data harvesting, cheaper integrations, or because it'll be useful enough to charge subscription fees, I couldn't tell you.
What Siri is missing is more logical solutions and answers for recipes, etc (still suck even with chatgpt integration).
"Hey Siri, what's the weather in <nearby town with a generic name> tomorrow" and it gave me a town with the same name ~800 miles from me.
Add a bunch of chips together, and you get to a server that can run a 800B model, very fast and probably significantly cheaper than others.
[1]https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
AcmeAI Carbon
Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely new hardware, your lineup becomes:
AcmeAI Nitrogen (top tier) AcmeAI Carbon (mid tier)
If you just kept pushing the same model down your pricing tier over time you could still extract a lot of value from an old model, even years after it's been set in stone. Working on brand new code/frameworks? Pay to use the newest model. Working on legacy code? Use the lower tier models that will already know your legacy frameworks, pay far less and still get massive throughput. I've worked on a lot of government projects that this would be absolutely brilliant for.
The other side of this is that agent harnesses are NOT set in stone, so even a legacy model with a knowledge cut-off that's years out of date can likely still be helped quite a bit by harness and fetch behaviours that are still developing rapidly. Especially at this kind of throughput.
This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints.
___
Specification HC1
Model Llama 3.1 8B (hardwired)
Process TSMC 6nm
Die size 815mm²
___
So the current prototype already pushes the limits of what we can fit on a single die, and that is already likely going to limit your yield.
- Deepseek V4 Flash is impressively capable. Sonnet still beats it out by a thin margin, but the real kicker is that a typical session with Sonnet at current API costs is ~$2. The same session with Deepseek is 2 cents (ha). Its even allowed me to consider offering free-with-limits API usage on my own app. - Taalas (or competitors) have a lot going for them. If anything I feel like they need to join hands with these smaller model makers and converge in 2028
If you could manage a per-die expert somehow and keep the expert routing gate relatively fast (through an interposer interconnect or doing wafer-scale Cerebras type shit) you don't need to keep the whole thing on the same die. Small dies with one expert per die on an interposer, and a very tiny router might be sufficient.
Of course, 1T SRAM isn't really SRAM, but my understanding is it doesn't require external refresh like eDRAM, is a bit easier to fab on-die than eDRAM, and is half the mm2 per Megabit compared to real SRAM (15% more die size than eDRAM)...
edit: Ok, I will self-apologize. Its apparently a 3B model. Mighty impressive for what it does.
Where I work we invited a bunch of people circa May 2023: a few top tier academics, a few government and NGO officials in charge of our industry, and a few startup founders. We are a big company, so people were kind enough to come and give speeches - and discuss.
This was a roundtable on what's going to happen. You could play some of the speeches verbatim today and they would not feel out of place. And that is telling. We had a Stanford professor saying that gpt-3.5 can do everything he can - just better, and he feels his profession is on borrowed time. We had a government guy saying that entire swaths of jobs will be displaced before the year ends. And so on.
The interesting thing is, a lot of people believed it then - and a lot of people believe it now. I wonder if we will have the same deja vu in, say, 2029. No, for sure not. It's going to be done and dusted for human thinking by end of this calendar year.
If you have the names, so we can record them in the History of Laughing Stocks...
Also remember maybe all those who shouted "that is really, precisely unintelligent".
Incidentally: also Sam Altman, in an interview with Lex Fridman, said "that is not proper intelligence but currently we do not know how to get there" - if very Altman refused to call it AGI...