I wonder if this is a global study. Cynically, it feels like a way to get a few more billions of tokens of free alignment training material. They explicitly say that they will “use the results to inform how we develop our models and products”.
257 karma · joined February 9, 2023
I wonder if this is a global study. Cynically, it feels like a way to get a few more billions of tokens of free alignment training material. They explicitly say that they will “use the results to inform how we develop our models and products”.
3D NAND flash can indeed routinely store hundreds of GB per die, so that’s proven. The question is about all the peripheral circuitry needed. Each attention block would need its own KV cache (i.e. SRAM or DRAM somewhere), plus DAC/ADC inputs and outputs, unless they figure out a way to keep it analog all the way (really cool but unlikely).
I think this field is very interesting, at least from a technology point of view. Whether it works out or not will sadly be a matter of economics more than physics I fear.
Flash NAND can routinely be bought with 4 bits per cell, perhaps even 5 soon (QLC and PLC drives). Since it has been proven to store 4 digital bits at production scale, I’m willing to bet an analog architecture running an LLM should be able to yield the performance analog of a 4-8 bit quantized model. Where in that 4-8 range is pretty crucial, but it depends on the specific design.
[1] https://www.nature.com/articles/s41928-023-01010-1 [2] https://www.nature.com/articles/s41586-022-04992-8
Btw, can guarantee that they are not ready to demonstrate that yet. They’re using 2D FLASH with 30M weights per die [1], so to get to 1T they will need… 33,333 dies. Interesting scaling problem to say the least
They are using GlobalFoundries’ 28nm node for the floating gate transistors, afaiu, which they then bond onto a TSCM 5nm digital I/O wafer.
Analog computation’s principles are sound. It’s mostly doing matrix vector multiplications though. The rest is digital.
Here is the top 10 right now:
| Rank | Model | YapIndex | YapTax$ |
| ---: | ---------------------------------------- | ---------: | ------: |
| 1 | openai/gpt-5.6-sol (reasoning) | 18.5 ±4.8 | 0.51 |
| 2 | openai/gpt-5.6-sol | 19.2 ±4.4 | 0.51 |
| 3 | openai/gpt-3.5-turbo | 22.7 ±4.8 | 0.02 |
| 4 | openai/gpt-5.6-luna | 27.8 ±9.8 | 0.15 |
| 5 | openai/gpt-5.4 (reasoning) | 40.7 ±10.7 | — |
| 6 | openai/gpt-5.4 | 40.7 ±9.0 | — |
| 7 | moonshotai/kimi-k2-0905 | 44.7 ±4.8 | 0.05 |
| 8 | mistralai/mistral-small-2603 (reasoning) | 46.2 ±31.5 | 0.03 |
| 9 | openai/gpt-4 | 51.2 ±20.6 | 1.39 |
| 10 | openai/gpt-5.3-codex | 64.8 ±9.9 | — |
I also checked how the Claude models did specifically: | Rank | Model | YapIndex | YapTax$ |
| ---: | --------------------------------------- | ----------: | ------: |
| 23 | anthropic/claude-opus-4.5 | 97.0 ±28.9 | 1.52 |
| 25 | anthropic/claude-opus-4.5 (reasoning) | 99.2 ±29.3 | 1.44 |
| 42 | anthropic/claude-3.5-sonnet | 199.7 ±24.5 | 2.53 |
| 57 | anthropic/claude-sonnet-4.5 (reasoning) | 278.7 ±41.7 | 1.65 |
| 61 | anthropic/claude-sonnet-4.5 | 285.0 ±39.4 | 1.63 |
| 71 | anthropic/claude-opus-4.6 | 330.3 ±59.1 | — |
| 72 | anthropic/claude-haiku-4.5 | 333.2 ±26.7 | 0.64 |
| 73 | anthropic/claude-haiku-4.5 (reasoning) | 335.2 ±26.8 | 0.65 |
| 76 | anthropic/claude-opus-4.6 (reasoning) | 342.3 ±30.5 | — |
| 87 | anthropic/claude-3.5-haiku | 401.2 ±25.4 | 0.59 |
| 88 | anthropic/claude-sonnet-4.6 (reasoning) | 422.8 ±23.3 | — |
| 89 | anthropic/claude-sonnet-4.6 | 425.0 ±26.3 | — |
No other Anthropic model seems to be there for now. I would have been very curious to see the more recent ones.It's really unusable for anything other than code. And I have to remove its incomprehensible comments 50% of the time before committing anyway. After interacting with it, "slop vomit" is truly the most fitting description. I have to admit I have lost my temper and spontaneously referred to its output as vomit more than once. Seems like I'm not the only one.
Point is, sometimes people just have different experiences from you.
This substack article also comes with additional graphs, a much better story flow (data is progressively introduced and explained before reaching the final plot), and was posted 2 days before the OP. I agree with GP that it is significantly superior to the OP (which is likely AI slop). Thanks for posting it.
For my usage, I would very much prefer if those $/task were being spent in thinking and experimenting, and the actual output would be as short and maintainable as possible. “maintainability” is a vague target of course, but it’s at least somewhat correlated with code size.
I wonder why nobody has tried to optimize for actual code size or complexity metric, or at least why I haven’t seen more benchmarks that display this. GPT5.5 just keeps pushing more and more pointless indirection into every function it writes in my main project, it’s borderline negative productivity.
P.S. I’d be curious to see Cursor’s composer models in there, they seem to be among the best performing low cost models: https://artificialanalysis.ai/articles/cursor-composer-2-5-c...
The most absurd part is that everyone in academia knows that publish or perish is tremendously damaging to real research. Yet we’re all hostage of this system that we created in the name of “merit” and “efficiency”.
We need a different system to identify and reward talented hard-working people. Back in the day it all relied on actual interpersonal interaction and subjective judgment, but there were also much fewer researchers worldwide.
The issue with “just” photons and electrons is that you need something else to force them to behave like you want. And photons are large and non-interacting, really the opposite of what you want for computing. Great for communications of course.
And you could write nice sci-fi about subatomic transistors, but forget making them in this reality.
Still, it’s known that your skin color is the main thing that matters, which is why Australians have the worst melanoma statistics. I guess mine tans fast enough to keep up with the seasons here.
To draw a parallel with physics, about which we as a society (and me as an individual) know a lot more, we are gonna define mathematical objects and laws whose behavior maps well to certain subsystems of the brain that we ourselves have defined. Physics had it easy in this sense: it turned out to be remarkably simple to describe the universe to a great degree of precision. Still, we all recognize that the physics mapping we have to this day isn't perfect; and it's even possible that it will never be perfect. It may be fundamentally impossible to reduce some systems to simpler mathematical objects which we can reason about. It does seem to be generally possible, however, to find a reasonable approximation. This again is what I think Rovelli's point is about: science is the process of finding a good approximation which has predictive power. And what is there in the brain that's fundamentally so different and that we're never gonna be able to explain? Why does everyone keep insisting that consciousness is special?
I do agree that that only when (if) we do get the accurate mathematical description, then we're gonna be able to properly discuss the hard problem. But my hunch is that once we do have all the tools it will just dissipate from scientific discussion, similarly to how the measurement problem in quantum mechanics is slowly undergoing the same transition from "this is a fundamental problem" to "we were just asking an invalid question due to misunderstanding and old ways of thinking". Obviously I have no proof of this beyond intuition, but I more or less agree with every sentence of this article, and that shapes my intuition.
To anticipate a possible question about my definition: I don’t have a strict one. I’m almost completely with Rovelli on this one. I think the day we find a proper definition of the concept we’ll have done the first step is solving the (one and only) “easy” problem of consciousness. But I’m open to hearing your own definition since I feel like I just can’t grasp your concerns. I must be missing something.
There’s very little research work needed to make this happen; it’s all about engineering some satellite buses and having them fly in close formation to get a “data center”. And this group of satellites in sun-synchronous orbit would relay to a comms constellation e.g. starlink itself) and operate as a global scale data center. The heat management and orbital mechanics are all straight forward really.
[1] https://quoteinvestigator.com/2010/05/01/misbehave/#e6c0268a...
I may think stridently (debatable) but I generally believe it is best to always try to meet in the middle if the goal is genuine discussion. This is my attempt at that.
I had never heard of you, and this article appeared very biased to me. I found the information ecology piece superior, shame that it went unnoticed; I will try to go through all of them. I admire the breadth of topics you’re covering and appreciate the many sources. They’re clearly written in your own voice and that is great to see, I guess I mostly reacted to not being fully aligned with your view.
[1] https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
I do think that safety is important. I'm particularly concerned about vulnerable people and sycophantic behavior. But I think it's better not to be a luddite. I will give a positively biased view because the article already presents a strongly negative stance. Two remarks:
> Alignment is a Joke
True, but for a different reason. Modern LLMs clearly don't have a strong sense of direction or intrinsic goals. That's perfect for what we need to do with them! But when a group of people aligns one to their own interest, they may imprint a stance which other groups may not like (which this article confusingly calls "unaligned model", even though it's perfectly aligned with its creators' intent). People unaligned with your values have always existed and will always exist. This is just another tool they can use. If they're truly against you, they'll develop it whether you want it or not. I guess I'm in the camp of people that have decided that those harmful capabilities are inevitable, as the article directly addresses.
> LLMs change the cost balance for malicious attackers, enabling new scales of sophisticated, targeted security attacks, fraud, and harassment. Models can produce text and imagery that is difficult for humans to bear; I expect an increased burden to fall on moderators.
What about the new scales of sophisticated defenses that they will enable? And for a simple solution to avoid the produced text and imagery: don't go online so much? We already all sort of agree that social media is bad for society. If we make it completely unusable, I think we will all have to gain for it. If digital stops having any value, perhaps we'll finally go back to valuing local communities and offline hobbies for children. What if this is our wakeup call?
- Text isn't selectable on the page.
- The tooltip in the "day 1" to "day 14" cards gets cut off by the border (I see this mistake ALL the time with AI-generated frontends btw)
- It's sparse and very long. I think the information could be condensed in half the size, and it would improve the presentation. This is personal preference though.
- The playbooks' "mark complete" are not persisted on reload or navigation.
All in all, it's functional and quite decent. I agree with the other people saying it looks generic, but I disagree on it being necessarily a bad thing for this kind of product.
I know nothing about pools so I can't comment on the accuracy of the playbooks. It's nice that there's so many of them, but given the LLM vibe of the text I'm slightly suspicious.