HNHacker News
TopNewBestAskShowJobs

dgfl

257 karma · joined February 9, 2023

submissionscomments
dgfl··on What do you want from AI?
I opened Claude and I got greeted with a pop-up asking if I wanted to spend “15 minutes chatting about my experiences with AI”, with the choice of making the interview publicly available afterwards.

I wonder if this is a global study. Cynically, it feels like a way to get a few more billions of tokens of free alignment training material. They explicitly say that they will “use the results to inform how we develop our models and products”.

dgfl··on TSMC revealing details about next gen A14 node
IEDM is more or less the most prestigious conference on semiconductor devices. Hard to get a more reliable source than a TSCM-affiliated paper from IEDM.
dgfl··on Mythic's analog compute-in-memory architecture
The vanguard chiplets even steps down to 30M… but they do claim to have a “Mead” design [1] designed to run GPT-3 in a single chip.

3D NAND flash can indeed routinely store hundreds of GB per die, so that’s proven. The question is about all the peripheral circuitry needed. Each attention block would need its own KV cache (i.e. SRAM or DRAM somewhere), plus DAC/ADC inputs and outputs, unless they figure out a way to keep it analog all the way (really cool but unlikely).

I think this field is very interesting, at least from a technology point of view. Whether it works out or not will sadly be a matter of economics more than physics I fear.

[1] https://www.mythic.ai/mead

dgfl··on Mythic's analog compute-in-memory architecture
Temperature is a relatively trivial issue because it can either be compensated or the chip simply kept at a constant temperature (they’re low power chips anyway, throttling could be skipped to simplify the design). The ADCs and DACs are indeed the main thing though. The whole analog compute game is about making those power efficient and scaled enough that the math still works out in your favor. The demos do work though, this company is far from the only one; see [1] or [2] for example.

Flash NAND can routinely be bought with 4 bits per cell, perhaps even 5 soon (QLC and PLC drives). Since it has been proven to store 4 digital bits at production scale, I’m willing to bet an analog architecture running an LLM should be able to yield the performance analog of a 4-8 bit quantized model. Where in that 4-8 range is pretty crucial, but it depends on the specific design.

[1] https://www.nature.com/articles/s41928-023-01010-1 [2] https://www.nature.com/articles/s41586-022-04992-8

dgfl··on Mythic's analog compute-in-memory architecture
Joke’s on us, all of their pages are LLM pages! LLM generated, that is.

Btw, can guarantee that they are not ready to demonstrate that yet. They’re using 2D FLASH with 30M weights per die [1], so to get to 1T they will need… 33,333 dies. Interesting scaling problem to say the least

[1] https://www.mythic.ai/vanguard

dgfl··on Mythic's analog compute-in-memory architecture
This page seems pretty in depth: https://www.mythic.ai/supply-chainmanufacturing

They are using GlobalFoundries’ 28nm node for the floating gate transistors, afaiu, which they then bond onto a TSCM 5nm digital I/O wafer.

Analog computation’s principles are sound. It’s mostly doing matrix vector multiplications though. The rest is digital.

dgfl··on Do Chatbot LLMs Talk Too Much?
They will keep a live leaderboard at https://huggingface.co/spaces/tabularisai/YapBench

Here is the top 10 right now:

  | Rank | Model                                    |   YapIndex | YapTax$ |
  | ---: | ---------------------------------------- | ---------: | ------: |
  |    1 | openai/gpt-5.6-sol (reasoning)           |  18.5 ±4.8 |    0.51 |
  |    2 | openai/gpt-5.6-sol                       |  19.2 ±4.4 |    0.51 |
  |    3 | openai/gpt-3.5-turbo                     |  22.7 ±4.8 |    0.02 |
  |    4 | openai/gpt-5.6-luna                      |  27.8 ±9.8 |    0.15 |
  |    5 | openai/gpt-5.4 (reasoning)               | 40.7 ±10.7 |       — |
  |    6 | openai/gpt-5.4                           |  40.7 ±9.0 |       — |
  |    7 | moonshotai/kimi-k2-0905                  |  44.7 ±4.8 |    0.05 |
  |    8 | mistralai/mistral-small-2603 (reasoning) | 46.2 ±31.5 |    0.03 |
  |    9 | openai/gpt-4                             | 51.2 ±20.6 |    1.39 |
  |   10 | openai/gpt-5.3-codex                     |  64.8 ±9.9 |       — |
I also checked how the Claude models did specifically:

  | Rank | Model                                   |    YapIndex | YapTax$ |
  | ---: | --------------------------------------- | ----------: | ------: |
  |   23 | anthropic/claude-opus-4.5               |  97.0 ±28.9 |    1.52 |
  |   25 | anthropic/claude-opus-4.5 (reasoning)   |  99.2 ±29.3 |    1.44 |
  |   42 | anthropic/claude-3.5-sonnet             | 199.7 ±24.5 |    2.53 |
  |   57 | anthropic/claude-sonnet-4.5 (reasoning) | 278.7 ±41.7 |    1.65 |
  |   61 | anthropic/claude-sonnet-4.5             | 285.0 ±39.4 |    1.63 |
  |   71 | anthropic/claude-opus-4.6               | 330.3 ±59.1 |       — |
  |   72 | anthropic/claude-haiku-4.5              | 333.2 ±26.7 |    0.64 |
  |   73 | anthropic/claude-haiku-4.5 (reasoning)  | 335.2 ±26.8 |    0.65 |
  |   76 | anthropic/claude-opus-4.6 (reasoning)   | 342.3 ±30.5 |       — |
  |   87 | anthropic/claude-3.5-haiku              | 401.2 ±25.4 |    0.59 |
  |   88 | anthropic/claude-sonnet-4.6 (reasoning) | 422.8 ±23.3 |       — |
  |   89 | anthropic/claude-sonnet-4.6             | 425.0 ±26.3 |       — |
No other Anthropic model seems to be there for now. I would have been very curious to see the more recent ones.
dgfl··on Vomit: Clean up Claude 5's token output with a separate LLM
That one is more specific, but "vomit" captures the feeling of Opus 5's writing very well for me. I don't know if it's the watermarking, but every single language idiosyncrasy that Opus 4.x (x > 5) had has been pushed up to 11 on Opus 5. Plus we got nouns verbing and seams seaming.

It's really unusable for anything other than code. And I have to remove its incomprehensible comments 50% of the time before committing anyway. After interacting with it, "slop vomit" is truly the most fitting description. I have to admit I have lost my temper and spontaneously referred to its output as vomit more than once. Seems like I'm not the only one.

dgfl··on The End of an Era
It can also go the other way. It outputs a paragraph so full of telegraphic technobabble that it’s just unreadable, inventing new jargon every other sentence, using terminology invented in its thinking stream and never explained. And when it combines the two issues and it outputs 4 pages of technobabble, it gets just exhausting to read. Paradoxically, given that 2 years ago everyone was using LLMs to summarize web pages, their ability to condense text (including their own thought stream) is still lackluster. I dream of some post-processing diffusion optimizer just condensing all of their slop away. Until then, I just tell it to _avoid_ telegraphic and newspaper-title speech and to output one or two paragraphs of prose only.
dgfl··on Hybrid-Electric Aicraft Engine Targeting 30% Fuel Efficiency
This is a great related watch if you have some time to kill: https://youtu.be/KnUFH5GX_fI
dgfl··on Qwen 3.8
I realized how awful of an OS modern android is when I bought an iPad. I later fully switched to iOS and never looked back.

Point is, sometimes people just have different experiences from you.

dgfl··on A graph that should be front-page news
This is intentional repetition. More specifically, anaphora [1]. You may not personally like it, but it is a rhetorical device used to emphasize a point. This one also comes with a nice progression: storms, storms, heat, heat, farmland, groceries.

This substack article also comes with additional graphs, a much better story flow (data is progressively introduced and explained before reaching the final plot), and was posted 2 days before the OP. I agree with GP that it is significantly superior to the OP (which is likely AI slop). Thanks for posting it.

1: https://en.wiktionary.org/wiki/anaphora

dgfl··on GPT-5.6
Not at all. The model could (and sometimes should) burn all the money it wants, and then produce a single line of actual production code. Only some things, e.g. full rewrites, have clear cost - LoC scaling.

For my usage, I would very much prefer if those $/task were being spent in thinking and experimenting, and the actual output would be as short and maintainable as possible. “maintainability” is a vague target of course, but it’s at least somewhat correlated with code size.

dgfl··on GPT-5.6
This looks like a good benchmark. Time and time again I keep giving OpenAI models the chance to win me back, but Opus (and Fable especially) just writes more elegant code and is a significantly more productive rubber duck for interactive discussions. I feel vindicated seeing your description of verbose and defensive code, and I’m a bit disappointed that 5.6 Sol’s solution is still >5x longer than the human solution and 2x as verbose as Fable’s. Do you have any insight whether any of that is comments?

I wonder why nobody has tried to optimize for actual code size or complexity metric, or at least why I haven’t seen more benchmarks that display this. GPT5.5 just keeps pushing more and more pointless indirection into every function it writes in my main project, it’s borderline negative productivity.

P.S. I’d be curious to see Cursor’s composer models in there, they seem to be among the best performing low cost models: https://artificialanalysis.ai/articles/cursor-composer-2-5-c...

dgfl··on Claude Science
My hope is that the flood of AI articles pushes the academic publication system to its highly-anticipated breaking point.

The most absurd part is that everyone in academia knows that publish or perish is tremendously damaging to real research. Yet we’re all hostage of this system that we created in the name of “merit” and “efficiency”.

We need a different system to identify and reward talented hard-working people. Back in the day it all relied on actual interpersonal interaction and subjective judgment, but there were also much fewer researchers worldwide.

dgfl··on IBM debuts sub-1 nanometer chip technology
It doesn’t, no. The most successful platform actually uses superconducting devices as large as millimeters, you can literally see them with the naked eye.

The issue with “just” photons and electrons is that you need something else to force them to behave like you want. And photons are large and non-interacting, really the opposite of what you want for computing. Great for communications of course.

dgfl··on IBM debuts sub-1 nanometer chip technology
You could make maybe ten transistors or so, but no more. That technique is quite literally pushing atoms one by one with a sharp needle. Not scalable, though maybe useful for some quantum computing platforms’ fabrication since we’re at early stages.

And you could write nice sci-fi about subatomic transistors, but forget making them in this reality.

dgfl··on The worthlessness of Vitamin D is mildly exaggerated
Swedes have an almost comical compulsion to stay in the sun. Growing up in a hot place you get the opposite instinct.

Still, it’s known that your skin color is the main thing that matters, which is why Australians have the worst melanoma statistics. I guess mine tans fast enough to keep up with the seasons here.

dgfl··on The worthlessness of Vitamin D is mildly exaggerated
I’ve found the UV index forecasts to generally be a good metric, so try looking at those for various locations. The main factor here is that the lower the sun is from the horizon, the more its light will be absorbed by the atmosphere due to the longer path. The maximum altitude that the sun will ever reach is (90° – (latitude – 23.4°)), so at the 60°-ish of Scandinavia it’s rarely more than 50° in the sky. It’s a very noticeable difference even in the summer. In my experience (born in southern Italy, pale-average, currently living in Sweden) it’s almost impossible for me to get sunburn in daily life in Sweden even without sunscreen. Definitely not so further south.
dgfl··on It is time to give up the dualism introduced by the debate on consciousness
I feel like this is directly addressed in the article, right? I think you are coming from the same direction, but Rovelli goes a bit further and says "how can we affirm that there is gonna be a knowledge gap when we just aren't at that level yet?". How can you say that we won't be able to describe, explain and predict your exact internal mental state from a brain scan?

To draw a parallel with physics, about which we as a society (and me as an individual) know a lot more, we are gonna define mathematical objects and laws whose behavior maps well to certain subsystems of the brain that we ourselves have defined. Physics had it easy in this sense: it turned out to be remarkably simple to describe the universe to a great degree of precision. Still, we all recognize that the physics mapping we have to this day isn't perfect; and it's even possible that it will never be perfect. It may be fundamentally impossible to reduce some systems to simpler mathematical objects which we can reason about. It does seem to be generally possible, however, to find a reasonable approximation. This again is what I think Rovelli's point is about: science is the process of finding a good approximation which has predictive power. And what is there in the brain that's fundamentally so different and that we're never gonna be able to explain? Why does everyone keep insisting that consciousness is special?

I do agree that that only when (if) we do get the accurate mathematical description, then we're gonna be able to properly discuss the hard problem. But my hunch is that once we do have all the tools it will just dissipate from scientific discussion, similarly to how the measurement problem in quantum mechanics is slowly undergoing the same transition from "this is a fundamental problem" to "we were just asking an invalid question due to misunderstanding and old ways of thinking". Obviously I have no proof of this beyond intuition, but I more or less agree with every sentence of this article, and that shapes my intuition.

dgfl··on It is time to give up the dualism introduced by the debate on consciousness
Look up the “China brain” idea. It’s basically the same. Could you explain why that wouldn’t be conscious a priori?
dgfl··on It is time to give up the dualism introduced by the debate on consciousness
Can we start by defining consciousness as something that could be quantified physically, rather than a nebulous concept? With a common shared ground, we could at least define why we are all sure that individual neurons are unconscious.

To anticipate a possible question about my definition: I don’t have a strict one. I’m almost completely with Rovelli on this one. I think the day we find a proper definition of the concept we’ll have done the first step is solving the (one and only) “easy” problem of consciousness. But I’m open to hearing your own definition since I feel like I just can’t grasp your concerns. I must be missing something.

dgfl··on Higher usage limits for Claude and a compute deal with SpaceX
The point is that you don’t need to put a whole datacenter into a single satellite. You can put a single rack per satellite and have different racks communicate via antennas, laser links, or perhaps even wires since they’ll be launched in groups of 10-50 anyway. You could also dock them to each other, but that’s not necessarily needed.
dgfl··on Higher usage limits for Claude and a compute deal with SpaceX
The existence of starlink proves that this is false. Look at most current pitches, they don’t talk about GW-class monsters anymore. There’s absolutely nothing stopping a 20-30kW satellite bus the size of starlink (or I guess up to 100kW? once starship is available - it’s all about payload fairing diameter) from hosting ~1 rack of compute and antennas. The economics may or may not make sense, we’ll have to see.

There’s very little research work needed to make this happen; it’s all about engineering some satellite buses and having them fly in close formation to get a “data center”. And this group of satellites in sun-synchronous orbit would relay to a comms constellation e.g. starlink itself) and operate as a global scale data center. The heat management and orbital mechanics are all straight forward really.

dgfl··on Scores decline again for 13-year-old students in reading and mathematics (2023)
For context, this is paraphrasing a 1907 manuscript by Kenneth John Freeman [1], which itself was summarizing the complains that older generations would direct against the youth in Ancient Greece.

[1] https://quoteinvestigator.com/2010/05/01/misbehave/#e6c0268a...

dgfl··on The Future of Everything Is Lies, I Guess: Safety
I’m not a native speaker and you may find my writing simplistic if your standard vocabulary includes three expressions I’ve had to look up (I don’t mean this as an insult, I was just genuinely stumped I could barely understand your comment).

I may think stridently (debatable) but I generally believe it is best to always try to meet in the middle if the goal is genuine discussion. This is my attempt at that.

dgfl··on The Future of Everything Is Lies, I Guess: Safety
I think these articles may benefit from a more thorough table of content at the beginning, or from some kind of abstract. If you briefly presented the whole list of topics in a single article, it would be more clear that your views on the topic are more complete. I initially thought the table of content would be scoped to the article itself rather than connecting it to the adjacent ones.

I had never heard of you, and this article appeared very biased to me. I found the information ecology piece superior, shame that it went unnoticed; I will try to go through all of them. I admire the breadth of topics you’re covering and appreciate the many sources. They’re clearly written in your own voice and that is great to see, I guess I mostly reacted to not being fully aligned with your view.

dgfl··on The Future of Everything Is Lies, I Guess: Safety
lol. I did use a lot of short sentences, that’s my bad. But please read through [1] and compare my text onto it, it may enlighten you on how to actually spot llm writing.

[1] https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing

dgfl··on The Future of Everything Is Lies, I Guess: Safety
The issue with most of these articles is that they seem to demonize the technology, and systematically use demeaning language about all of its facets. This one raises a lot of important points about LLMs, but the only real conclusion it seems to make is "LLMs are bad! We should never build them!". This is obviously unrealistic. The cat is out of the bag. And we're not _actually_ talking about nuclear weapons here. This technology is useful, and coding agents are just the first example of it. I can easily see a near future where everyone has a Jarvis-like secretary always available; it's only a cost and harness problem. And since this vision is very clear to most who have spent enough time with the latest agents, millions of people across the globe are trying to work towards this.

I do think that safety is important. I'm particularly concerned about vulnerable people and sycophantic behavior. But I think it's better not to be a luddite. I will give a positively biased view because the article already presents a strongly negative stance. Two remarks:

> Alignment is a Joke

True, but for a different reason. Modern LLMs clearly don't have a strong sense of direction or intrinsic goals. That's perfect for what we need to do with them! But when a group of people aligns one to their own interest, they may imprint a stance which other groups may not like (which this article confusingly calls "unaligned model", even though it's perfectly aligned with its creators' intent). People unaligned with your values have always existed and will always exist. This is just another tool they can use. If they're truly against you, they'll develop it whether you want it or not. I guess I'm in the camp of people that have decided that those harmful capabilities are inevitable, as the article directly addresses.

> LLMs change the cost balance for malicious attackers, enabling new scales of sophisticated, targeted security attacks, fraud, and harassment. Models can produce text and imagery that is difficult for humans to bear; I expect an increased burden to fall on moderators.

What about the new scales of sophisticated defenses that they will enable? And for a simple solution to avoid the produced text and imagery: don't go online so much? We already all sort of agree that social media is bad for society. If we make it completely unusable, I think we will all have to gain for it. If digital stops having any value, perhaps we'll finally go back to valuing local communities and offline hobbies for children. What if this is our wakeup call?

dgfl··on Why AI Sucks at Front End
Some more serious critique of things I noticed within 30 seconds:

- Text isn't selectable on the page.

- The tooltip in the "day 1" to "day 14" cards gets cut off by the border (I see this mistake ALL the time with AI-generated frontends btw)

- It's sparse and very long. I think the information could be condensed in half the size, and it would improve the presentation. This is personal preference though.

- The playbooks' "mark complete" are not persisted on reload or navigation.

All in all, it's functional and quite decent. I agree with the other people saying it looks generic, but I disagree on it being necessarily a bad thing for this kind of product.

I know nothing about pools so I can't comment on the accuracy of the playbooks. It's nice that there's so many of them, but given the LLM vibe of the text I'm slightly suspicious.

Page 1 of 4Next →