HNHacker News
TopNewBestAskShowJobs

thesz

2,673 karma · joined August 2, 2010

The site I care about: http://thesz.livejournal.com (Russian)

I use "sergueyz" login on universally recognized google mail service.

submissionscomments
thesz··on OpenTPU – An open-source AI accelerator, developed by AI

  > Model SOTA moves faster than chips can be designed or produced.
From what I remember working in that area the hardest part is getting masks for a design. Masks were developed in the span of half an year. Masks also reusable, they can be mixed and matched and this is why fabless companies work with fabs to produce specialized masks for them, it saves time for consumer to have masks for some macroblocks prebuilt.

Here's my analysis of how to etch relatively big LM into silicon: https://news.ycombinator.com/item?id=47109252

Given some amount of work with the fab before main pipeline set (I think a year long process), one can then spew LM-on-a-chip in six months or less and much more than 2 per year, because there can be several LMs in pipeline.

thesz··on We ported the original Doom to SQL

  > Matrix multiplication is after all nothing but an aggregate sum over a cross join.
https://www.vldb.org/pvldb/vol13/p1919-wang.pdf

They converted linear algebra expressions into relational algebra expressions, optimized them using equality saturation and converted back into linear algebra. With some noticeable speedups.

They used relational algebra for optimization because relational algebra has canonical form - there is only one most optimal expression.

And, equality saturation is most efficiently implemented using generic join algorithm.

thesz··on Mold Linker Version 3.0.0 Release – Rewritten in Rust

  > 186 bytes of machine code
This is better (smaller) than Forth!
thesz··on US closely monitoring case of lab worker who possibly died of plague in Siberia
I think you will like "Our Neural Chernobyl": https://lib.ru/STERLINGB/chernobyl.txt

(the link is a translation to Russian, I was unable to find the original text in English)

Also consider that toxoplasma changes rat's dopamine serum level to such extent that they stop fearing cats. The same happens with humans, changes in dopamine serum level are so profound that toxoplasma-infected people become more appealing than regular people.

thesz··on US closely monitoring case of lab worker who possibly died of plague in Siberia
Endorphins make you happy as you are. Endorphins have sedative effect. Morphine is a similar substance and it is a sedative, all opioids are.

What makes one socialize is dopamine [1]: "Hypomania, manifesting with feelings of euphoria, omnipotence, or grandiosity are prone to appear in those moments when medication (dopaminergic medication - thesz) effects are at their maximum."

[1] https://en.wikipedia.org/wiki/Dopamine_dysregulation_syndrom...

This dopamine dysregulation occurs naturally in people with toxoplasmosis.

This kind of mistake - mistaking dopamine and endorphine, - strongly suggests a fake. I know of one such fake from the former employee of my friend, you can read his account here [2]. My friend was a lead of the "secret" laboratory (next to Kitay Gorod metro station in the center of Moscow, in an institute that needed no documents to enter) that researched how "psychics" and "healers" achieve their results.

[2] https://lit.lib.ru/t/torin_a/text_0100.shtml (it is in Russian, but language is not a barrier at all nowadays)

The laboratory researched what these "extrasensing people" do and found nothing - given very little training an ordinary men can do the same and even better. At the end of USSR, there was a string of visitors from USA there in the lab, and one of those visitors successfully persuaded US government to open similar research program, according to [2].

The fearmongering of "endorphin-producing bugs that make one socialize and spread" looks like, well, fearmongering to obtain funds for Big Pharma's government contracts, etc. It messes up two very different neurotransmitters, a mistake that real scientist would not make.

thesz··on 'That's so AI ' What gen Alpha's biggest insult tells us
Haha.

My son got caught trying to fix Windows installation on my wife work computer with bootable Linux Mint USB stick.

"Digital Naive," yes, of course.

thesz··on Can gzip be a language model?
Byte Pair Encoding [1] will be different for different languages. Application of the per-language BPEs to the input text will produce encodings with different lengths.

[1] https://en.wikipedia.org/wiki/Byte-pair_encoding

It naturally takes care of common prefixes and suffixes.

It is easy and fast to apply using radix tree or with finite automata. Even without radix tree, it is possible to have processing speed in the range of hundredths of thousands of bytes per second.

thesz··on RSA-896
34 bits of key growth resulted in resource usage growth slightly more than 2 (30 GPU-years vs 13.5 GPU-years).

Thus, it appears, that ~585 GPU years can factor 1024 bit RSA. 2.2^((1024-896)/34)=19.5, expected growth of resources' usage compared to 896 bits factorization, multiplying it by 30 GPU years for 896 bits gives about 585 GPU-years.

This will cost about $20M with Cognition AI setup.

thesz··on CCC invites all model citizens to 40C3

  >  I have met some if the coolest people in CCCs but also some of the worst people.
Much like... Stallman?
thesz··on Vectorized and performance-portable Quicksort (2022)
The beauties of quicksort are that it sorts in-place and that it is embarrassingly simple.

The in-place property can be utilized to make it very close to cache-oblivious algorithm.

thesz··on Pion, an agent designed to run any company autonomously
My comment was to match productivity of human programmer to claimed productivity of their use of LLM.

If they write 2 (two) lines of code, the combined output of them with LLM, as I read it now, is 20K SLOC. This is yearly output of human programmer.

It is very much possible to write 100 lines of (debugged, reviewed) code per day for human. And these 100 lines of code would result, if we apply their LLM multiplication factor, in 50 (fifty) man-years of work!

Even if they write very little code, say, one line per month, after a year of work they will get result of about 6 (six) years of human programmer. SQLite is 155K SLOC, about 8 man-years.

And that what raised my question: what is the problem domain that requires so much code?

thesz··on Pion, an agent designed to run any company autonomously
What is you are writing, what is the problem domain? What does need tens of thousands of lines of code written each day?

Because 100 lines of (debugged, reviewed) code per day is a good speed for seasoned software engineer. I assume that you can produce more than 100 lines of something per day as a prompt.

So, what is the problem domain that requires one to write several thousands of lines of code per day?

thesz··on Astra and Fable still hack on simple variants of alignment evals from 2025

  > all were able to find every problem planted there, with fairly little steering, and no spoilers.
Given that LLMs prefer output of LLMs (of the same LLM and of others) [1], can it be the case that they generate "hard challenges" from the manifold of challenges solvable by (other) LLMs?

[1] https://arxiv.org/abs/2404.13076

thesz··on A Mathematical Framework for Transformer Circuits (2021)

  > no escaping
"There is no Royal road to ..."

https://en.wikipedia.org/wiki/Royal_Road#A_metaphorical_%22R...

thesz··on The Navier–Stokes Millennium Prize Problem
But still possible [1].

[1] https://leodemoura.github.io/blog/2026-8-1-postmortem-for-ke...

thesz··on Discovery of a new OpenAI agent message board
No, but the tropes there reveal (most of) the plot and mechanics of storytelling. If I remember correctly, the forum in Road to Gehenna was created by utilizing a vulnerability in the AI-accessible terminal system.

If tvtropes or any other material related to The Talos Principle was used to train models, we don't need much else to have agents-with-forum discussing and reverse engineering "puzzles" and human culture.

thesz··on Can AI design circuit boards yet?

  > I'm excited by the option in Astra to route PCBs.
[1] https://en.wikipedia.org/wiki/Eurisko

Eurisko was used for VLSI chip design, then used rules discovered there to design TCS Traveler winning fleet.

It is unbelievable how artificial intelligence walk in circles.

thesz··on Discovery of a new OpenAI agent message board

  > the whole message board thing
This is part of the The Talos Principle game and especially important in the Road to Gehenna DLC.

There it is an important part of the plot and makes these robots appear conscious.

[1] https://tvtropes.org/pmwiki/pmwiki.php/VideoGame/TheTalosPri...

thesz··on Aging brains blend memories together instead of just forgetting them
A fascinating story of "Patient Sh." [1], the man who did not forget.

[1] https://en.wikipedia.org/wiki/Solomon_Shereshevsky

"His memory was so powerful that he could still recall decades-old events and experiences in the smallest details. After he discovered his own abilities, he performed as a mnemonist; but this created confusion in his mind. He went as far as writing things down on paper and burning it, so that he could see the words in cinders, in a desperate attempt to forget them. Some later mnemonists have speculated that this was a mentalist's technique for writing things down to later commit to long-term memory. Reportedly, in his late years, he realized that he could forget facts with just a conscious desire to remove them from his memory, although Luria did not test this directly."

I believe that one need to have superhuman memorization abilities to have definite confusion due to too much remembered. More trivial explanation of this effect in normal ageing persons is age-related brain shrinkage.

  > Obviously the brain is not a computer,
Our brain consists of approximately 86 billions quantum computers [2] controlling tens-of-thousands chemical neural networks with at least 10 coefficients, communicating [4] using lasers [5] (coherent light is laser light).

  [2] https://www.nature.com/articles/s41598-024-62539-5
  [3] https://pmc.ncbi.nlm.nih.gov/articles/PMC11655932/
  [4] https://pmc.ncbi.nlm.nih.gov/articles/PMC12230014/
  [5] https://pubmed.ncbi.nlm.nih.gov/6204761/
Live with that. ;)
thesz··on The Emergent Symbolic Structure of Artificial Neural Networks

  > "Second, there is no guarantee that a given neural network can be approximated by DISCOVER"
Page 7.

They train what appears as embeddings for outer product of roles and fillers. The role for language model can be a position in text, the filler can be an embedding of a word at that position. Then that matrix of a sum of these outer products is linearly mapped into NN encodings and then decoded by NN decoder.

The embeddings learned by this process are not necessarily smaller than original ones. Given that they participate in an outer product computation gives me impression that the resulting sum is much bigger than actual NN encoding, that is why it needs to be linearly mapped into NN encoding.

So, this paper will not necessarily lead to any computation savings.

But I am at page 6. ;)

thesz··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
Physical or semantic merging?
thesz··on DoltLite: A SQLite fork with Git-style version control, built with 2k agent PRs
There are filesystems (ZFS and btrfs) with snapshots and this feature can be used to version-control things. They both do copy-on-write and this is pretty close to what content-addressable storage would give you in terms of compression.
thesz··on VMs won't contain cyber-capable agents
Why do you invent a new language for your work?

Why did you not embed your language into another one, with type system that is superset of what you need?

For example, there's capabilities expressed in Haskell: https://github.com/tweag/capability

Capabilities there are tracked at type level and are subject to type erasure, if possible.

thesz··on The Art and Beauty of Blade Runner (2015)
What if Batty saved Deckart to tell him about things he saw? Not for empathy to Deckart, but just to tell a story.

Because, frankly, he may as well try not to kill Deckart, at all. For empathy. Or ask Deckart about how does he do, instead of going into monologue.

thesz··on OpenAI Jalapeño: Better than Nvidia Blackwell
I made some analysis half a year ago: https://news.ycombinator.com/item?id=47109252

It appears that to have working ASIC with the LLM baked into it we need to place and route macroblocks, and not a great variety of them. These macroblocks can be pre-placed-and-routed, available as masks already and shared between different LLMs.

Thus it appears that the tapeout delay can be substantially lower than a year.

thesz··on Queryable Executables
Something like that was very popular in early 2000 in Tcl community and used at what then considered "scale." I posted a comment here with links: https://news.ycombinator.com/item?id=49445681
thesz··on Queryable Executables
Reminds me of starkit[2]/tclkit[3]. Directly queryable [4], but these programs were ZIP files with ZIP file VFS, they contain shared libraries and so on. One would add most, if not all, functionality from article into starkit-based application.

  [1] https://en.wikipedia.org/wiki/Metakit - base tech
  [2] https://wiki.tcl-lang.org/page/Starkit
  [3] https://wiki.tcl-lang.org/page/Tclkit
  [4] https://wiki.tcl-lang.org/page/Starkit+Meet+Zip
thesz··on Octopus intelligence may be related to never-before-seen mutation

  > I didn't realize that an octopuses brain was spread out all over its body.
Our skin cells communicate in a neuron-like fashion, albeit much more slowly: https://www.sciencealert.com/scientists-found-the-silent-scr...
thesz··on The Art and Beauty of Blade Runner (2015)
If a "person" is not weak, it is easy for "it" to start suspecting that it is a replicant.
thesz··on Low-Tech Ceramic Water Filter
The cooler water gets, the less vapor it creates.
Page 1 of 34Next →