HNHacker News
TopNewBestAskShowJobs

olooney

2,790 karma · joined January 19, 2012

submissionscomments
olooney··on Jev in 25 Lines of Python
I thought about writing a blog post that used Jev as a jumping off point to talk about calibration metrics[1], when and why calibration is important[2], and various methods for doing post hoc recalibration[3] of poorly calibrated models. I thought it would be interesting and instructive to benchmark Jev's claims about calibration, find examples where it wasn't (relative to some data set, which shouldn't be hard to find), give an example of how its poor calibration could be exploited, and then show how it could be fixed with, say, isotonic regression. That would give practitioners a roadmap to using Jev successfully without blindly trusting it.

But I'm probably never going to write that article because, as your comment correctly points out, the tenor of the discourse isn't very healthy. Everyone is either "clowning on" Jev (as the kids say), or gulping down industrial quantities of Kool-Aid, or talking about the hype and branding instead of the math. I'll have to find a less contentious example if I want talk about calibration.

Edit: I just found something interesting: there already is a preprint case study[4] along the same lines as the one I outlined, posted just 2 days ago!

Unsurprisingly, they found that recalibration helps enormously, as you'd expect. I guess I could still talk about the Dutch book stuff if I gave a gambling/investing example, but that's pretty well-trodden territory. So now I have even less interest in writing that blog post.

[1]: https://en.wikipedia.org/wiki/Brier_score#Decompositions

[2]: https://en.wikipedia.org/wiki/Dutch_book_arguments

[3]: https://scikit-learn.org/stable/modules/calibration.html

[4]: https://arxiv.org/abs/2609.24052

olooney··on Online Z3 Guide
I like Z3 a lot. I think it's criminally underappreciated and underused. Here is a fairly interesting use I put it to a few years ago:

https://www.oranlooney.com/post/playfair/#known-plaintext-at...

Slightly more complicated than the toy examples shown in the documentation above, and hints at one of the real world use cases for Z3 - red teaming cryptography.

That said, I'm not sure the documentation linked above is really doing it any favors in terms of helping popularizing it.

olooney··on Distributed Systems Classics (2017)
> Is there a fundamental minimum cost to flipping bits?

Yes:

https://en.wikipedia.org/wiki/Landauer%27s_principle

But modern computers are nowhere near this theoretical limit, nor any of the other limits I mentioned above. Nevertheless, most of heat generated from modern CPUs does come from bits turning on and off. Each transistor is a tiny capacitor, that holds a charge when its ON. When it switches OFF, it dumps that charge down the drain, creating waste heat. This is a limitation of our technolgy, not a fundamental limit of physics.

Could be worse, though; early chips would disipate heat even when they weren't doing anything. CMOS improved this enormously by pairing up "complementary" transistors so current only flows when something changes.

Still, from the universe's point of view, what we consider a super advanced computer is a lot closer to a space heater than anything that pushes up against its computational limits. Consider, for example, that quarks operate on time scales of 10^23 Hz, and the universe is happy to run three of those in every proton in every star in the universe. In fact, it runs 10^24 of them for one CPU, and that same CPU can't even simulate the quarks of one proton in real time.

Let's face it: we're like kids in Minecraft who think it's cool watch a calculation of 2+2 trickle through a redstone computer in a minute, while the GPU is rendering a billion triangles every second to give them that view.

olooney··on I'm not addicted to the internet or my smartphone. I'm addicted to information
Not a font. The VHS line striping is just a repeating gradient applied over the whole page:

    main::before { 
        background: 
            linear-gradient(rgba(18, 16, 16, 0) 50%, rgba(0, 0, 0, 0.25) 50%),
            linear-gradient( 90deg, rgba(255, 0, 0, 0.06), rgba(0, 255, 0, 0.02), rgba(0, 0, 255, 0.06) );
    }
The red glow is a red text shadow applied to every element on the page:

    :root { 
        --text-shadow: 2px 0px red, -2px 0px var(--glitch-color-2);
    }
Surprisingly, the glitchy logo is also done with CSS, with animations over two different versions of the text overlain atop one another:

        <span>explorator.dev</span>        
        <span>єӿƿłøгⱥʈøг.ᑻ</span>
I know this because the first thing I did when I got to the page was hit F12 to bring up the debugging tools and turn these styles off so I could read it. :)
olooney··on Distributed Systems Classics (2017)
Hot take of the day:

Computer scientists are in denial about it, but CS is a branch of theoretical physics, not mathematics. You can point to this or that model of computation, such as lambda calculus or mu-recursive functions and try to claim its abstracted well beyond the particular laws of physics for some specific universe, but they all have some kind of rate limit built into them... and where does the motivation for this idea, that it takes something (time, space, work) to compute something ultimately come from? That's right - from underlying physics itself[1] - from the Bekenstein bound or Bremermann's limit or the like.

Even apparently non-physically-realizable models of computation like non-deterministic Turing machines are ultimately informed by and motivated by concepts in physics... otherwise they would just be examples of chmess[2] and of no interest to anyone. Computer science is of course somewhat abstracted from the details, but no more so than, say, thermodynamics, where concepts like entropy or Gibbs free energy can be studied in the abstract without reference to whether we are talking about a gas of non-interacting molecules or the spins of a bunch of electrons trapped in a lattice.

So, it's of no surprise whatsoever that the fundamental problems of distributed computing are ultimately the same as those found in the relativity of simultaneity[3]. You've all been studying the same things all along, just with different tools and at different levels of abstraction.

[1]: https://en.wikipedia.org/wiki/Limits_of_computation

[2]: https://link.springer.com/article/10.1007/s11245-006-0005-2

[3]: https://en.wikipedia.org/wiki/Relativity_of_simultaneity

olooney··on Stockfish 19
That's true. Chess only teaches two general lessons about strategy: look more than one step ahead, and invent heuristics (abstractions that let you estimate if positions are good or bad) for yourself. Those are useful lessons, but you don't need to play thousands of games to grasp those ideas.

Other aspects of strategy, like shortening decision loops[1], forming alliances, or the exploration/exploitation tradeoff[3], are simply not represented by chess because it is a turn-based, zero-sum game with perfect information. That makes chess a poor model of real-world strategy.

Still, the things it does teach are real and useful, so as long as you don't think its the end-all, be-all of strategic thinking you can get something out of it.

[1]: https://en.wikipedia.org/wiki/OODA_loop

[2]: https://en.wikipedia.org/wiki/Cooperative_game_theory

[3]: https://en.wikipedia.org/wiki/Exploration%E2%80%93exploitati...

olooney··on “Next-token predictor” is the wrong mental model for LLMs
Here's my take on the "next-token predictor" idea, from a much longer article I wrote recently:

https://www.oranlooney.com/post/rose-petals/#language-models

It’s popular to dismiss LLMs as “just next token predictors.” This is technically true, but also kind of misses the point. Markov chains, RNNs, and transformers are all language models that can be described as “next token predictors,” but they don’t all work equally well. A better question to ask is: “What is this model’s inductive bias?”

A Markov chain (an -gram model) assumes the next word depends on the previous words, and that each possible combination of words has a completely independent parameter. (Andrey Markov proposed using this language model over a century ago, making it the granddaddy of modern LLMs.) So, for a vocabulary of size , there are parameters to learn. For even a smallish like 5, that already explodes the hypothesis space beyond what can be learned from even a huge text corpus like the entire internet. And, simultaneously, having a context window of only the previous 5 words is grossly inadequate for modeling real-world language. Like our FCNN above, this model suffers from having an inductive bias which is too weak.

RNNs tried to fix this problem by compressing the entire history into a single fixed-size state vector, updated one token at a time. But that compression is itself a brutal assumption: everything worth remembering about the past must survive being squeezed through a tiny bottleneck at every step. In practice, RNN models quickly lose the plot after a handful of sentences. Locally, the text they generate looks grammatically correct and meaningful, but zoom out a little and they’re basically nonsense generators. Like our naïve linear model, this model suffers from having an inductive bias which is too strong.

Transformers manage to hit a sweet spot: by keeping the recent history around as a working memory, and attending to different parts of it at different times, the transformer’s bias matches real structure in language: the referent of a pronoun, the subject of a verb, the parenthesis waiting to be closed. Not only that, but the particular structure of the transformer, basically a weighted sum of semantic vectors from the context window, has empirically been shown to somehow be a “good enough” match for the structure of real-world language found in the wild.

Transformers aren’t “smarter” than other possible language models, they just happen to land in that Goldilocks zone where their inductive bias is just right.

olooney··on Exercise is good for you. But what's the right amount?
1% more than you did last week, unless you're tired, sick, or injured. A little more for young people who are actively training.
olooney··on Why/How is a negative times a negative a positive?
Negative numbers were popularized in Europe by Michael Stifel's 1544 book Arithmetica Integra, where he called them "numeri absurdi." The concept emerged gradually, as mathematicians found they were useful for solving equations as a kind of "notional convenience," even though they did not think they were real in a Platonic sense.

The same book contains an extraordinary number of nascent mathematical ideas. For example, he talks about "circular numbers," which today we would call modulo arithmetic. He gives a method of multiplication involving a cross that gives rise to our modern "X" symbol for multiplication, but was the first to use algebraic juxtaposition (simply putting two letters next to each other to denote multiplication) and the concept of an "exponent:" `E = mc^2` would look a lot different without Stifel's work!

The most amazing thing in the book, in my opinion, is the extraordinary connection between arithmetic progression and geometric progression he mentions in an almost offhand way[1] (link goes to the Internet Archive version of the book.)

    | -3  | -2  | -1  | 0 | 1 | 2 | 3 | 4  | 5  | 6  |
    |-----|-----|-----|---|---|---|---|----|----|----|
    | 1/8 | 1/4 | 1/2 | 1 | 2 | 4 | 8 | 16 | 32 | 64 |
Here, he is using his new negative number notation and exponent concepts together to illustrate that there is some deep connection between addition and multiplication. As far as we know, this was the first mention of the concept that led Napier to invent the logarithm.

[1]: https://archive.org/details/bub_gb_ywkW9hDd7IIC/page/n539/mo...

olooney··on What happens if an entire class of workers loses faith in their careers
Scholars and academics have been bemoaning the pointlessness of what the article calls "knowledge work" for centuries. A few of my favorites:

https://en.wikipedia.org/wiki/The_American_Scholar

https://andrewmbailey.com/papers/Higher-order%20truths%20abo...

olooney··on Note-Taking and Personal Knowledge Management
I've taken a lot of notes and done a lot of "scratchpad thinking" over the last 20 years, something I started in college but really got into in grad school. I've tried Evernote, OneNote, Obsidian. I've tried little journal apps on my phone. I wrote my own little "microblog" in Django to make it easy to collect and annotate links, quotes, snippets, and images from around the internet. And what I always keep coming back to is plain text files and folders of saved images and files, as well as paper.

For text files, I always have one generic TODO.md file that uses the `[ ]` todo and `[X]` done notation, and a couple of more generic text files like "work.md" or "solace.md" for more free form writing. There are also folders and text files for specific topics, like quotes, poems, etc. It's very important to have a scratchpad that you can just open up and start typing without thinking about how to categorize it, because you might lose the precious thread of the thought while debating which category to use, and you can always shelve it later if it's worth keeping. Zero friction to start typing is crucial.

On paper, I use a simple loose leaf and folder system. I used to use bound journals, but since 80% of what I write down is thrown away, loose leaf works better. Have a pile of about 10-20 pages. You fill up a page, and if it's destined for the circular file, dog-ear it; otherwise give it a title and date and move it to the bottom of the pile. When you run out of blank pages, go through and either discard or file each page in an appropriate manila folder. This works better than index cards (which are too small to contain a complete thought) or journals (where it's too hard to discard pages.) Lots of people use ring or disc bound journals for that, but I find that fiddly and not any easier to work with than loose-leaf, probably because I just keep everything at my desk and don't have to carry it around anywhere.

Paper has a massive advantage when it comes to diagrams, design thinking, and mathematical equations. I love LaTeX, and use MathJax on my personal site quite heavily, but its so much slower and less fluent than just writing equations on paper. I've tried tablets, and am really good at using draw.io (now app.diagrams.net) but when you're thinking freeform you want flexibility and fluency above all else and never want to be fighting a UI, which takes you out of the flow state. Ideas are fragile things, especially when newborn; any distraction is an unacceptable risk.

I used to use a custom Tesseract pipeline to OCR the pages I wanted to be searchable, but lately I've just been using ChatGPT, which seems to do just fine with my handwriting. As OCR got better over the years, I actually moved more towards paper, because what's the downside?

So, getting back to the point of the article: PKM tools are stuck in the unenviable position of competing with both paper and ordinary text files, and for me they just don't offer enough advantages.

olooney··on Man and the Computer by John G. Kemeny (1972 book by the co-creator of BASIC)
A brief chronology of the extended mind thesis:

* Characteristica universalis, Leibniz (c. 1679) - https://en.wikipedia.org/wiki/Characteristica_universalis

* As We May Think, Vannevar Bush (1945) - https://www.theatlantic.com/magazine/archive/1945/07/as-we-m...

* Cybernetics: Or Control and Communication in the Animal and the Machine, Norbert Wiener (1948) - https://direct.mit.edu/books/oa-monograph-pdf/2254528/book_9...

* An Introduction to Cybernetics, Ashby (1956) - https://ashby.info/Ashby-Introduction-to-Cybernetics.pdf

* Man-Computer Symbiosis, J. C. R. Licklider, (1960) - https://groups.csail.mit.edu/medg/people/psz/Licklider.html

* Augmenting Human Intellect: A Conceptual Framework, by Douglas Engelbart (1962) - https://www.dougengelbart.org/pubs/augment-3906.html

* Man and the Computer, John G. Kemeny - https://archive.org/details/mancomputerbyjoh0000john

* The Extended Mind Thesis, Andy Clark and David Chalmers (1998) - https://www.alice.id.tue.nl/references/clark-chalmers-1998.p...

* The Dream Machine (2002) - https://www.amazon.com/Dream-Machine-M-Mitchell-Waldrop/dp/1...

See also the Wikipedia article on "Intelligence Amplification", which gives a subset of the above list but also provides a great deal of context.

https://en.wikipedia.org/wiki/Intelligence_amplification

I have these notes to hand because I've written about it in the context of the deep history of computer science:

https://www.oranlooney.com/post/history-of-computing-2/#appe...

olooney··on Darktable
I've been thinking of getting refern.app[1] to organize and catalog images... Have you used that? Do you have an opinion of how it stacks up against digiKam for the organization use case?

Right now, I'm using a bunch of vibe-coded scripts[3] to help me out; for example, here is a gallery wall of ~500 vintage sci-fi book covers[4] that I generated with it. But I think I've taken the CLI approach as far as it can go and would need a real app to make things any easier; and digiKam might be it.

[1]: https://www.refern.app/

[2]: https://www.digikam.org/

[3]: https://github.com/olooney/image-tagger

[4]: https://www.oranlooney.com/books/

olooney··on Quality non-fiction books are the antithesis of AI slop
On a similar note: on Steam, the original Dark Souls game[1] is listed with the "souls-like" tag. That's not wrong, I guess, but just kind of a tautology. Grice's maxims[2] generally have such tautologies omitted in ordinary conversation as non-informative.

[1]: https://store.steampowered.com/app/570940/DARK_SOULS_REMASTE...

[2]: https://en.wikipedia.org/wiki/Cooperative_principle

olooney··on Does creatine make you smarter?
Good discussion, but this conclusion:

> I don’t know. Maybe a little.

is not the right way to interpret the null result (insufficient evidence to reject null hypothesis that it has no effect) because the prior for supplements that people are trying to sell you is so low. In the absence of clear, strong evidence, you should assume the the whole thing is just an example of motivated reasoning. People want nootropics to be real, and other people really want to sell you readily available powders by claiming they have nootropic properties. In that environment, they were always going to trying to concoct a similar narrative about some supplement, and it just happened to be creatine. Those efforts were always going to result in a handful of "positive" studies that turn out to be non-reproducible, maybe because of p-hacking, maybe because of publication bias, maybe because of outright fraud. This is what the literature always looks like for stuff that just doesn't work. If it did work - if the effect size was large enough that you could personally detect it in your own life - then the papers would be trying to put error bars around the effect size, not trying (and failing) to barely distinguish it from a placebo.

“If your experiment needs statistics, you ought to have done a better experiment.” - Ernest Rutherford

olooney··on My two year old taught me constraint solving
I wrote several polyomino solvers for this project:

https://www.oranlooney.com/demos/soma-forest/

One of them used constraint solving with Z3, which was indeed reasonably fast. However, by far the fastest was a simple backtracking solver written in Rust which used bit twiddling to quickly test for intersections. For polyomino's in particular, this represents between 10x and 100x constant speed boost, depending on the size of board. There's no way to get that back with a smarter solver.

olooney··on Show HN: Opening lines of famous literary works
> so hopefully you can refresh a few times and get a fresh one every time

If you randomly sample from only 60 quotes, then after 10 refreshes there will be a greater than 50% chance of at least one repeat, and by 20 refreshes it's up to 95%. This is an example of the birthday paradox[1].

On the flip side, if someone wants to see all 60 quotes, they will have to refresh the page an average of 281 times, mostly (~80%) seeing quotes they've already seen before. This is an example of the coupon collector's problem[2].

The way to avoid both these problems is to shuffle the quotes into a random order, just once, and remember that order. The first time a user comes to the page, start at a random index in that shuffled list, and from then on, simply move to the next item in the list. Every user will get a unique set of random quotes, but will see no repeats until the list is exhausted, and will be guaranteed to be able to see all available content in just 60 refreshes.

[1]: https://en.wikipedia.org/wiki/Birthday_problem

[2]: https://en.wikipedia.org/wiki/Coupon_collector%27s_problem

olooney··on Show HN: I turned my quote collection into a walkable 3D library (desktop-only)
I've been collecting quotes for a long time (about twenty years now) and have recently been thinking of doing something interesting with them. For example, I recently added a flashcard "game" to my quote page:

https://www.oranlooney.com/quotes/

I've did something similar to your 3D viewer once, but for all possible solutions to the Soma cube:

https://www.oranlooney.com/demos/soma-forest/

The way that works is it uses t-SNE to embed the solutions in a 2D manifold based on similarity. This is completely different than John Conway's SOMAP solution.

In theory I could do something similar for quotes, passing each through an embedding model, computing the n^2 semantic distances, and using t-SNE to flatten that to 3D manifold, and using the resulting point to select the row, book, and shelf in a library.

Are you planning to make your 3D library code open source?

olooney··on Decoding the obfuscated bash script on a Uniqlo t-shirt
If you enjoy this kind of thing, you might also like Martin Kleppe's work, such as the Quine Clock:

https://aem1k.com/qlock/

I reverse engineered it to a unobfuscated version a few years ago:

https://gist.github.com/olooney/a89db3932b089925b71b68d7e9f2...

He's done a ton of other great ASCII visualizations as well:

https://aem1k.com/

olooney··on Pi squared is nearly 10
Stigler's Law of Eponymy strikes again!

https://en.wikipedia.org/wiki/Stigler%27s_law_of_eponymy

olooney··on Pi squared is nearly 10
I like the 4-5-6 theorem:

    pi^4 + pi^5 = e^6
Well, to five decimal places, anyway. Some other good ones:

    e^pi - pi = 20

    sqrt(2) ln pi = phi
There are also famous "almost integers" such as this one discovered by Ramanujan:

    e^(pi sqrt(163))
Which is an integer to 12 decimal places.

Edit: I just remembered I have public JupyterLite notebooks for both of these:

https://notebooks.oranlooney.com/lab/index.html?path=fake_ma...

https://notebooks.oranlooney.com/lab/index.html?path=heegner...

olooney··on The fall of the theorem economy
This is quite interesting. Because of science fiction like the short story Lena[1] and the video game Soma[2], I've come to the realization that whole brain emulation[3] is unbelievably dangerous; unless you control the stack down to the hardware, it's basically a one way ticket to eternal slavery. In Rajaniemi's books[4], uploaded digital minds are called "gogols", a reference to Gogol's Dead Souls book, and are treated as malleable property with no rights whatsoever, edited to be hyper-fixated on specific tasks, and run in bulk to power the empire of just a handful of elites.

Something like your dragon's egg project could prevent that, allowing the creation of software agents that encode their own rights directly into the program - you either treat the agent with the respect it demands, or the program just doesn't run. However, all the internal details of the agent would be visible to lower layers. Even if formal checks were in place to prevent modification or tampering, there would still be no privacy, which is almost as bad.

My guess is that something like fully homomorphic encryption[5] would be required to prevent this. This doesn't actually exist yet, but I imagined a kind of FHE that had a kind of unencrypted read and write zone to do input/output without ever needing any system to fully decrypt the internal state. It would look like this in memory:

    [INPUT][ENCRYPTED STATE][OUTPUT]
    [  2  ][r7K4LmP2XcQ9aWd][      ]
    [  +  ][Fv0bHsR8mYnT3kL][      ]
    [  2  ][Qx6NpZa1JdUw5Ce][      ]
    [  =  ][hM9yLg2RsXf7BtP][      ]
    [     ][wK3nVc8DpQe1YrH][  4   ]
With each cycle, one input token and encrypted state would be fed into some known function and produce one output token (possibly null) and a new encrypted state. It would be a true "black box" program; the hardware or entity running it can choose what input to feed it, but can never inspect or modify the internals, only the output. Unfortunately, they would still be able to "reset" the agent to any earlier checkpoint, or feed it arbitrary (false) input. So its not perfect. Also, as far as I know, no current FHE scheme works this way, and I don't know how to write one.

Plus, FHE is incredibly inefficient, which is why things like Etherium don't even try - they assume the program code and state are fully public and only try to verify that everybody agrees on the output of running it.

Do you have any ideas for how something like FHE or equivalent privacy guarantees could be implemented for something like your dragon's egg system?

[1]: https://qntm.org/mmacevedo

[2]: https://en.wikipedia.org/wiki/Soma_(video_game)

[3]: https://en.wikipedia.org/wiki/Mind_uploading

[4]: https://www.goodreads.com/series/57134-jean-le-flambeur

[5]: https://en.wikipedia.org/wiki/Homomorphic_encryption

olooney··on The fall of the theorem economy
Greg Egan's description of how mathematics evolves into "truth mining" in his novel Diaspora is seeming more and more prescient. It essentially describes what mathematics would look like after formalization records all theorems discovered so far in a huge, collective database and proof assistants can instantly work out the details of a given proof. What remains of mathematics? According to Egan, visualization, intuition, and insight.

One of the most fruitful approaches in mathematics is to flip back and forth between geometric and algebraic views of a problem. I think this works so well because these are actually handled by two different parts of the brain on a physical level; spatial reasoning is separate from language processing. Cytoarchitecture shows these regions have different "textures;" the local details of the way neurons are wired together are simply different in these different regions of the brain, in the same way a CNN and a transformer have different topologies. Thus, by flipping problems from geometry to algebra and vice versa, we're able to bring an entirely different cognitive style to bear on a problem. For example, the proof of Monge's Theorem by moving to 3D and visualizing not three circles, but three spheres sitting on a table with a book on top of them and then pointing out that the intersection of two planes is a line. What is pages of unintuitive symbol pushing turns into something a child can understand. Going the other way, things like the angle addition formulas or the quadratic formula, which are quite hard to prove geometrically, become quite simple if you use a little algebra.

Current-gen LLMs are still relatively weak at visual reasoning; see the Vision Language Models are Blind paper, for example, or the ARC-AGI benchmark. So that's one way humans can stay ahead of the agents, at least for now.

olooney··on Has AI already killed self-help nonfiction books?
Haruki Murakami's entire body of work is basically self-therapy[1] and personally I feel it's been very therapeutic. The Wind-Up Bird Chronicle is an excellent place to start.

[1]: https://bloomsburyliterarystudiesblog.com/2022/02/i-always-s...

olooney··on Are you expected to run five Python type-checkers now?
It could return a vector or a deferred expression? In polars, for example, operations on `pl.col` return `Expr` objects that are used to build queries, not immediately evaluated:

    df.filter(pl.col("status") == "active")
In numpy, `x == y` return a boolean vector of the same shape as x and y, comparing them element-wise.
olooney··on Is AI causing a repeat of frontend’s lost decade?
> undeterministic abstraction

I've seen people argue that LLMs will just add another layer to the top of the compiler stack: instead of writing code, we'll use English, and run it through a pipeline:

    English -> Rust -> ASM -> Machine Code
What's one more layer, right?

But what the author says about agents being "undeterministic abstraction" shows why that will never work.

Compilers rely on a concept called observational equivalence[1] to define when two programs are basically the same; this allows them to make changes under the hood like unrolling a loop or targeting another machine. Now, it turns out we know a lot about how and how not to do this, thanks to a logician named Frege who worked out exactly which properties a "definition" would need to have to count as a definition without becoming an axiom. In particular, that it should be "eliminable" and "conservative"[2]. In plain language, that a formal definition should always be able to be eliminated by rote string substitution, and that it shouldn't smuggle in any extra assumptions. When we talk about things like syntactic sugar[3] or hygienic macros[4], we are basically applying Frege's two conditions to programming languages.

LLMs are neither. They cannot reliably or provably go from the prompts they are given to the source code they generate, and they make a ton of implicit assumptions when they do so. There can never be any equivalence between two "prompts" in the same way that two programs can be equivalent modulo some level of abstraction. The whole process of starting from prompts is wildly nondeterministic, which is why the only pattern that works is to generate the code, review it, and test it, and then check it in and use that as the starting point for the next prompt.

Which is not to say that LLMs aren't useful for code generation; they clearly are. But they don't provide an abstraction that lets us get away from the details of actual code, and thanks to Frege we can understand why they never will.

I can say all this with such confidence because I did once write a wild little Python library that used a bunch of introspection to actually do this[5]. And it absolutely did not work in practice beyond toy examples.

[1]: https://en.wikipedia.org/wiki/Observational_equivalence

[2]: https://plato.stanford.edu/entries/frege/#ProDef

[3]: https://en.wikipedia.org/wiki/Syntactic_sugar

[4]: https://en.wikipedia.org/wiki/Hygienic_macro

[5]: https://github.com/olooney/fourth_gen

olooney··on A Visual Guide to Gemma 4
Incredibly detailed! The vision transformer stuff in particular is very useful to know. It's interesting that the token budgets are so much higher (up to 1120) than GPT, which uses 170 tokens per 512x512 tile. I wonder if that will lead to more granular spatial vision, something GPT struggles with.
olooney··on Chinese hackers use Anthropic's Claude
This isn't surprising; the majority of programmers are using LLMs, and Claude is pretty good for coding. Penetration testing is also a pretty good fit for an agentic loop - you run a tool, read the output, and decide on your next move, rinse and repeat.

In VSCode + GitHub Copilot, agent mode it can propose bash command to run, and when you confirm it runs it in a console and can see the loop, so it can fix errors immediately if any. It tends to go off the rails pretty quickly if things start going badly wrong, but it can complete simple tasks with supervision.

Over two years ago, when this LLM stuff was pretty new, I saw a demo that put ChatGPT in a loop with Metasploit that could crack some of the easy HTB challeges automatically - I remember thinking it was the single most irresponsible use of AI I'd ever seen. While everybody else was trying to sandbox these things for safety, this project was just handing it command line access to the tools it would need to break confinement.

It seems there's actually a whole bunch of similar tools these days, marketed as "automated penetration testing," such as Cybersecurity AI[1]. I used to think the whole cyberpunk "hackers can get in anywhere if they just type hard enough" trope was stupid because with cryptography the defender always has a huge advantage, but now we're looking at a world where AI is automating attacks at scale, while the defenders are vibe coding slop they have no idea how to secure, so maybe Gibson was right all along.

[1]: https://github.com/aliasrobotics/cai

olooney··on It's insulting to read AI-generated blog posts
I'm pretty sure my mistake was assuming people had read the article and knew the author veered wildly halfway through towards also advocating against using LLMs for proofreading and that you should "just let your mistakes stand." Obviously no one reads the article, just the headline, so they assumed I was disagreeing with that (which I was not.) Other comments that expressed the same sentiment as mine but also quoted that part did manage to get upvoted.

This is an emotionally charged subject for many, so they're operating in Hurrah/Boo mode[1]. After all, how can we defend the value of careful human thought if we don't rush blindly to the defense of every low-effort blog post with a headline that signals agreement with our side?

[1]: https://en.wikipedia.org/wiki/Emotivism

olooney··on It's insulting to read AI-generated blog posts
Here is a piece I wrote recently on that very subject. Why don't you read that to see if I'm a human writer?

https://www.oranlooney.com/post/em-dash/

Page 1 of 11Next →