HNHacker News
TopNewBestAskShowJobs

kqr

19,093 karma · joined January 2, 2015

Quant, systems thinker, anarchist.

I write at https://entropicthoughts.com

My inbox is hn[at]xkqr.org

submissionscomments
kqr··on Ephemeral Testing
It's for chiseling out the design, not verifying the functionality.
kqr··on Ephemeral Testing
This is an interesting idea. It's effectively what a good software engineer already does in their head, except it's doing it with a real compiler.

Richard Gabriel wrote something that has really stuck with me:

> Abstractions must be carefully and expertly designed, especially when reuse or compression is intended. However, because abstractions are designed in a particular context and for a particular purpose, it is hard to design them while anticipating all purposes and forgetting all purposes, which is the hallmark of the well-designed abstractions.

This is one of my favourite quotes on abstraction, because “anticipating all purposes and forgetting all purposes” is such a good summary of what goes into abstraction design.

A language model in an agentic harness cannot (yet) do this at the same level as a good software engineer, but the advantage they have is speed of token generation, so they can actually build the things the engineer tries to imagine, and verifying those is easier. Very cool!

kqr··on We are going to kill "unalive"
I agree with the overall point, but I think a large fraction of the usage of these terms is about in-group signaling rather than conscious efforts to appease the algorithm.

What I mean is that a few people did figure out that appeasing the algorithm generates more views, and these are the people that get many followers. The followers themselves probably don't think that deeply about it; they just want to be like the accounts they follow (because they think everyone else in their social circles want to), and copy their use of language.

It's probably similar to how I started using semicolons more after reading Mary Shelley, except more gross.

kqr··on Infidel goes wild
If you can build something like that and get it to work, I'd be impressed and very excited! People before you have failed.
kqr··on Infidel goes wild
What makes you believe people haven't tried? I have yet to find an example that actually works.

The language models have a tendency to not stick to the script (inventing world state that does not exist), be too helpul (provide spoilers), or get stuck on the same confusions as a human would.

kqr··on Infidel goes wild
Someone commented that it's fun to see Zarf is still around, so I want to take the opportunity to say that the entire community around text adventures is likewise still thriving, most visibly on the IntFiction.org forums.

People seem to look at text adventures as relics of a specific era, and it's true that it's hard to make money off of them these days, but there are several high-quality releases per year, and the quality seems to only go up. These games are often playable for free, because they're released for competitions that stipulate the submission must be free.

-----

It can be a little hard to get into text adventures, because the parser interface is designed to accept inputs according to a specific convention (all three of "cut tapestry", "cut the tapestry", "look behind tapestry. cut it." would work, but it's not a guarantee "slice tapestry" would, and almost never would "cut hanging fabric" work).

This convention must be learned. Despite the marketing in the 1980's, it does not come naturally to people. Learning that convention means forcing yourself through a few games. But once you have learned that convention, an ocean of games open up to be played, many of them very enjoyable. I highly recommend taking the time to grind through a few games to get familiar with the convention.

I recommend not starting with the old classics because they were in part made to be frustrating and difficult and played over a long time. Instead, look at modern text adventures that follow more recent advances in game design. A good source is the Interactive Fiction Database, with e.g. a search like this: https://ifdb.org/search?searchfor=tag%3AParser+published%3A2...

-----

Text adventures are also a nice way to dip one's toes into game development. They force the developer to think about game design without having to spend time or money on graphics or music.

We are spoiled choice when it comes to development systems for text adventures these days. Most popular development systems consist of two parts: programming language, and separately, standard library. The library half contains generic definitions of things like rooms, scenery items that cannot be picked up, clothing that can be worn, animals that can be interacted with but not talked to, etc. (The exception is perhaps Inform 7, which tightly couples its compiler to the standard library. In the case of the others, one is fairly free to swap out the standard library for anything else.)

Examples include:

- Inform 7 lets developers write code in a weird subset of English. Some people swear by it, but people who come from programming backgrounds sometimes find it hard to learn. It does have a great IDE though, with automapping, an index of all objects, a "Skein" that records a game tree and lets one play it back and check for inconsistencies, etc.

- Inform 6 is a completely different language, reading more like object-oriented code. There is an alternative standard library called PunyInform which lets you make tiny games that run well on old hardware. (But which can also be used as a creative constraint.) It doesn't have the same level of tooling, but it's been around for a long time and is well supported on many platforms. The compiler is a simple, dependency-free C program one runs `cc -O2 -o inform *.c` to compile.

- Dialog is a newer language and library for making text adventures that's stable, but receives frequent updates still. It's inspired by the rules-based approach of Inform 7, but uses a more consistent, programmer-friendly syntax. It's got a kick-ass debugger that reloads definitions live and supports a highly iterative development cycle. There's also a recent, partly vibe-coded "Dialog Tool" made for it that gives it an Inform 7-like Skein functionality.

Then there are others I haven't tried, like TADS, which, like Inform 6, is closer to traditional programming. The original Infocom games were written in Lisp macros that got translated into the Lisp-like ZIL language. These days, we have an extremely well-documented open source implementation of ZIL, too.

In case anyone is curious, I have written two articles introducing Inform 6 and Dialog respectively as general programming languages outside of their use for text adventures. (But this is not a recommended use case. Just a fun way to learn where the language ends and the standard library picks up.)

- https://entropicthoughts.com/advent-of-code-on-z-machine

- https://entropicthoughts.com/advent-of-code-in-dialog

kqr··on We're going to need default hard budget caps on pretty much everything
This goes beyond dollar charges. For production code to be reliable, everything needs to have a hard limit.

Queue lengths, request sizes, response wait duration, message payload size, authentication attempts, allocation rates -- there's always some upper number beyond which the system is so messed up you'd rather it crashes.

> An argument against this is that businesses don’t want their hosted applications to start throwing errors because some budget was exceeded. I expect that most businesses and individuals would prefer errors to a surprise $10,000+ bill.

Indeed. If you want a surprise $10,000 bill that's still not an argument against a hard cap -- just set it at $9,999,999 instead, or wherever you don't want the surprise bill. There's always a number that indicates something has gone insane. There's always a sensible upper limit to any operation.

kqr··on Why the Bronze Age Collapsed
Most of what I know about life in the bronze age comes from The Wise-Woman's Dog[1], a slightly fantastical game that leans into what life was like for regular people at the time. It's made by someone who researches that sort of thing for a living, and is full of interesting footnotes too.

It's a fun experience and a useful way to make a place and a time really feel alive.

(Not strictly on topic but having played it really lends texture to TFA and other comments posted here.)

[1]: https://ifdb.org/viewgame?id=bor8rmyfk7w9kgqs

kqr··on Backblaze drive stats for Q2 2026
I published the same criticism in relation to one of their earlier summaries. Backblaze responded with something along the lines of (paraphrasing liberally)

> We're giving you this data for free and providing a brief summary of what you should expect to see in it. We do not get paid to teach our readers advanced statistics, nor is that what we want to do. We encourage you to take the raw data and explain to people how you think it should be analysed.

which seems like a fair response. I'm very, very grateful that they release the data freely the way they do. It reflects well on them and more companies should do that.

kqr··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
I think the distinction is between "improving the g factor" and "adapting a given level of g to perform certain types of work better".
kqr··on When did Google get so weird?
I think you overestimate how much people want an actual truthful answer. I think people are often looking things up to silence a nagging voice in the back of their heads, or settle an argument to be able to move on. In those cases, no decision is riding on the answer, and a quick answer, even when factually incorrect, could be argued to serve the purpose just as well as (or better than) a slow but more correct answer.

I realised this for myself when I read Sivers on being without internet[1]. He discovered a process to silence the nagging voice without searching:

> When I’m yearning to search, I ask myself why.

> - What answer am I hoping to hear?

> - What answer would be a surprise?

> - What would I do in each case?

[1]: https://sive.rs/off23

kqr··on What is the size of Yemen? (2024)
This reminds me of what Deming used to say: there is no correct value for any measurement, there are only the outcomes of methods chosen to measure something.
kqr··on What is the size of Yemen? (2024)
Imagine how much of published science is statistically significant only due to clerical errors and mistakes in calculation...

(See e.g. The misreporting of statistical results in psychology journals; Bakker, Wicherts; Behaviour Research Methods; 2011.)

kqr··on The Board Game of the Alpha Nerds (2014)
I've never played Diplomacy, but I try to turn every other board game I play into Diplomacy[1] and it's great fun. For me, the reason to play physical board games over other kinds of games is precisely the in-person negotiation and problem-solving that gets people closer together.

[1]: https://entropicthoughts.com/free-banking-monopoly

kqr··on My weird new hobby: Wandering around Tokyo on Google Maps
I have done this for places I visited a lot in my early childhood but since have never returned to. It's a bit eerie to see the memories come back to life. Often I remember just enough of the place to be able to find my way between what were the main attractions to me at that age. Somehow, despite forgetting all the details of the path, something in the back of my head compels me to follow the right road even in confusing intersections. I find it fascinating.
kqr··on Artificial intelligence now beats some of the best human forecasters
Providing liquidity?
kqr··on Pangram – AI detector for text and images
I didn't want to shell out $20 a month for the general thing, so I spent $90 on data collection and built my own for code comments specifically. It runs locally in your browser with a relatively small classification model trained on old-school stylometric features. You can try that before turning to Pangram for uncertain cases, if you wish.[1]

It's easy to get high accuracy numbers if you're testing on long (50+ words) texts. Much harder when the documens are short, as code comments tend to be.[2]

[1]: https://xkqr.org/aicomment

[2]: https://entropicthoughts.com/better-ai-comment-classifier

kqr··on Pangram – AI detector for text and images
If I were your editor, I'd strike the adverb. It weakens the sentence.
kqr··on Let's make quality the norm again
> Quality brands are economically incentivized to sell-out [...] because customers take some time to catch on, and short-term financials are too often more important

That probably happens a lot, but there's another mechanism by which it could happen, that requires no malice: the brand gets a quality reputation, which increases demand for its products, and they are unable to produce quality products at the rate needed to satisfy that demand. Then it's easy to reason that it'd be a net benefit for everyone to loosen the quality requirements a tiny bit, in order to satisfy the demand of more customers. People used to the quality will barely notice any difference, and then twice as many people can get that experience.

That works once, but when applied again, and again, and again, and again it becomes a problem, even with zero ill intent.

kqr··on Rope, twine and thread: Invisible technologies of the Stone Age
Indeed. I maintain a mental list of underrated conceptual leaps throughout history and rope is in there. Also on the list is writing, statistics, and the world wide web.

The wheel is overrated. It isn't that innovative (having been invented many times throughout history by various cultures) and to be properly useful it needs environmental adaptations. Controlled fire is neither under- nor overrated, I think.

kqr··on Rope, twine and thread: Invisible technologies of the Stone Age
Not mentioned in TFA: cordage is a critical component in slings, an early, powerful ranged weapon.

With some sort of pulley, rope allows a single human to excert large amounts of force trough leverage, often in ways more convenient than other forms of mechanical leverage. (Such as actual levers.)

Rope can also be used for 1:1 power transfer.

kqr··on What will our economic future look like?
Aren't "not rich" and "poor" very different things?
kqr··on How well do agents use test/verification techniques?
This surprised me too. I suspect the robots only got a "use mutation testing" prompt and independently decided that it must mean manual mutation testing, or possibly that they ran in a sandbox where an automated mutation testing tool was not installed.

But I'm surprised Dan Luu didn't make a remark about this. Surely he must know the difference!

kqr··on How well do agents use test/verification techniques?
I believe it would. Automated mutation testing is the test coverage metric that is nearly impossible to cheat.
kqr··on How well do agents use test/verification techniques?
I thought it pretty clear that the code was generated by the same agent that received the testing prompt, so there were no constraints on the code structure, and the testing strategy was known at the time the structure was generated.
kqr··on How well do agents use test/verification techniques?
This is answered by the large plot early in TFA. All tests were made using the same model, and adding no special instructions at all provided performance above the average.
kqr··on How well do agents use test/verification techniques?
I don't know what I'm most impressed by: (a) the testing expertise, (b) the time spent looking into the reasoning mistakes generated by LLMs, or (c) the insane amounts of money this must have cost!
kqr··on Show HN: Fly By – retro biplane flying game
Indeed. Trim wheels when rolled upward trim for higher speed, i.e. nose down.

Wait, this agrees with GP's experience but not mine. Did OP invert it?

kqr··on Show HN: Fly By – retro biplane flying game
Like many side scrolling airplane games, the object controlled in this game behaves a lot more like a rocket than an airplane. I suggest learning a little about the physics of airplanes[1] to improve the feeling of flying.

I made a prototype[2] once (not working on mobile, unfortunately) to show what can be done, and it takes getting a lot of details right![3]

[1]: http://av8n.com/how/

[2]: https://xkqr.org/flightle/

[3]: https://entropicthoughts.com/sidescrolling-flight-simulator

kqr··on Show HN: Fly By – retro biplane flying game
What is the wind dynamics? I would have expected to climb when the headwind increased but I didn't. I also didn't manage to stall, only invert when I went out of bounds.
Page 1 of 34Next →