It is fair to complain that subscription AI usage limits are opaque, but that is at best tangentially related to them being subsidized relative to API prices.
178 karma · joined April 22, 2026
It is fair to complain that subscription AI usage limits are opaque, but that is at best tangentially related to them being subsidized relative to API prices.
At the time that essay was written, Make was the best runner available, and the only portable one. There are better options now.
When one's uncompressed RINEX files are 200 GB/day, it's worth some effort to shrink the files, especially if it means they are faster to read.
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
Arguably so, but then one would lose the ability to natively name 16-bit integer types because "short" would be 32 bits.
An earlier comment addresses x86-64. AArch64 (pedantically, the A64 instruction set used for AArch64's 64-bit execution mode) is similar, in that addresses are 64 bits wide but ALU instructions typically encode a width bit, called "sf", that selects either 32- or 64-bit data registers and arithmetic. See, for example, https://arm.jonpalmisc.com/latest_aarch64/add_addsub_ext .
Maybe an LLM did the search for what expressions to put in the blocks vs tail of the vectorized collision detection, but that seems like a good place to use an LLM as long as all of the expected bit relations are verified as checked properly.
Well priced when compared to other models of similar price, eh?
Are we allowed to call this slop, even if the output is not directly from an LLM?
How should a reader infer the bye order for a UCS-2 or UTF-16 file without a BOM? It seems like one would have to read until finding a code point that would be illegal under one ordering (but files might not include such a code point).
Similarly, a UTF-8 BOM is a useful flag to distinguish UTF-8 from other text encodings. You are right that the ambiguity goes away if those other encodings do, but people don't want to rewrite their legacy files. Some people don't want to use two bytes for common non-ASCII characters, so they are really attached to ISO-8859 or Windows-1252 or koi8r or whatever. CJK languages have their own encodings that are more efficient for their languages. UTF-8 is great for English speakers, but it's a compromise for everyone else, so they might reasonably want incompatible systems for their own use. UTF-8 BOM is a good "magic" sequence to detect encoding as long as people have non-UTF-8 files.
If you have billions of vectors, you should use a dedicated implementation: nearest-neighbor algorithms in high-dimensional space can be tricky, and there are a lot of trade-offs. My case is amenable to a simple implementation because I don't need huge scale. That's also why I did not need or want an out-of-process server.
I thought my explanation for the second one was already clear: "Programs should avoid, if possible, [doing X] too often" is a truism because "too often" implies that some reduction is possible. One should leave out the ", if possible," -- although deciding what is "too often" can be challenging and sometimes a matter of taste. (Is reducing allocation frequency by 5% worth doubling the CPU usage or memory usage or code complexity? Maybe in some cases, but often not.)
My baseline was (vibe coded) full-text search with SQLite, and we landed on long chunks with overlap, with a really simple nearest-neighbor search: quantize the index to a sign bit per scalar, which makes it super cheap to estimate dot products, then calculate better (8-bit quant index times native precision for the query) dot products to sort the top documents. A vector database would make sense for a much larger corpus, but I currently have fewer than two million rows. Claude vibe-coded it to use an OpenAI-speaking local inference server and made semantic search an optional augmentation for the full-text search.
Which advice do you stand by? Obviously, very short content doesn't need chunking, so let's consider a document that fills 75% of the input context.
When chunking, your cost overhead (per token) goes up as the number of new tokens per chunk goes down. That's an argument for longer chunks, although the averaging/smearing point argues for not going too long.
Embedding calculations are effectively prefill: on my cheapo local inference system (32 GB AMD R9700 + 8 GB AMD RX 7600), the older 8 GB card goes about 80% as fast as the bigger card for Qwen3-Embedding-4B (a bit over 19 chunks/second on my usual corpus, blog posts+comments that are mostly well under 32K tokens). So I would suggest that anyone who is limited by CPU embedding models could benefit from even a small local GPU.
For your blog post, I would suggest an explanation of the chunking modes, either in the blog post or as a hyperlink to the docs about them. "truncate" and "sentence" are fairly clear, whereas the others are not. (If "mean" just computes the mean of the embeddings, that seems like a poor choice. The arithmetic at https://www.johndcook.com/blog/2026/09/16/coffee-milk-latte/ might work for single words, but seems likely to break down at the document level. "recursive" and "fixed" are opaque, at least to me.)
If/when I index my team's documents, I will consider a content-aware chunking that fits as many sentences, paragraphs or sections as possible into each chunk, with overlap determined by the level at which the chunk finishes. Content-agnostic chunking is easier to code and more generic, but indexing should respect a document's internal structure.
For languages like English, there's also usually a lot of redundancy within a text, so 512 tokens might not give a very clear indication of the context. Lots of documents have similar introductions (like "#include <foo.h>\n") that make short contexts and truncation particularly harmful.
Also, "Nothing in the document past that point can ever be retrieved, and nothing anywhere told you." This is user-hostile behavior, even if they didn't want to admit to users that the auto-embedding support was poor.
Finally, the paragraph later on about truncation being "what you already have" reads like Claude talking to the developer, not like a vendor talking to users. But sure, maybe this is a good default for a database searching page titles, chat logs and Xeets?
And I have to think hard to guess what it probably meant by "Comparability is a property of the reference the results are traceable to, not of the number". I don't think "[X and Y] are different scopes by canonical bytes" even makes sense.
A year ago, LLMs were not useful for me as a programmer. Now they are: the models are better, they can use long contexts more effectively, and the harnesses are better at helping the models. Nowadays my job is mostly not programming, but LLMs let me organize and prepare tools in spare time rather than needing days or weeks of attention. I would not trust them on a 200k+ line project -- and Claude Opus 5 has issues even on 50k LOC projects -- but they absolutely can help given good direction and a narrow enough scope.
Their claim is consistent with Google's experience, but somewhat at odds with (say) that of FreeBSD, OpenBSD and NetBSD.
I've been reading John D. Cook for years (maybe decades? "The Endeavour" is one of my oldest bookmarks), and this post was no more written by AI than his oldest posts.
People who study this. This glitch was caused by unusual, structured, widespread distortions in the ionosphere. There's no known mechanism that would cause enough distortion to make the errors 100 times as large. There would also be other symptoms of any unknown cause.
https://claude.ai/share/5cc4163e-9cbf-47cd-9749-4bb55bca42fe if you don't mind discursions into safety of aviation navigation systems. Caution: Do not take the error bounds in the second half as relevant to errors for unaugmented single-frequency users; those users will see much bigger errors.
GPS has always had the P(Y) code, which is an encrypted signal, at a higher chip rate than the C/A code that has been broadcast on L1 forever. The P(Y) code has its own interesting history of (semi-)codeless processing. If the US military is lucky, the history books will close on that by 2030: https://www.gps.gov/codelesssemi-codeless-gps-access-commitm... . (Currently, 21 GPS satellites broadcast L5.)
If one has a lot of patience, one can read the details of how it works at https://elibrary.icao.int/product/299828 (along with SBAS, the ICAO name for what WAAS does; the constellations they can augment; GRAS, a weird hybrid between GBAS and SBAS that apparently went nowhere; ILS; VOR; DME; and VHF marker beacons). Annex 10 Vol I probably doesn't say explicitly that triple redundancy is required, because if you had a single receiver with an MTBF of something like 10 billion hours then you could use just that one.
This particular 30-foot error should go away almost entirely for future receivers that implement DFMC SBAS or ARAIM.