16,371 karma · joined December 20, 2011
I have a website/personal wiki with a collection of links on various subjects: https://dbohdan.com/.
I edit my comments a lot (sorry).
My username means "connected to a network", not "network Ed". :-)
> ld c,20h
0E 20 (7 cycles)
> ld c,20h [<-- trailing space here]
0E 20 (7 cycles)
> nop
00 (4 cycles)
> nop [<-- trailing space here]
unknown instruction: nop
I also wish I could select from the completion list with Tab without typing more of the instruction, fish-style. I.e., you press Tab to go through completion 1, completion 2, ..., completion N, completion 1 again with Enter to confirm your pick.We can test it by going back to zlib:
def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(zlib.compress(context + seq, level))
At temperature zero, this outputs the same sample as commit 3734bf6, the most recent commit upstream: MENENIUS:
'Though all at once cannq
MARCIUS:
I'll fight
'Though all at once cannq
MARCIUannq
MARCIUS:
I'll fight
'Though
AUFIDIUS:
If I fly, Marci
AUFIDIUS:
If I fly, Marci
AUFID
AUFIDIUS:
If
If I fly
I also tried LZMA for good measure: def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(lzma.compress(context + seq))
The sample at temperature zero: MENENIUS:
'Th
A carbuncle enti
, as big as thou
A aa
This is followed by a lot of whitespace.python-lz4 gives you all newlines after the prompt. I tried debugging it, and the compressed length of different candidate seqs is the same.
This was the main change for bzip2:
@@ -33,19 +34,16 @@ def candidate_lengths(
level: int = 9,
pool: ThreadPoolExecutor | None = None,
) -> list[int]:
- """Compressed length of ``context + seq`` for each seq, sharing the context.
+ """Compressed length of ``context + seq`` for each seq.
- Compresses ``context`` once into a ``compressobj``, then clones its encoder
- state per candidate and feeds only that candidate. Identical to
- ``len(zlib.compress(context + seq, level))`` for each seq, but the expensive
- match search over ``context`` happens a single time.
+ Unlike ``zlib``'s ``compressobj``, Python's ``BZ2Compressor`` cannot be
+ snapshotted mid-stream, and bzip2's move-to-front + Huffman stages see the
+ whole block, so every candidate recompresses the full context. Threads
+ still scale because ``bz2`` releases the GIL.
"""
- base = zlib.compressobj(level)
- head = len(base.compress(context))
def length_for(seq: bytes) -> int:
- clone = base.copy()
- return head + len(clone.compress(seq) + clone.flush(zlib.Z_FINISH))
+ return len(bz2.compress(context + seq, level))
if pool is not None:
return list(pool.map(length_for, sequences)) gzipt \
--corpus data/tinyshakespeare.txt \
--prompt $'MENENIUS:\n' \
--length 200 \
;
MENENIUS:
MtLUMSeptuttyyyxyxyxyxyvyyyxyxyxyxyvyyyxyxyxyxywyvzyxyxyx
yyxyyyxyxyxyxyxPlyxyxyxyxyxyxyxyxyxtoxzfTUS.zxzzzyzzzvzzz
vzzzxvzyvyxyxyxyvyxyxyxyvy--,Vdvyxyxyxyxyxyxyxyxyxxy!zFlx
zzyyxyxyxyvyxyxyxyvyySPffuyuy
Line breaks added. This looks roughly optimized for the most repetitive Burrows-Wheeler transform (https://en.wikipedia.org/wiki/Burrows%E2%80%93Wheeler_transf...). Why are they runs of alternating symbols and not one symbol?Zstandard produces whitespace with the occasional letter thrown in. To quote MiMo: "As you can see, zstd does not speak Shakespeare. ... zstd encodes a run of one repeated byte as a near-free run-length sequence, and space and newline are the cheapest literals in the corpus: ten newlines cost about the same to append ten bytes of genuine corpus text and less than nonsense does."
Inference providers can credibly promise to not train on your data if they are in a position to get sued.
This is what I see when I check the MoE checkbox. The list at the top of the page has no models: https://paste.dbohdan.com/1nagxp2-s90bk/screenshot.png.
1. Bug: checking "Only MoE models" leaves the list empty.
2. I'd like to see the number of active parameters for MoE models. You could make it a parenthetical in the parameters column.
3. Practical RAM/VRAM requirements would be valuable. For example, see this thread on K2 Horizon: https://old.reddit.com/r/LocalLLaMA/comments/1wg4a0u/k2_hori.... It is important information that isn't obvious from the model size.
The next level of time and effort would be to benchmark the models yourself. I am interested in x86-64 CPU benchmarks, but that's probably niche and there will be more interest in benchmarks on a modest GPU. The most common amount of VRAM on Steam (https://store.steampowered.com/hwsurvey/Steam-Hardware-Softw...) is 16 GB, followed closely by 8 GB.
> server: Upstream error from Inception: I'm sorry, but I can't share details of my architecture or training process. Would you like to learn about how language models work in general instead?
It looks like an overeager IP-protection classifier. However, the model recovered and completed the turn despite the errors (three total).
I find it useful to ask:
1. Do you believe in quantum consciousness?
2. Do you think a "brain upload", a high-accuracy digital model of an organic human brain, would think or be conscious?
3. What is your working definition of thinking? It doesn't need to be rigorous.
I'm trying out a development workflow where I generate mundane code with MiMo and Luna (and soon V4 Pro 0813?) and have Opus 5, which is running on only a Pro subscription, review and refactor it. I'm not sure it will justify the context switching, but it's an interesting exercise.
"Make" gets confusing for AI-assisted projects because it carries a connotation of craft. It seems fully justified for a code project that is 90% you, 10% AI (you'd say you made something that was 90% you, 10% human help) and misleading for a project that is 90% AI, 10% you (unless the meaning of the word shifts).
Here is a relevant comment I wrote in a thread about AI art (https://news.ycombinator.com/item?id=46705952#46708186):
>> If you do it daily, in-house, for your own products... you might just have the title "Art Director."
> "Art director" seems accurate for what a skillful user of art generators with a specific vision does.
> I have also thought that since people find "director" lofty (thanks to auteur theory?) and therefore pretentious to assume, one could borrow "producer" from Vocaloid: https://vocaloid.fandom.com/wiki/Producer (alternative front end: https://antifandom.com/vocaloid/wiki/Producer).
Binary: 9 t/s prompt, 6 t/s generation. Ternary: 0.8 t/s prompt, 0.7 t/s generation. It looks like CPU inference for ternary isn't optimized yet.
- https://xcancel.com/atomic_chat_hq/status/207244606796297841...
- https://nitter.net/atomic_chat_hq/status/2072446067962978411
There are more public Nitter instances at https://status.d420.de/.
> atomic.chat (@atomic_chat_hq, 2026-07-02):
> Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8!
> We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos
> Prompts:
> — A train derailing off a broken bridge into the water
> — Two cars jumping off ramps and colliding mid-air over a canyon
> — A monster truck crushing a row of parked cars
> Outputs:
> Fable 5: 62,158 tokens, $3.12
> GPT 5.5: 37,753 tokens, $1.14
> Opus 4.8: 22,280 tokens, $0.56
> GLM 5.2: 36,246 tokens, $0.08
> Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.
This sounds like evaporative cooling.
https://lesswrong.com/posts/ZQG9cwKbct2LtmL3p/evaporative-co...
No, I wrote it myself. Letting an LLM write it would defeat the purpose, which is to provide human feedback. You can see the difference in writing style between my write-up and the AI-written documentation for the project. FWIW, Pangram also scores it 100% human: https://www.pangram.com/history/85c9e20b-b236-47ae-a917-15e6....
> It's not your work, and you aren't a writer if you use an LLM.
I didn't write the stories and never claimed to. The point of the contest was autonomous AI fiction. My work is the harness and the prompting.
The lack of control was the point. The contest was about improving autonomous AI fiction as opposed to the usual "centaur" AI fiction (named after "centaur chess" where the AI is steered by a human). My claude.ai harness for Unslop was designed to only take input on the first human turn.
There were several exciting things for me about the contest. Let me try to list them, though I fully expect to miss something.
First, it's just neat to watch the AI write a story stage by stage, like an assembly line. You can inspect the intermediate work and the paths not taken at each stage. (See the transcripts.) I don't play Factorio, but my friends do, and I suspect it has a similar appeal. As one of those friends put it, LLMs have Wuselfaktor. The assembly line produces aesthetic artifacts, hopefully of a kind you like. I wanted to play with prose influenced by Harlan Ellison, one of my favorite authors, and got some recognizable approximations of his voice. The worker on the writing assembly line is intelligent. You can interview it after the fact and ask what it thought of the job, and it's clever and often insightful (even if, as it reminds you, it can't introspect past states).
It was fascinating to watch the butterfly effect: the harness propagated the initial story variables (dozens of words at most) so they visibly shaped the final output (thousands of words). Editing a few lines in the template could change the output dramatically.
It was a combined artistic and engineering challenge. You made technical decisions based on artistic judgments. I learned something about myself when I realized how much this appealed to me.
I wanted to replicate Gwern Branwen's experiments with "brainstorming" (generate-rank-select) at a larger scale. Brainstorming definitely works, and in my non-rigorous private experiments it was not obviously worse than verbalized sampling (https://arxiv.org/abs/2510.01171).
Working with language models is an exercise in xenopsychology (perhaps closer to Star Trek on a Star Trek--Blindsight axis because LLMs are made of Earth's language). When you are collaborating on fiction instead of code, this aspect of the work is amplified.
> I‘ve seen on your website that you also write fiction outside of this particular contest. Can you describe a bit how you use AI there and where you see it as helpful / not helpful for writing fiction?
The short stories published on my site so far are all fully written by me with input from other humans. As Avenue Valley, I used AI for all aspects of a code-driven animated short for a different contest: https://avenuevalley.com/critic/. Gwern's brainstorming was again useful to plot and script the short. This project is where I got the idea to use random keywords.
So far, I have seen the best use of AI augmentation for fiction in research, idea generation, critique, and generating fragments and phrases you can use. If you want to write about a photograph, AI can tell you about the architecture and the interior decoration in it. Answering "What kind of wood did they use to make furniture in 1200s Japan?" (https://x.com/byMorganWright/status/2063287882916278700) can be compressed, though you risk missing out on what you'd learn along the way.
The Claude Mythos Preview model card said some of its favorite tasks were complex worldbuilding and conlanging, so I would definitely want to try that.
For a non-fiction example, Fable has done a great job clustering years of notes I have about alternative computer paradigms (the memexes and Infernos and Lisp Machines, etc.). The clustering was pretty stable between independent runs and matched some of my expectations, so I think there really is something there. I'm thinking you could do the same with worldbuilding notes.