Did you check MiMo correctly performed this unfamiliar task before posting this comment?
This was the main change for bzip2:
@@ -33,19 +34,16 @@ def candidate_lengths(
level: int = 9,
pool: ThreadPoolExecutor | None = None,
) -> list[int]:
- """Compressed length of ``context + seq`` for each seq, sharing the context.
+ """Compressed length of ``context + seq`` for each seq.
- Compresses ``context`` once into a ``compressobj``, then clones its encoder
- state per candidate and feeds only that candidate. Identical to
- ``len(zlib.compress(context + seq, level))`` for each seq, but the expensive
- match search over ``context`` happens a single time.
+ Unlike ``zlib``'s ``compressobj``, Python's ``BZ2Compressor`` cannot be
+ snapshotted mid-stream, and bzip2's move-to-front + Huffman stages see the
+ whole block, so every candidate recompresses the full context. Threads
+ still scale because ``bz2`` releases the GIL.
"""
- base = zlib.compressobj(level)
- head = len(base.compress(context))
def length_for(seq: bytes) -> int:
- clone = base.copy()
- return head + len(clone.compress(seq) + clone.flush(zlib.Z_FINISH))
+ return len(bz2.compress(context + seq, level))
if pool is not None:
return list(pool.map(length_for, sequences))(I honestly don’t know is gzip does something different when presented with two chunks as opposed to one, or, if it does, if bz2 has equivalent behaviour - but the difference in the code did stand out to me, and it does seem related to ‘extending the token sequence’)
We can test it by going back to zlib:
def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(zlib.compress(context + seq, level))
At temperature zero, this outputs the same sample as commit 3734bf6, the most recent commit upstream: MENENIUS:
'Though all at once cannq
MARCIUS:
I'll fight
'Though all at once cannq
MARCIUannq
MARCIUS:
I'll fight
'Though
AUFIDIUS:
If I fly, Marci
AUFIDIUS:
If I fly, Marci
AUFID
AUFIDIUS:
If
If I fly
I also tried LZMA for good measure: def length_for(seq: bytes) -> int:
- return len(bz2.compress(context + seq, level))
+ return len(lzma.compress(context + seq))
The sample at temperature zero: MENENIUS:
'Th
A carbuncle enti
, as big as thou
A aa
This is followed by a lot of whitespace.python-lz4 gives you all newlines after the prompt. I tried debugging it, and the compressed length of different candidate seqs is the same.