HNHacker News
TopNewBestAskShowJobs

imurray

1,578 karma · joined November 22, 2009

https://iainmurray.net/

https://imurray.bsky.social/

https://mastodon.social/@imurray / @imurray@mastodon.social

https://twitter.com/driainmurray

Research interests include statistics and machine learning.

submissionscomments
imurray··on A simple way to scale pixel art games
See also:

https://www.scale2x.it/

https://johanneskopf.de/publications/pixelart/

imurray··on Show HN: Offline audiobook from any format with one CLI command
> And do you know a good speech to text model?

OpenAI's whisper, code+model are available, and multiple projects have built on it. You could try this wrapper: https://github.com/m-bain/whisperX -- or for short utterances on a smart-phone https://github.com/futo-org/whisper-acft

imurray··on Microtome
Machine learners might think of "Microtome publishing", who used to print JMLR (the Journal of Machine Learning Research -- community run and free for all authors and readers). It turns out they don't bother to publish on paper any more:

http://www.mtome.com/Publications/JMLR/jmlr.html

imurray··on MobileLLM: Optimizing Sub-Billion Parameter Language Models for On-Device Use
The paper says they tried that: https://arxiv.org/abs/2402.14905

Deep link to the relevant snippet in html version: https://ar5iv.labs.arxiv.org/html/2402.14905#S3.SS5

"So far, we trained compact models from scratch using next tokens as hard labels. We explored Knowledge Distillation (KD)... Unfortunately KD increases training time (slowdown of 2.6−3.2×) and exhibits comparable or inferior accuracy to label-based training (details in appendix)."

imurray··on Building a data compression utility in Haskell using Huffman codes
Nicely done; thanks.
imurray··on Building a data compression utility in Haskell using Huffman codes
That code in the EDIT is suboptimal. It doesn't saturate the Kraft inequality. You could make every codeword two bits and still encode 4 symbols, so that would be strictly better.
imurray··on Building a data compression utility in Haskell using Huffman codes
> It would be interesting to see a uniquely decodable code that is neither a prefix code nor one in reverse.

More interesting than I thought. First the adversarial answer; sure (edit: ah, I see someone else posted exactly the same!):

    a 101
    b 1
But it's a bad code, because we'd always be better with a=1 and b=0.

The Kraft inequality gives the sets of code lengths that can be made uniquely decodable, and we can achieve any of those with Huffman coding. So there's never a reason to use a non-prefix code (assuming we are doing symbol coding, and not swapping to something else like ANS or arithmetic coding).

But hmmmm, I don't know if there exists a uniquely-decodable code with the same set of lengths as an optimal Huffman code that is neither a prefix code nor one in reverse (a suffix code).

If I was going to spend time on it, I'd look at https://en.wikipedia.org/wiki/Sardinas-Patterson_algorithm -- either to brute force a counter-example, or to see if a proof is inspired by how it works.

imurray··on FUTO Keyboard
This open source android keyboard supports pinyin Chinese input: https://github.com/osfans/trime
imurray··on AES-GCM and breaking it on nonce reuse
You're right, GHASH has 128 bits output. I'd wrongly assumed it was 96 from my quick reading of the conclusion on p29:

> unless an implementation only uses 96-bit IVs that are generated by the deterministic construction: The total number of invocations of the authenticated encryption function shall not exceed 2^32, including all IV lengths...

Which would be the conclusion if the hash was always reduced to 96 bits. What it actually goes on to say though is:

> For the RBG-based construction of IVs, the above requirement, in conjunction with the requirement that r(i)≥96, is sufficient to ensure the uniqueness requirement in Sec. 8

So, it's a bound. If there are at least 96 random bits, it should all be ok. But it strangely leaves open the possibility, without saying either way, that a longer iv, with r(i)>96 random bits might allow generating more iv's. As you point out, it will depend on the properties of GHASH (and potentially on how the result is used downstream from there). At this point I don't know, but the spec says:

> For IVs, it is recommended that implementations restrict support to the length of 96 bits, to promote interoperability, efficiency, and simplicity of design.

So personally, not being an expert, I'd follow the advice and use 96 bit iv's. And if using random iv's, re-key before using the key a billion times. I'd certainly want a reference to a careful analysis before assuming that I could ever use an AES-GCM key with longer random iv's more than that.

imurray··on AES-GCM and breaking it on nonce reuse
> When I use AES-GCM I just use a bigger nonce and use a random one.

I don't think nonces bigger than 12 bytes will help. My quick reading of the AES-GCM spec is that when using a nonce that's not 96 bits (12 bytes), it is hashed to 96 bits. So either the nonce (called iv in the spec) is carefully constructed from a counter and set to exactly 96 bits, or the number of invocations is limited. The spec still restricts use of a key to 2^32 total uses for random nonces of any bigger length (resulting in a re-use probability of about 1e-10):

https://nvlpubs.nist.gov/nistpubs/Legacy/SP/nistspecialpubli...

imurray··on NanoGPT: The simplest, fastest repository for training medium-sized GPTs
See also (from the same author) https://github.com/karpathy/llm.c — "LLMs in simple, pure C/CUDA with no need for 245MB of PyTorch or 107MB of cPython."
imurray··on Statement on CVE-2024-27322
> A few years back I have heard from a lot of people working in ML communities that they are surprised that `numpy.load` is able to execute arbitrary code.

This is correct, before version 1.16.3 (April 2019) `numpy.load` was unsafe by default, unless explicitly specifying `allow_pickle=False`. However, to be clear, that unsafe default was then fortunately changed. Loading numpy arrays with `numpy.load` should now be safe (unless there are yet-to-be-found bugs in that code).

imurray··on Claude 3 beats Google Translate
Google translate uses TPUs, and has done since they swapped to neural models: https://cloud.google.com/blog/products/ai-machine-learning/a...
imurray··on The Man in Seat 61
I've found advice from seat61.com useful a few times.

He recommends https://raileurope.com/ -- When I used them back when they were loco2, I was impressed by the customer service. I needed to change my train ticket, so emailed them and they sorted it all out for me with minimum fuss, emailing me replacements. At the time it was a lot easier than dealing with the local railway companies in countries where I didn't speak the language. I don't know if they would be as good now that they aren't a startup though.

History of loco2/raileurope: https://www.seat61.com/websites/who-are-raileurope.htm

imurray··on Library of Juggling
It's nice how it breaks down the patterns, and explains everything step by step.

If you would like 3D in-browser animations of a subset of passing patterns, you could look here: https://passist.org/patterns These only have terse generated explanations, because they all follow a common theory, which is https://passing.zone/4-handed-siteswap/

There's a community of jugglers who pass clubs who now document their new patterns with videos (and notation) at https://passing.zone/ rather than making animations.

For another "library" of sorts, here's a rather random assortment of (mostly) juggling related documents: https://jugglingedge.com/pdf/

imurray··on Mic Test
Thanks! Adapted into a single command-line version for an alias/script that toggles:

    pactl unload-module module-loopback 2>&1 | grep -q 'Failed' && pactl load-module module-loopback > /dev/null
I find the browser version useful too, as sometimes problems are browser-specific. E.g., the snap version of firefox has messed up permissions.
imurray··on Mic Test
This is the quick hack I use to check my microphone is working with my browser: https://homepages.inf.ed.ac.uk/imurray2/tmp/echo_recorder.ht...

It's a self-contained page, using just javascript without sending any data anywhere. You can take a copy of the .html file and use it locally, or host it somewhere you control.

imurray··on Ubuntu Pro Shenanigans
As a reward for installing ESM/Ubuntu Pro, that version of ffmpeg will break vlc: https://bugs.launchpad.net/ubuntu/+source/ffmpeg/+bug/204274...

Manually upgrading the source .deb to the (ABI compatible) ffmpeg 4.4.4 and compiling gets all the security fixes and seems to work. It looks like backporting just the security fixes was done incorrectly in Ubuntu's ESM release.

My main concern with ESM/Ubuntu Pro is that patches don't get as much community scrutiny and help as other updates. In this case it's broken a fairly major piece of software on what's supposed to be a super-stable LTS release.

imurray··on Yorick is an interpreted programming language for scientific simulations
> Although length-1 axes already being in the data sounds worrying: you have an array that doesn't depend on an axis, but there's an axis anyway to show where it would go if it did?

Example:

    A / A.sum(axis=3, keepdims=True)
Make all of the vectors along axis=3 sum up to one. There are other ways of doing it, but this way seems fairly clear to me. The shape of the denominator is the same as the numerator, except for a 1 in position 3. Unfortunately we have to specify `keepdims`, because the default of `False` removes the dimension being summed over, which doesn't work in general. `keepdims=True` is the behavior in Matlab/Octave, so the example becomes

    A ./ sum(A, 4)
with 4=3+1 because Matlab is 1-based like Fortran.
imurray··on Yorick is an interpreted programming language for scientific simulations
Thanks for the pointer. I can believe that a language that looks so different will find that different patterns and primitives are natural for it.

My experience from writing a lot of array-based code in NumPy/Matlab is that broadcasting absolutely has made it easier to write my code in those ecosystems. Axes of length 1 have often been in the right places already, or have been easy to insert. It's of course possible to create a big mess in any language; it seems likely that the NumPy code you saw could have been neater too.

In machine learning there can be many array dimensions floating around: batch-dims, sequence and/or channel-dims, weight matrices, and so on. It can be necessary to expand two or more dimensions, and/or line up dimensions quite carefully. Einops[1] has emerged from that community as a tool to succinctly express many operations that involve lots of array dimensions. You're likely to bump into more and more people who've used it, and again it seems there's some overlap with what Rank does. (And again, you'll see uses of Einops in the wild that are unnecessarily convoluted.)

[1] https://einops.rocks/ -- It works with all of the existing major array-based frameworks for Python (NumPy/PyTorch/Jax/etc), and the emerging array API standard for Python.

imurray··on Yorick is an interpreted programming language for scientific simulations
I don't know APL, but from what I found quickly, APL's "conformability" rules[1] only seem to be a subset of the Yorick/NumPy broadcasting? Here's a Python example of broadcasting that APL doesn't seem to support(?): subtracting arrays of shape (3,1,7,5) and (1,4,1,5) to get an array of shape (3,4,7,5).

    from numpy.random import randn
    randn(3,1,7,5) - randn(1,4,1,5)
Inspired by other comments in this thread: The second line can also be run in current versions of Octave/Matlab as-is, BUT they didn't support it before NumPy, and neither did/does FORTRAN[2]. Mathworks resisted introducing this general form of broadcasting into Matlab for a long time. Octave then did it anyway (version 3.6.0 in 2012), and Matlab eventually followed (version R2016b). Before late 2016, Matlab code required `bsxfun(@minus, A, B)` instead of `A-B` to get broadcasting, and before bsxfun (introduced in R2007a) it was even more awkward (code often used `repmat`, or indexing tricks to explicitly expand arrays in memory to match shape).

[1] https://aplwiki.com/wiki/Conformability

[2] suggestion from 2022 to introduce broadcasting to Fortran: https://github.com/j3-fortran/fortran_proposals/issues/252

imurray··on Google is picking ChatGPT responses from Quora as correct answer
Doesn't look like it: https://www.youtube.com/watch?v=imVKSskCB4I#t=3m50s
imurray··on Juggling Lab GIF Server
Thanks so much for the IJDb Colin E!
imurray··on Juggling Lab GIF Server
The JIS database of clubs was already out of date and not responding to update requests ~25 years ago. For an actively-maintained list of events and clubs there is: https://www.jugglingedge.com/ -- which is where most of the rec.juggling and IJDb community went when IJDb shut down. There are probably more active forums now, but it's still the cleanest and most comprehensive list of juggling clubs and events that I know of.
imurray··on Juggling Lab GIF Server
I found an archive of another juggling animator by Paul Klimek, the first known discoverer of siteswap, who called it "quantum juggling": https://web.archive.org/web/20140105002226/https://quantumju...

(Works in Chrome but not firefox for me. I remember it working in firefox, but it's apparently succumbed to bit-rot or the internet archive wrapping.)

imurray··on Juggling Lab GIF Server
Siteswaps can be juggled with any number of hands, such as the 4 hands of two jugglers: https://passing.zone/4-handed-siteswap/

A really nice web-app for searching and animating passing siteswaps is: https://passist.org/

For example: https://passist.org/siteswap/9968926?jugglers=2

In case of confusion: the diagram with arrows is a "causal diagram", showing the throw that is forced by each throw. So linking arrows are not the same ball/club, and the arrows can go backwards in time(!).

passist.org will animate solo siteswaps: https://passist.org/siteswap/915?jugglers=1 -- but 915 is animated really high by default. Click on the animation, then the cog, then increase the juggling speed, rather than the animation speed, to bring the pattern down lower.

There's also an android app that does similar stuff: https://github.com/namlit/siteswap_generator

imurray··on Optimization Without Derivatives: Prima Fortran Version and Inclusion in SciPy
You might be interested in David MacKay's old conjugate gradient implementation in C that also only uses derivatives: https://www.inference.org.uk/mackay/c/macopt.html

(It was written in another time, don't expect it to be nice. My 20-year-old Octave/Matlab wrappers on that page have almost certainly bit-rotted, don't expect them to work.)

imurray··on Google Translate “Get well [Swedish firstname]” translates to “fuck you”
I've also found that DeepL's consistently better or equal when it applies for my personal usage. Some users will care that DeepL doesn't support as many languages, doesn't have TTS, and doesn't offer transliterations (such as pinyin for Chinese).
imurray··on JAX – Augments numpy and Python code with function transformations (2019)
While I didn't like the form of the title, I assumed it was referring to other tutorials with similar titles. Indeed there exists "you don't know JS" and "you don't know bash".
imurray··on Browser extension that lets you follow accounts on foreign Mastodon instances
For those that don't want to install a browser extension (or use an app), I wrote a bookmarklet that lets you toggle between viewing things in your Mastodon instance and a foreign one.

https://homepages.inf.ed.ac.uk/imurray2/code/mastodon.html

← PreviousPage 2 of 13Next →