'changeme' is valid base64
blog.3fx.ch
blog.3fx.ch
$ printf AAAA | base64 --decode | od -tx1
0000000 00 00 00
0000003
$ printf AAAAAAAA | base64 --decode | od -tx1
0000000 00 00 00 00 00 00
0000006
$ printf AQEB | base64 --decode | od -tx1
0000000 01 01 01
0000003
$ printf AQID | base64 --decode | od -tx1
0000000 01 02 03
0000003
$ printf main | base64 --decode | od -tx1
0000000 99 a8 a7
0000003
$ printf scrabble | base64 --decode | od -tx1
0000000 b1 ca da 6d b9 5e
0000006
$ printf 12345678 | base64 --decode | od -tx1
0000000 d7 6d f8 e7 ae fc
0000006
Since '+' and '/' are also used as symbols in Base64 encoding (for binary 111110 and 111111, respectively), we also have: $ printf 1+2+3+4+5/11 | base64 --decode | od -tx1
0000000 d7 ed be df ee 3e e7 fd 75
0000011
$ printf "\xd7\xed\xbe\xdf\xee\x3e\xe7\xfd\x75" | base64
1+2+3+4+5/11 $ echo "changeme" | base64 --decode | hexdump -C
00000000 72 16 a7 81 e9 9e |r.....|
00000006
$ echo "*lock*" | base64 --decode | hexdump -C
base64: invalid input $ echo "llockk" | base64 --decode | hexdump -C
00000000 96 5a 1c |.Z.|
00000003https://en.m.wikipedia.org/wiki/Base64
(I got burned by this a few months ago where Chrome can produce HAR files that have base64 encoded bodies without padding characters, and .NETs built in decoder expects them)
> the probability that changeme is actually valid base64 encoding must be very low
This makes no sense. The characters are part of the base64 alphabet and you have an acceptable number of them.
You wouldn't claim that the probability that deadbeef is a valid base 16 number must be very low. Probability doesn't come into it, the characters are part of the alphabet, that's all there is to it.
Thanks!
How did you work that out?
I'm a bit confused, I thought any string with only lowercase letters was "valid base64" (more precisely, I thought "valid base64" is equivalent to "string consists only of the 64 special characters we're using to represent digits 0-63").
It gets more complicated if you're using base64 to represent a number of bytes that isn't a multiple of 3, those would be unlikely to happen randomly. Those will usually end in a number of '=' signs to pad their length to a multiple of 4 and indicate how many bytes are missing. Although apparently there are also versions o base64 that don't include the padding.
Try running `base64 -d <<< a` and it will fail, but `base64 -d <<< aaaa` works (4*6 == 24 bits, which is divisible by 8 and gives 3 bytes of output)
As an example, Elixir's stdlib includes a base64 decoder; passing "padding: false" will give the "decode anything" behavior you're describing.
The author is surprised that the dummy string is valid base64. Other comments on this thread discuss whether all lowercase strings are valid and whether implementations differ.
But ... Isn't the point of using base64 (rather than hex or something) to make a large share of easily typeable strings valid (which implies they can be shorter/faster/cheaper)? I.e. someone upstream of you _wanted_ "changeme" to be valid, and the misunderstanding is less about the technical details of base64, and more about a hidden conflict in desires between people making design choices separated by time, space, and perhaps organization.
I use a VS Code extension for TODOs (Todo Tree) , which highlights TODO comments and aggregates them in a centralized list, which is useful. I think it can configured to include custom words, might be useful for your personal "changeme"s
I wonder what the TK equivalent in German publishing is, here I'd expect a TK bigram to happen too often.