HNHacker News
TopNewBestAskShowJobs

sigwinch28

867 karma · joined July 20, 2018

https://sigwinch.uk

Don't bother with the broken plaque-ridden keyservers. Use WKD instead:

    gpg --auto-key-locate wkd --locate-keys joe@sigwinch.uk
submissionscomments
sigwinch28··on Bad code is kudzu
I will now refer to senior engineers with a bone to pick with a codebase as “hungry goats”.
sigwinch28··on Base84 deserves a place in file names
I think bit-aligned encodings (i.e. schemes where the output alphabets have a size that is a power of two) are the sweet-spot for encoding schemes due to simplicity of encoding and decoding compared to non-bit-aligned schemes like base84.

Base16 is verbose, but each symbol in the alphabet carries exactly 4 bits (hence why a single byte in hex is two symbols). This is nice for encoding bytes. Base64 requires two non-alphanumeric symbols such as + and / or - and _ because A-Za-z0-9 only provides 62 characters. Each output symbol encodes exactly 6 bits of the original input. But when we're encoding whole bytes we sometimes need to use padding (when the input is not a multiple of 3 bytes and that ambiguity is harmful in the context that decoding occurs in). Then there's Base32, which is what currently seems like a sweet spot to me and is what I'm considering using for identifiers in my own systems. It only requires 32 characters, so can easily take a range of A-Z0-9 or a-z0-9 excluding visually-similar characters. Like Base64 it can also require padding in some settings.

I quite like the general scheme used in the Bech32 and Bech32m framing idea, which allows for a human-readable prefix before a `1` and then the base32 data follows. This can be done because `1` is excluded from the Bech32 alphabet. This prefix can be used to differentiate between kinds of identifier, for example:

https://github.com/bitcoin/bips/blob/master/bip-0350.mediawi...

But to address the article directly, I'm not sure the complexity is worth the gains in the table in https://github.com/jedisct1/zig-base84#encoding, which states that we save ~6.5% of characters, and for 128 input bytes, the resulting Base64 string is 171 characters, while the Base84 string is 161-166 characters.

I also dislike that the chosen alphabet breaks text selection; GitHub very carefully picked their current token formats so that the whole token is selected with a double-click:

> One other neat thing about _ is it will reliably select the whole token when you double click on it. Other characters we considered are sometimes included in application word separators and thus will stop highlighting at that character. Try out double clicking this-random-text versus this_random_text!

https://github.blog/engineering/platform-security/behind-git...

For example, compare double-clicking on the article's Base84 alphabet:

> ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789!#$%&'()+,-;=@[]^_`{}~

with the Bech32 Base32 alphabet:

> qpzry9x8gf2tvdw0s3jn54khce6mua7l

Regardless though we still have the fundamental issue that bytes encoded as Base16, Base32, Base64, or even Base91/Base84 will be _longer_ when encoded. Base32 can encode up to 159 bytes before hitting a filename limit of 255 bytes. I can't remember whether filenames are 16-bit on Windows (UCS-2? UTF-16?), but I speculate on average that maybe you could save a lot of bytes by first ensuring that filenames are UTF-8 before encrypting them, since filenames on a lot of computeres, especially ones running in the west, likely contain a lot of latin characters. You could even switch between encodings to get the best bit-packing in filenames.

At the moment it looks like turbocrypt forbids encrypted filenames that are too long: https://github.com/jedisct1/turbocrypt/blob/4905241d271e84a0...

I think turbocrypt could switch to using a surrogate file when an encrypted filename exceeds 255 bytes to allow for encrypting any filename permitted by the filesystem, while still preserving the property that the same filename in a different directory has the same encrypted filename. If an encrypted filename is too long to fit on the filesystem, maybe its hash could be stored as the filename instead (which is extremely unlikely to collide). Then we could store the full encrypted filename in a `.name` file in the same directory. We could do similar for directory names that are too long. We could then use a less dense but bit-aligned encoding scheme like Base32 with an all-lowercase alphabet without any punctuation to make filename text selection straightfoward. We would avoid all case sensitivity traps that can occur with things like Base64 and Base84. Let's pick two filename prefixes, say "tcf1" for "turbocrypt filename" and "tch1" for "turbocrypt filename hash". Then we could have a directory layout like this:

  tch1q9x8gf2tvdw0 # a file whose encrypted filename exceeds 255 bytes; this filename is a hash of the encrypted filename, e.g. SHA-256.
  tch1q9x8gf2tvdw0.name # a file whose contents are the full encrypted filename of tch1q9x8gf2tvdw0.
  tcf1tvdhc6mua7l9x8s3qprz # a file whose encrypted filename does not exceed 255 bytes.
gocryptfs follows a similar idea for long filenames: https://github.com/rfjakob/gocryptfs-website/blob/master/doc...

Once a system to support long filenames is implemented, the size of the alphabet used for encoding the filenames (like Base84) becomes less important; Base16, Base32, Base64, or another bit-aligned encoding scheme could be used.

As a very small added bonus, you could even implement a bit-aligned codec such as base16 or base64 using SIMD via a bunch of swizzling, shifting, bitmasking, and AND/XORing, making it very fast should the need arise.

sigwinch28··on Appreciating Exif
Conversely, as a hobbyist photographer, I want to do the exact opposite for most photos I take.

I would like my camera info, especially the body, lens, focal length, and settings in the image. I recently discovered that software like Darktable can even take a gpx file and photo timestamps to add coordinates to photos taken on a camera without a GNSS receiver.

sigwinch28··on Spanish track was fractured before high-speed train disaster, report finds
Measurement trains filled with cameras and LIDAR

For example, in the U.K.:

https://en.wikipedia.org/wiki/New_Measurement_Train

sigwinch28··on C++ says “We have try... finally at home”
A writable file closing itself when it goes out of scope is usually not great, since errors can occur when closing the file, especially when using networked file systems.

https://github.com/isocpp/CppCoreGuidelines/issues/2203

sigwinch28··on SQL Anti-Patterns
Or it’s simply an indicator of a schema that has not been excessively normalised (why create an addresses_cities table just to ensure no duplicate cities are ever written to the addresses table?)
sigwinch28··on Belkin shows tech firms getting too comfortable with bricking customers' stuff
Via Wikipedia:

> The developer offered full refunds to the game for macOS and Linux owners regardless of how long they had the game.

https://en.wikipedia.org/wiki/Rocket_League#Free-to-play_tra...

https://www.rockpapershotgun.com/rocket-league-ending-mac-an...

sigwinch28··on U.K. orders Apple to let it spy on users’ encrypted accounts
I’m from the U.K. and I consider the government’s actions around digital privacy to be somewhere between incompetent and malicious.
sigwinch28··on Jaywalking legalized in New York City
>waiting for 2 different lights just to get to the opposite corner.

A solution sometimes seen in London is a “Pedestrian Scramble”, where pedestrians are explicitly given full (and even diagonal) access to a junction with all other traffic stopped.

https://en.wikipedia.org/wiki/Pedestrian_scramble

sigwinch28··on Ask HN: How to store and share passwords in a company?
With SSO, the party running the SSO decides what the authentication policy is.

For example, where the authentication request is coming from (on-site, managed device), what methods are being used (hardware second factor, Authenticator app).

These are all things that the SSO can check at time of authentication, before a token or session key gets issued to the user. Also, all of these things can be checked again when doing any auth flows for the various linked services.

So with stolen SSO credentials, they might be worth diddly squat to you if you didn’t think to also be on-site or on a managed company device (physically or virtually).

sigwinch28··on Airlines are running out of 4-digit flight numbers
Reuse of identifiers seems to be a theme in aviation https://news.ycombinator.com/item?id=37401864
sigwinch28··on A skeptic's first contact with Kubernetes
If we replace “YAML” with “JSON” and then talk about naive text-based templating, it seems wild.

That’s because it is. Then we go back to YAML and add whitespace sensitivity and suddenly it’s the state-of-the-art for declaring infrastructure.

sigwinch28··on Elsevier embeds a hash in the PDF metadata that is unique for each download (2022)
A lot government funding stipulates open access publication of some form.

https://en.wikipedia.org/wiki/Open-access_mandate

sigwinch28··on Google releases smart watch for kids
A smartwatch for kids could be so good if it was designed in a way to be educational, but most importantly, which respects a child’s privacy utmost, even from their parents in terms of tracking.

For example, a maps app, to always get the kid home if they’re lost. Medication reminders. Fitness tracking. Emergency SOS. A calendar to remind them about family birthdays and upcoming holidays. School timetables. Medical ID. Payment cards or passes for travel (in Western Europe a lot of schoolchildren commute by themselves, especially on public transport) and spending their allowance. Let the kid choose to notify their family of their location as and when they want to. Empower them to use tech to their advantage but put their privacy first.

Children are going to end up as adults in this world regardless of whether we teach them, so we should be teaching them the benefits and warning them of the many bad actors. We should be teaching our children the skills they need to navigate the modern world. This includes technology and abusive/controlling relationships.

I believe a good responsible smartwatch for kids can exist. Alas, this is Google and helicopter parents exist, so this product is not it.

sigwinch28··on Ticketmaster breach affects more than half a billion users
At 2.6 megabytes per dollar, it is at least cheaper than the price of a (very legal) kdb license, which can hover around 3 bytes per dollar.

Comparing apples and oranges here but I like thinking about the monetary value assigned to a byte.

sigwinch28··on Withdraw most of my ownership in favor of Mark
This seems to be about GitHub’s annoying CODEOWNERS feature where every matching user in the CODEOWNERS file is forcefully added as a reviewer of opened PRs.

To my knowledge this “feature” can’t be toggled independently and in my experience often drastically reduces the signal-to-noise ratio of GitHub notifications for people in a CODEOWNERS file.

I wish GitHub allowed this to be configured. You either get this functionality and enforced code owner approval, or neither.

sigwinch28··on The OpenAI board was right
There are clauses in many Hollywood contracts which prevent the use of an actor’s likeness.

See what happened in Back to the Future Part 2: https://en.m.wikipedia.org/wiki/Back_to_the_Future_Part_II#R...

sigwinch28··on Meteor seen in Portugal
https://youtu.be/MgNItWdfEIU?si=QGZbjRIh50InkiHV

Context: https://en.wikipedia.org/wiki/Your_Name

sigwinch28··on Raspberry Pi Ltd is considering an IPO
> What's the risk here?

A monopolistic stock exchange could, off the top of my head: increase fees, engage in rentseeking behaviour, impose unfair rules, discriminate (against companies and traders) or reduce the quality of service.

Listing on different (or multiple) exchanges ensures that they engage in proper competition.

> compared to the actual extractions you are subjected to in the UK

An apt example of why it's healthy to have competition. A working professional or business could relocate to a country where less of their wealth is taken from them.

sigwinch28··on Apple and Google deliver support for unwanted tracking alerts in iOS and Android
At least for AirTags before today, this functionality allegedly only triggers when the device is not near one of the owner’s iPhones/iPad/Macbooks.
sigwinch28··on Mount Everest: Climbers will need to bring poo back to base camp
They leave oxygen tanks littered up there.

They leave human bodies up there, too, but those are usually moved out of view.

sigwinch28··on Should error messages apologize? (2013)
I try to do this when writing error messages I expect to be seen by other engineers. I also try to state why this is an error condition.

“ERROR: Database query returned 0 rows”

versus

“ERROR: Database query returned 0 rows but need 2 or more rows for this operation. Ensure $other_etl has successfully ingested the data.”

sigwinch28··on Alaska Airlines flight 1282 NTSB preliminary report [pdf]
It is the rule in Europe, which is mentioned in (II)(C) in the link. I failed to link it properly.
sigwinch28··on Alaska Airlines flight 1282 NTSB preliminary report [pdf]
More altitude means more time to work the problem.
sigwinch28··on Alaska Airlines flight 1282 NTSB preliminary report [pdf]
The rule (edit: in Europe) is now 25 hours for aircraft over a certain weight, though it is not (currently) retroactively applied to existing equipment.

https://www.federalregister.gov/documents/2023/12/04/2023-26...

sigwinch28··on Amazon's Ring to stop letting police request doorbell video from users
Storing CCTV footage on-prem is at odds with preserving said footage. It’s usually easier for an intruder to destroy or steal the physical storage device on site compared to gaining access to an off-site storage device in an unknown or practically inaccessible location. This is especially true for residential properties.

Do you hide it in a wall or in an attic somewhere?

sigwinch28··on Why are we templating YAML? (2019)
> your config becoming more and more complex until it inevitably needs its own config, etc. You wind up with a sprawling, Byzantine mess.

We're already there with Helm.

People write YAML because it's "just data". Then they want to package it up so they put it in a helm chart. Then they add variable substitution so that the name of resources can be configured by the chart user. Then they want to do some control flow or repetitiveness, so they use ifs and loops in templates. Then it needs configuring, so they add a values.yaml configuration file to configure the YAML templating engine's behaviour. Then it gets complicated so they define helper functions in the templating language, which are saved in another template file.

So we have a YAML program being configured by a YAML configuration file, with functions written in a limited templating language.

But that's sometimes not enough, so sometimes variables are also defined in the values.yaml and referenced elsewhere in the values.yaml with templating. This then gets passed to the templating system, which then evaluates that template-within-a-template, to produce YAML.

sigwinch28··on Why are we templating YAML? (2019)
I would recommend implementing a similar API to Grafana Tanka: https://tanka.dev

When you "synthesise", the returned value should be an array or an object.

1. If it's an object, check if it has an `apiVersion` and `kind` key. If it does, yield that as a kubernetes object and do not recurse. 2. If it's an array or any other object, repeat this algorithm for all array elements and object values.

This gives a lot of flexibility to users and other engineers because they can use any data structures they want inside their own libraries. TypeScript's type system improves the ergonomics, too.

sigwinch28··on Why are we templating YAML? (2019)
I don't see any practical difference w.r.t. cybersecurity between "I blindly applied this pile of YAML to my production kubernetes clusters without looking at it" and "I blindly downloaded and ran this computer program on my CI runner without looking at it".

A supply chain attack on the former means that your environment is compromised. So does the latter.

sigwinch28··on Why are we templating YAML? (2019)
Relevant: https://noyaml.com/

YAML and its ecosystem is full of footguns and ergonomics problems, especially when the length of the document extends beyond the height of a user's editor or viewport. Loss of context with indentation, non-compliant or unsafe parsers, and strange boolean handling to name a few.

It becomes even worse when people decide that static YAML data files should have variable substitution or control flow via templating. "Stringly-typed programming" if you will. If we all started writing JSON text templates I think a lot of people would rightly argue we should write small stdlib-only programs in Python, Typescript, or Ruby to emit this JSON instead of using templated text files. Then it becomes apparent that the YAML template isn't a static data file at all, but part of a program which emits YAML as output. We're already exposing people to basic programming if we're using YAML templates. People brew a special kind of YAML-templated devops hell using tools like Kustomize and Helm, each of which are "just YAML" but are full of idiosyncracies and tool-specific behaviour which make the use of YAML almost coincidental rather than a necessity.

Yes, sometimes people would prefer to look at YAML instead of JSON, in which case I suggest you use a YAML serialization library, or pipe output into a tool like `yq` so you can view the pretty output. In a pinch you could even output JSON and then feed it through a YAML formatter.

The Kubernetes community seems to have this penetrating "oh, it's just YAML" philosophy which means we get mediocre DSLs in "just YAML" which actually encode a lot of nuanced and unintuitive behaviour which varies from tool to tool.

Look at kyverno, for example: it uses _parentheses_ in YAML key names to change the semantics of security policies! https://kyverno.io/docs/writing-policies/validate/ . This is different to (what I think are the much better ideas of) something like kubewarden, gatekeeper, or jspolicy, which allow engineers to write their policies in anything that compiles to WASM, OPA, and Typescript/Javascript respectively.

We engineers, as a discipline, have decades of know-how building and using general purpose programming languages with type checkers, linters, packaging systems, and other tools, but we throw them all away as soon as YAML comes along. It's time to put the stringified YAML templates away and engage in the ecosystem of mature tools we already know to perform one simple task they are already good at: dumping JSON on stdout.

Let's move the control flow back into the tool and out of the YAML.

Page 1 of 8Next →