HNHacker News
TopNewBestAskShowJobs

acatton

1,998 karma · joined February 14, 2016

Just some random engineer. https://antoine.catton.fr

Currently employed by Google, Inc. as a Site Reliability Engineer.

All my opinions are my own and NOT the views of my current or past employers.

submissionscomments
acatton··on Building static binaries with Go on Linux
The secret with sqlite is to use "-tags sqlite_omit_load_extension", if you don't use any extension. (which is 99% of the users)

This is explained in https://www.arp242.net/static-go.html

acatton··on RegreSSHion: RCE in OpenSSH's server, on glibc-based Linux systems
Force of habit. No particular reason, "4kiB feels like a nice number", cargo culting. Choose one :) .

It doesn't matter if you have more than 256 bits, as your key file gets hashed with SHA256 at the end[1]. It could be 5GiB it would be the same. So yes, you're right to mention that more bits don't add more security.

[1] https://github.com/Tarsnap/spiped/blob/2194b2c64de65eed119ab...

acatton··on RegreSSHion: RCE in OpenSSH's server, on glibc-based Linux systems
* The tool is not obscure, it's packaged in most distributions.[1][2][3] It was written and maintained by Colin Percival, aka "the tarnsnap guy" or "the guy who invented scrypt". He is the security officer for FreeBSD.

* spiped can be used transparently by just putting a "ProxyCommand" in your ssh_config. This means you can connect to a server just by using "ssh", normally. (as opposed to wireguard where you need to always be on your VPN, otherwise connnect to your VPN manually before running ssh)

* As opposed to wireguard which runs in the kernel, spiped can easily be set-up to run as a user, and be fully hardened by using the correct systemd .service configuration [4]

* The protocol is much more lightweight than TLS (used by stunnel), it's just AES, padded to 1024 bytes with a 32 bit checksum. [5]

* The private key is much easier to set up than stunnel's TLS certificate, "dd if=/dev/urandom count=4 bs=1k of=key" and you're good to go.

[1] https://packages.debian.org/bookworm/spiped

[2] https://www.freshports.org/sysutils/spiped/

[3] https://archlinux.org/packages/extra/x86_64/spiped/

[4] https://ruderich.org/simon/notes/systemd-service-hardening

[5] https://github.com/Tarsnap/spiped/blob/master/DESIGN.md

acatton··on RegreSSHion: RCE in OpenSSH's server, on glibc-based Linux systems
Yearly reminder to run your ssh server behind spiped.[1] [2] [3]

[1] https://www.tarsnap.com/spiped.html

[2] https://news.ycombinator.com/item?id=29483092

[3] https://news.ycombinator.com/item?id=28538750

acatton··on CentOS Linux 7 will reach EOL on Sunday
Funnily enough, years ago, I migrated from Debian as my daily driver to (at the time) "Fedora Core" on my desktop.

My first question was "what's the replacement for aptitude", and people pointed me to "yum shell". It was not as good, but I got used to it, and went with it.

If you run "aptitude" on debian, without any argument, you end up in a TUI, you can use it to install or remove packages from your system, and then see the "preview" of the change, and apply/cancel the change. The same way people use "yum shell".

I'm used to new "dnf shell", so I don't miss aptitude anymore, but I think aptitude is what you're looking for.

acatton··on How I Made Google's "Web" View My Default Search
Most search engine popping up are just a rebranding of bing and yandex. There is a good article about it: https://seirdy.one/posts/2021/03/10/search-engines-with-own-...

Regarding your question of cost, if you buy the cheapest hardware possible, and manage it yourself, because you don't want to pay AWS' premium, for storage (again, assuming to discard 90% of the pages, and only do english) you'll be at least at €80k (12 × €3000 disk servers with 300 × €150 HDDs) of upfront cost, then, if you colocate at Hetzner, just for storage, you'll bet at €500/month + ~€3k of electricity that you need to pay yourself.

My gut feeling is that this would cost (only 10% of the English-speaking web) ~€150k of upfront cost and €5k/month to run. Assuming you buy the cheapest everything, and do your own admin sys. And this is not forecasting growth, serving ads, etc...

acatton··on How I Made Google's "Web" View My Default Search
I don't know how Kagi manages it. I suspect like duckduckgo, they index a little bit on their own, and use a GBY [1] as a back up. According to Seirdy[1], they use Brave in the background. Brave just burn their cryptomoney to build a search engine. Don't get me wrong, I like what their doing, but it was easy for them to start since they bootstraped from Cliqz' index[2]

And for the second part, you do need to store the relationship between keywords and pages, that's what I was talking about. You cannot store a relationship between "types of water" and "reddit.com" you need to store it between "type of water" and "reddit.com/r/hydrohomies/..."

[1] https://seirdy.one/posts/2021/03/10/search-engines-with-own-...

[2] https://brave.com/blog/brave-search/

acatton··on How I Made Google's "Web" View My Default Search
I wish you luck and I hope you succeed, but you make it sound much much easier than what it would be.

First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "sushi" entry will be ~19GiB large, assuming you use UUIDs. If you think 90% of this is spam, and you only index the non spam (detecting spam/seo is far from trivial, but let's say you figure it out), you still need ~2GiB just for mapping "sushi". With 755,865 words in the English dictionary, according to wikipedia[2], you'll need ~1.5 PiB (yes, pebi/peta, 1,536 TiB) just to store relationships for English pages. This is assuming you don't support other languages, you discard 90% of pages, and you don't cache the content of pages for re-indexing.

In addition to this, you also need to store the meta-data for each pages (vote counts from your voting system, whether it's serving different content, etc...). The order of magnitude has to be in the O(100TiB) from my conservative gut feeling. (still assuming you discard 90% of the web, and I'll assume you aggregate the metadata on the domain, not on the individual pages)

The second challenge is your ranking. Now that you've become the dominant search engine with your awesome ranking system, you will become the main target for swaths of motivated click-farms which are exploiting workers from low income countries. They will be trying to register accounts, vote and game your ranking. You can most likely detect this behaviour, but their behaviour will be very similar to a significant portion of your real users. So you'll be fishing in a pond with a rocket launcher, and some of your legitimate users will be collateral victims. Otherwise, you'll spend most of your time playing a cat-and-mouse game with the SEO spammers instead of improving your search engine and fixing bugs.

I'm also falling in the trap "i could rewrite that in a weekend" sometimes, but for a search engine, I would love to see decent competition, but it's near impossible.

[1] https://news.ycombinator.com/item?id=30925402

[2] https://en.wikipedia.org/w/index.php?title=List_of_dictionar...

acatton··on Design docs at Google (2020)
There is a review process, but "yes, your comments are relevant, but they're just nit picking, can you just approve it so that we can start working on this project?"
acatton··on Design docs at Google (2020)
I do not share the same experience as the author, as someone working for the mentioned company.

There is a wide range of types of design docs, and none of them are useful. I've rarely seen any useful design doc at Google. I feel that design docs are for engineers who are too much process-oriented.

Here are the types of design docs I've encountered in the wild:

* The promo design doc: it's not really explaining what this is trying to solve, it's more stating that this project is awesome, and makes the company better. The logical conclusion is that the author of this doc should be promoted.

* The turbo-encabulator[1] design doc: this is a technobabble design doc which is full of terms never encountered before, and which is not understandable unless you're a senior member of the team. I'm sometimes not even sure the senior members of the team understand it...

* The new-grad design doc: this a design doc with no substance, but as the person just graduated from university, they felt compelled to make is as long as possible, to prove... I don't know what... It is not conveying any information. They most likely copied/pasted huge chunks of the code they've already written, to fill most of the ~70 pages of the doc.

* The made-up-facts design doc: this a design doc full of "everybody knows that", "they all say". Of course, it's not as obviously done as some politicians do it. But the design doc will push their design with "this follows good practices", "this software is slow, therefore..." who defined the good practice? why is it a good practice? what is slow? was it measured? is it an end user feeling?

This is 99% of the design docs I've seen out there. Of course, exceptions exists, but they're very rare, in my experience. I'm shocked that the author pushes for this practice... But again, they were no engineer, they were a director, I guess design docs make sense for their position, for which I'm still trying to figure out the value these folks bring.

[1] https://en.wikipedia.org/wiki/Turbo_encabulator

acatton··on The London office where swinging pendulums keep cyber threats at bay
Random numbers are used to generate your TLS certificate key or your SSH key. If one can predict how a TLS or SSH key was generated they could impersonate the key holder.

Here is an explanation of what was possible when a Debian packager mistakenly introduced a patch which reduced the SSL certificate keyspace in 2008: https://jblevins.org/log/ssh-vulnkey

The possibilities of keys which were generated by this random number generator was so small, that a brute-force attack on keys was feasible.

That being said, for years, random number generators have been using random signals coming to your computer (key strokes, network packets, ...) and feeding them into a sponge function. You don't need lava lamps or pendulums to generate random numbers, it's just for the press.

https://www.2uo.de/myths-about-urandom/

acatton··on Tell HN: "Default" FileZilla download bundled with adware
https://www.deceptive.design/types/preselection
acatton··on Indian government moves to ban ProtonMail after bomb threat
> ProtonMail is the best choice if you want an end-to-end encrypted email platform

Is this native advertisement? There are many other choices, some of them being as good if not better. A few of them are small mom-and-pop businesses which didn't take on investment, and won't enshittify ( https://en.wiktionary.org/wiki/enshittification )

* https://www.privacytools.io/privacy-email * https://github.com/pluja/awesome-privacy?tab=readme-ov-file#...

acatton··on Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced
https://googlesystem.blogspot.com/2013/05/error-404-not-foun...
acatton··on Quickemu: Quickly run optimised Windows, macOS and Linux virtual machines
TLS prevents a different kind of attack, the MitM one which you describe.

GPG signing covers this threat model but much more, the threats include:

* The server runs vulnerable software and is compromised by script-kiddies. They, then, upload arbitrary packages on the server

* The cloud provider is compromised and attackers take over the server from the admin cloud provider account.

* Attacker use a vulnerability (from SSH, HTTPd, ...) to upload arbitrary software packages to the server

GPG doesn't protect against the developer machine getting compromised, but it guarantees that what you're downloading has been issued from the developer's machine.

acatton··on Quickemu: Quickly run optimised Windows, macOS and Linux virtual machines
Just a security reminder from the last time this got posted[1]

This tool downloads random files from the internet, and check their checksum against other random files from the internet. [2]

This is not the best security practice. (The right security practice would be to have the gpg keys of the distro developers committed in the repository, and checking all files against these keys)

This is not downplaying the effort which was put in this project to find the correct flags to pass to QEMU to boot all of these.

[1] https://news.ycombinator.com/item?id=28797129

[2] https://github.com/quickemu-project/quickemu/blob/0c8e1a5205...

acatton··on Tesla blamed drivers for failures of parts it long knew were defective
VHS was a superior format to Betamax. I mean... their sales numbers proved it, right?
acatton··on Tesla blamed drivers for failures of parts it long knew were defective
Yes, but they are doing another layer of quality control on whatever comes out of Bosch. Because at the end of the day, they can't blame Bosch for the bad quality of cars sold under their brand.
acatton··on Tesla blamed drivers for failures of parts it long knew were defective
Exactly, it worked for the Boeing 737 Max... for a while...
acatton··on Show HN: My Go SQLite driver did poorly on a benchmark, so I fixed it
There are more "Alert blocks" available https://docs.github.com/en/get-started/writing-on-github/get...

This is github specific, not part of commonmark AFAIK

acatton··on I benchmarked six Go SQLite drivers
The modernc one is the one I was refering to.

https://gitlab.com/cznic/sqlite

acatton··on Bash one-liners for LLMs
I get excited when hackers like Justine (in the most positive sense of the word) start working with LLMs.

But every time, I am let down. I still dream of some hacker making LLMs run on low-end computers like a 4GB rasbperry pi. My main issues with LLMs is that you almost need a PS5 to run the them.

acatton··on I benchmarked six Go SQLite drivers
In my personal opinion and usage, the performance doesn't matter. Only one driver is written in pure go, and can be easily statically compiled and/or cross-compiled.

> modernc, modernc.org/sqlite, a pure Go solution. This is a newer library, based on the SQLite C code re-written in Go.

Unless I'm mistaken, this is not a re-write in Go. This is a transpilation of the the SQLite C library into go, using https://gitlab.com/cznic/ccgo

acatton··on Tesla FSD Timeline
You're lucky that it's "FSD" which has like 4 possible meaning.

Sometimes it's acronyms with 12 different possible meaning, where 4 of them could apply to the given context...

acatton··on 23andMe updates their TOS to force binding arbitration
Not only if you're a consumer. There are multiple cases in Germany of Oberlandesgerichten (~= "Circuit courts") voiding arbitration clauses in B2B contracts as well.

Subway (the sandwich chain) is a good example of that. They were kinda screwing their franchisees and were forcing them to do arbitration in NYC, even for German franchisees. This was voided by the northern German "circuit court"[1]

[1] https://www.omsels.info/wp-content/uploads/OLG-Schleswig-Urt...

acatton··on Stripe live dashboard
Firefox on iPhone uses Webkit. It's basically a Safari skin.

https://en.wikipedia.org/w/index.php?title=Firefox&oldid=118...

> It is the first Firefox-branded browser not to use the Gecko layout engine as is used in Firefox for desktop and mobile. Apple's policies require all iOS apps that browse the web to use the built-in WebKit rendering framework and WebKit JavaScript, so using Gecko is not possible.

acatton··on Google's plan to stop Apple from getting serious about search
https://en.wikipedia.org/w/index.php?title=Archive.today&old...
acatton··on Google’s dominance under siege: Antitrust trial threatens sweeping changes
Because Google's index is much larger.

According to estimates[1], Google indexes ~60 billions pages and Bing is only ~4 billions. If these numbers are close to reality, Bing index is 7% of Google's.

IMHO, this leads to all the problems people complain about Bing. The algorithm is shit because they have less data to fine tune it, people use it less because the algorithm is shit or the results don't show up, so they cannot collect data to fine tune the algorithm more, ...

[1] https://www.worldwidewebsize.com/

acatton··on Google’s dominance under siege: Antitrust trial threatens sweeping changes
Its search index, in my personal opinion.

A quarter of the websites on the internet are behind Cloudlfare/Akamai with anti-scraping enabled. Of course they want to end up on Google, so Google/Bing get a green light from Cloudflare.

IMHO, that basically makes it impossible for anybody else to spin up an index and compete with Google Search. They cannot just scrape, they also have to work around anti-scraping measure by these CDNs. They have to maintain these workarounds, while trying to scrape as much as google has scraped so far, which is already a feat in itself.

Once you have this, you have everything: a critical mass of users, ads, data set for AI, ...

I dare anybody to scrape Reddit or Stackoverflow from scratch today.

acatton··on Bookworm – the new version of Raspberry Pi OS
raspi-firmware is semi-officially maintained. But Raspberry Pi OS also has scripts to set flags on the eeprom like "vcgencmd" and "rpi-eeprom-update"

What I was saying these could be maintained in Debian.

← PreviousPage 2 of 9Next →