This is explained in https://www.arp242.net/static-go.html
1,998 karma · joined February 14, 2016
Currently employed by Google, Inc. as a Site Reliability Engineer.
All my opinions are my own and NOT the views of my current or past employers.
This is explained in https://www.arp242.net/static-go.html
It doesn't matter if you have more than 256 bits, as your key file gets hashed with SHA256 at the end[1]. It could be 5GiB it would be the same. So yes, you're right to mention that more bits don't add more security.
[1] https://github.com/Tarsnap/spiped/blob/2194b2c64de65eed119ab...
* spiped can be used transparently by just putting a "ProxyCommand" in your ssh_config. This means you can connect to a server just by using "ssh", normally. (as opposed to wireguard where you need to always be on your VPN, otherwise connnect to your VPN manually before running ssh)
* As opposed to wireguard which runs in the kernel, spiped can easily be set-up to run as a user, and be fully hardened by using the correct systemd .service configuration [4]
* The protocol is much more lightweight than TLS (used by stunnel), it's just AES, padded to 1024 bytes with a 32 bit checksum. [5]
* The private key is much easier to set up than stunnel's TLS certificate, "dd if=/dev/urandom count=4 bs=1k of=key" and you're good to go.
[1] https://packages.debian.org/bookworm/spiped
[2] https://www.freshports.org/sysutils/spiped/
[3] https://archlinux.org/packages/extra/x86_64/spiped/
[4] https://ruderich.org/simon/notes/systemd-service-hardening
[1] https://www.tarsnap.com/spiped.html
My first question was "what's the replacement for aptitude", and people pointed me to "yum shell". It was not as good, but I got used to it, and went with it.
If you run "aptitude" on debian, without any argument, you end up in a TUI, you can use it to install or remove packages from your system, and then see the "preview" of the change, and apply/cancel the change. The same way people use "yum shell".
I'm used to new "dnf shell", so I don't miss aptitude anymore, but I think aptitude is what you're looking for.
Regarding your question of cost, if you buy the cheapest hardware possible, and manage it yourself, because you don't want to pay AWS' premium, for storage (again, assuming to discard 90% of the pages, and only do english) you'll be at least at €80k (12 × €3000 disk servers with 300 × €150 HDDs) of upfront cost, then, if you colocate at Hetzner, just for storage, you'll bet at €500/month + ~€3k of electricity that you need to pay yourself.
My gut feeling is that this would cost (only 10% of the English-speaking web) ~€150k of upfront cost and €5k/month to run. Assuming you buy the cheapest everything, and do your own admin sys. And this is not forecasting growth, serving ads, etc...
And for the second part, you do need to store the relationship between keywords and pages, that's what I was talking about. You cannot store a relationship between "types of water" and "reddit.com" you need to store it between "type of water" and "reddit.com/r/hydrohomies/..."
[1] https://seirdy.one/posts/2021/03/10/search-engines-with-own-...
First of all, you're going to drown in hardware costs, if you run your own hardware. If you run on AWS, you will be the largest AWS customer. 2 years ago, when Google was still displaying result counts, I got 1.3 billions results for "sushi"[1]. This means that if you use a reverse index to lookup your results, the "sushi" entry will be ~19GiB large, assuming you use UUIDs. If you think 90% of this is spam, and you only index the non spam (detecting spam/seo is far from trivial, but let's say you figure it out), you still need ~2GiB just for mapping "sushi". With 755,865 words in the English dictionary, according to wikipedia[2], you'll need ~1.5 PiB (yes, pebi/peta, 1,536 TiB) just to store relationships for English pages. This is assuming you don't support other languages, you discard 90% of pages, and you don't cache the content of pages for re-indexing.
In addition to this, you also need to store the meta-data for each pages (vote counts from your voting system, whether it's serving different content, etc...). The order of magnitude has to be in the O(100TiB) from my conservative gut feeling. (still assuming you discard 90% of the web, and I'll assume you aggregate the metadata on the domain, not on the individual pages)
The second challenge is your ranking. Now that you've become the dominant search engine with your awesome ranking system, you will become the main target for swaths of motivated click-farms which are exploiting workers from low income countries. They will be trying to register accounts, vote and game your ranking. You can most likely detect this behaviour, but their behaviour will be very similar to a significant portion of your real users. So you'll be fishing in a pond with a rocket launcher, and some of your legitimate users will be collateral victims. Otherwise, you'll spend most of your time playing a cat-and-mouse game with the SEO spammers instead of improving your search engine and fixing bugs.
I'm also falling in the trap "i could rewrite that in a weekend" sometimes, but for a search engine, I would love to see decent competition, but it's near impossible.
[1] https://news.ycombinator.com/item?id=30925402
[2] https://en.wikipedia.org/w/index.php?title=List_of_dictionar...
There is a wide range of types of design docs, and none of them are useful. I've rarely seen any useful design doc at Google. I feel that design docs are for engineers who are too much process-oriented.
Here are the types of design docs I've encountered in the wild:
* The promo design doc: it's not really explaining what this is trying to solve, it's more stating that this project is awesome, and makes the company better. The logical conclusion is that the author of this doc should be promoted.
* The turbo-encabulator[1] design doc: this is a technobabble design doc which is full of terms never encountered before, and which is not understandable unless you're a senior member of the team. I'm sometimes not even sure the senior members of the team understand it...
* The new-grad design doc: this a design doc with no substance, but as the person just graduated from university, they felt compelled to make is as long as possible, to prove... I don't know what... It is not conveying any information. They most likely copied/pasted huge chunks of the code they've already written, to fill most of the ~70 pages of the doc.
* The made-up-facts design doc: this a design doc full of "everybody knows that", "they all say". Of course, it's not as obviously done as some politicians do it. But the design doc will push their design with "this follows good practices", "this software is slow, therefore..." who defined the good practice? why is it a good practice? what is slow? was it measured? is it an end user feeling?
This is 99% of the design docs I've seen out there. Of course, exceptions exists, but they're very rare, in my experience. I'm shocked that the author pushes for this practice... But again, they were no engineer, they were a director, I guess design docs make sense for their position, for which I'm still trying to figure out the value these folks bring.
Here is an explanation of what was possible when a Debian packager mistakenly introduced a patch which reduced the SSL certificate keyspace in 2008: https://jblevins.org/log/ssh-vulnkey
The possibilities of keys which were generated by this random number generator was so small, that a brute-force attack on keys was feasible.
That being said, for years, random number generators have been using random signals coming to your computer (key strokes, network packets, ...) and feeding them into a sponge function. You don't need lava lamps or pendulums to generate random numbers, it's just for the press.
Is this native advertisement? There are many other choices, some of them being as good if not better. A few of them are small mom-and-pop businesses which didn't take on investment, and won't enshittify ( https://en.wiktionary.org/wiki/enshittification )
* https://www.privacytools.io/privacy-email * https://github.com/pluja/awesome-privacy?tab=readme-ov-file#...
GPG signing covers this threat model but much more, the threats include:
* The server runs vulnerable software and is compromised by script-kiddies. They, then, upload arbitrary packages on the server
* The cloud provider is compromised and attackers take over the server from the admin cloud provider account.
* Attacker use a vulnerability (from SSH, HTTPd, ...) to upload arbitrary software packages to the server
GPG doesn't protect against the developer machine getting compromised, but it guarantees that what you're downloading has been issued from the developer's machine.
This tool downloads random files from the internet, and check their checksum against other random files from the internet. [2]
This is not the best security practice. (The right security practice would be to have the gpg keys of the distro developers committed in the repository, and checking all files against these keys)
This is not downplaying the effort which was put in this project to find the correct flags to pass to QEMU to boot all of these.
[1] https://news.ycombinator.com/item?id=28797129
[2] https://github.com/quickemu-project/quickemu/blob/0c8e1a5205...
This is github specific, not part of commonmark AFAIK
But every time, I am let down. I still dream of some hacker making LLMs run on low-end computers like a 4GB rasbperry pi. My main issues with LLMs is that you almost need a PS5 to run the them.
> modernc, modernc.org/sqlite, a pure Go solution. This is a newer library, based on the SQLite C code re-written in Go.
Unless I'm mistaken, this is not a re-write in Go. This is a transpilation of the the SQLite C library into go, using https://gitlab.com/cznic/ccgo
Sometimes it's acronyms with 12 different possible meaning, where 4 of them could apply to the given context...
Subway (the sandwich chain) is a good example of that. They were kinda screwing their franchisees and were forcing them to do arbitration in NYC, even for German franchisees. This was voided by the northern German "circuit court"[1]
[1] https://www.omsels.info/wp-content/uploads/OLG-Schleswig-Urt...
https://en.wikipedia.org/w/index.php?title=Firefox&oldid=118...
> It is the first Firefox-branded browser not to use the Gecko layout engine as is used in Firefox for desktop and mobile. Apple's policies require all iOS apps that browse the web to use the built-in WebKit rendering framework and WebKit JavaScript, so using Gecko is not possible.
According to estimates[1], Google indexes ~60 billions pages and Bing is only ~4 billions. If these numbers are close to reality, Bing index is 7% of Google's.
IMHO, this leads to all the problems people complain about Bing. The algorithm is shit because they have less data to fine tune it, people use it less because the algorithm is shit or the results don't show up, so they cannot collect data to fine tune the algorithm more, ...
A quarter of the websites on the internet are behind Cloudlfare/Akamai with anti-scraping enabled. Of course they want to end up on Google, so Google/Bing get a green light from Cloudflare.
IMHO, that basically makes it impossible for anybody else to spin up an index and compete with Google Search. They cannot just scrape, they also have to work around anti-scraping measure by these CDNs. They have to maintain these workarounds, while trying to scrape as much as google has scraped so far, which is already a feat in itself.
Once you have this, you have everything: a critical mass of users, ads, data set for AI, ...
I dare anybody to scrape Reddit or Stackoverflow from scratch today.
What I was saying these could be maintained in Debian.