8 karma · joined June 5, 2021
1. Backend: a computer in my basement executes a ~10,000 lines Nodejs program that implements the data acquisition pipeline, the majority of these lines deal with web scraping and API fetches, followed in line count by regex and LLM processing, and finally by database indexing. The language models that I use are all large 200B+ parameters and are called via SaaS APIs; they can execute web searches to integrate missing data from the product description pages or from the APIs exposed by the VPS providers; the models are queried multiple times for each question and their responses "vote" for the correct answer until a quorum is reached; this almost completely eliminates data inaccuracies. The final product of this Nodejs program is an SQLite database file that contains the indexed data for the VPS services, which I manually copy to the web server computer (a VPS).
2. Web server: this is a lightweight VPS that runs two processes: a thin Nodejs web application that I quickly wrote without frameworks and runs off simple HTML templates, and a basic C++17 database that serves the web application. Originally the Nodejs application incorporated its own search functions that would linearly swipe thru the data arrays to produce SERPs, but as the number of indexed services grew past 300,000 VPS, the latency became unacceptable (> 1s) and easily DoS-able so I rewrote the search functions in C++17, dropping the query time to ~10ms. This custom DB allows me to easily control low-level details, such as:
2.1 to quantize numerical data to 1 or 2 bytes using a quadratic lossy compressor; the compressor is bijective and totally ordered, so the stored quantized data can be compared as-is against the quantized query parameters, reducing memory use and throughput;
2.2 to compress long strings using a trained ZSTD dictionary, which is convenient for storing millions of offer URLs in memory on a tiny VPS avoiding disk access (URLs are largely repetitive across offers from the same providers, and compress to a 1:10 ratio with a 2kB dictionary);
2.3 to swipe thru presorted indexes, each presorted by a specific criterion, so that the matched data requires no runtime sorting; in comparison tests, MariaDB would consume half of its query time on sorting;
2.4 to efficiently avoid disk access: the SQLite file produced by the backend is read only once, fully, at the start of the process, and subsequent queries are served from memory. The database is read-only and no API permits file write access or memory modification.
A few users query the search engine every day, and it currently generates snack money thru affiliate links. It is a proof of concept more than anything, and I had a lot of fun profiling the database trying out different algorithms to squeeze some decent performance out of my tiny VPS with a shared CPU.
An authority with the effective power to cripple the finances and business practices of Google for the purpose of reducing its market share to an arbitrarily lower percentage would be immediately detrimental to furthering the development of Google's products, and consequently detrimental to the utility of its users, while in the longer term it wouldn't be obvious that conceding the lesser competitors to scramble for this stolen market slice would yield better alternative products, nor that users would ultimately benefit from having more lesser competitors.
This undue confidence or trust we bestow in regulated banks comes at the expense of the wider public: when we lose our deposit due to insolvency, everybody is forced to pay for it by the monetary expansion of the ECB, which covers our loss by printing (or by digitally creating) more euro, and distributes this minted money to the affected bank, allowing withdrawals to resume. This intervention of the ECB dilutes the purchasing power of every unrelated person holding euro, regardless of their country of residence, and regardless on how meticulous they are in choosing a reliable bank. Even the CFA franc in West Africa suffers from this enforced depreciation, being pegged against the euro at a fixed ratio.
Thus, the account holders gain an artificially strong confidence in the bank of their choice, regardless of whichever bank that might be, regardless of how much risk exposures the bank takes, regardless of what financial instruments it edges against, and regardless of the size of its fractional reserve compared to its liabilities. It is an insurance whose only purpose is to allow the bank to take undue risk with volatile instruments and to collectivize any resulting losses against the public, while deluding the account holders with the ancillary excuse that their deposits are safe no matter what.
The alternative to having 21 year old Joe running a million-dollar crypto exchange website from a laptop in his basement is to apply fucking due diligence in choosing our counterparty, and prosecute whenever proper fraud and intentional deception take place.