1,738 karma · joined March 3, 2019
As it stands, responding to user requests on Github just feels exactly like work.
However, I fully recognize how such a process might be abused and the need for a very firm legal hand to avoid accidentally becoming an authoritarian state.
Given the significant weight and danger of death for other drivers, it'll be years until legislators allow their "safety drivers" to be eliminated from the equation. This makes the AI system more akin to enhanced cruise control than robotic trucking.
The two decades during which human oversight of automated systems will be mandatory would be long enough for me to finish off my career getting paid to drive while I sit in a cab writing code, periodically checking over the status of my lead truck and the two or three slaved trucks following me.
If it wasn't for high-powered graphics cards AI would still be a niche university research subject.
Considering the pride Russians proclaim regarding their role in defeating a WWII dictator, it's ironic they refuse to lift a finger to handle the one they raised and enabled at home.
Chris Sawyer's ability to create addictive building games that remain fun to play long after their contemporaries have ended up in the dustbin of history is superhuman, in my humble opinion. Add to the fact that he did it all in Assembly, and it's hard not to place his achievements on a bit of a pedestal.
On the flip side, if you're a data analyst or developer who has a large database with one or more text columns they want results from in a more flexible way than using "LIKE/ILIKE" SQL queries, it's probably easier and faster to create an FTS index/table in that database to get them 90% of the way there.
Without caching, the cost of operating the site would dramatically escalate.
Scraping is a separate subject, but once you write one you can generally reuse relevant portions for many others. If you can get adept at a scraping framework like Scrapy you can do it fairly quickly, but there aren't many tools that work out of the box for every site you'll encounter.
Once you've written the spider, it's generally able to be rerun for updates unless the site code is dramatically altered. It really comes down to how brittle the spider is coded (i.e. hunting for specific heading sizes or fonts or something) instead of grabbing the underlying JSON/XHR that doesn't usually change frequently.
Here's a git repo someone can modify to do a cross comparison on a specific dataset, if they are interested. It doesn't seem to indicate the RMDBs are outclassed in a small-scale FTS implementation.
If you can start with Postgres to have a relational database with the benefit of Full Text Search (i.e. avoid Elastisearch) as well as JSON fields (i.e. avoid MongoDB) then you end up simplifying initial hardware/software requirements while retaining the ability to migrate to those solutions when user demand requires it.
So many developers seem to build with the idea that they'll become the next FAANG when actual (or reasonably forecasted) user load doesn't remotely require such a complex software stack.