317 karma · joined May 10, 2019
The constraint I, and I bet many here, have is just how much data there is. 3GB like in the 2014 article is one .pdf
Enterprise level data store is measured in hundreds of GB for a single customer, and you'll get murdered on data egress costs if you try to search an entire corpus, if you can even get through it all before the request times out or the customer decides after 5 minutes that enough is enough.
You'd need a true distributed filesystem to even start attempting what the authors suggest at any scale outside of your local machine.
This was my take as well.
My company recently started using Dspy, but you know what? We had to stand up an entire new repo in Python for it, because the vast majority of our code is not Python.
Edited to add: what's hypothetical about Alice being happy to run a coffee shop or Bob satisfied being a 90th percentile engineer (measured how?)? Plenty of these people exist, I've met them!
It's really about finding the entry points into your program from the outside (including data you fetch from another service), and then massaging in such a way that you make as many guarantees as possible (preferably encoded into your types) by the time it reaches any core logic, especially the resource heavy parts.
If something gets mathy I'll use LaTex.
> PReQuaL does not balance CPU load, but instead selects servers according to estimated latency and active requests-in-flight
So, still load balancing
>Are these tools measuring the “right” thing? It doesn’t matter!
Metaphors only go so far. Try to see what I'm really saying here: quality has a cost. Don't shoot yourself in the foot by preemptively reducing quality on account of some ill-conceived notion about how the relationship between product owners and engineering works.
On another note, your value as a human is not the same as the economic value you produce, and it's too easy to conflate the two