Genuine question: is a ML prediction slower than to fetch a row from a database and/or more difficult to scale?
Genuine question: is a ML prediction slower than to fetch a row from a database and/or more difficult to scale?
Specifically, this model is an ensemble of decision trees.
This involves
1. Row lookup
2. N (tree depth) inequality checks on fields in the row for M trees
3. Weighted sum over M trees
> XGBoost cannot serve predictions concurrently because of internal data structure locks. This is common to many other machine learning algorithms as well, because making predictions can temporarily modify internal components of the model.
A DBMS that does mostly point-query reads — fetching single "cold" rows from a large storage, with high concurrency — is an IO-bottlenecked workload. Given a sufficiently mainframe-y "DMA from disk straight through to the NIC" architecture (which modern microcomputer servers are rapidly approaching), the CPU is just "guiding" that transfer — acting as a PCIe switchboard operator — rather than actually performing it. The only things you need to worry about are JBOD IO-queue depth and (multi-aggregated) NIC total ring-buffer size. This means that you can go very far with "vertical" scaling for these workloads, in the sense that you can shove more disks and NICs into the same machine and it'll be able to serve more RPS, without really increasing CPU load. For IO-bottlenecked workloads, you only start to suffer when your PCIe bandwidth is exhausted. This is why systems built for "SAN" or "DB" or "hyperconverged" use-cases like this always use dual-socket lower-cores-per-socket CPUs (lately: dual Epyc 7532s) rather than single higher-core-count CPUs — as each socket gets its own PCIe channels, so the lower-powered dual-socket CPUs actually maximize PCIe bandwidth per machine.
Meanwhile, any kind of ML prediction is going to be a CPU- (or maybe GPU-) bottlenecked workload. If you have 64 CPU cores in a machine, then you can only be doing 64 ML predictions at a time on that machine. You might be able to resolve any individual prediction "quickly"; but unlike an IO-bottlenecked workload, that "quickly" doesn't involve the worker thread ever blocking + yielding the CPU to other workloads while waiting on anything; it's solid 100%-pinned CPU-cycles on that core, all the way through the prediction. There is no advantage in oversubscribing the machine with workloads; running KN predictions across N cores makes each query take (slightly more than) K times as long to run. You may as well just make each machine the smallest size it can reasonably be to run the model (as each machine then gets its own IO memory and IO "breathing room", and you avoid the price-premium of using latest-gen parts), stamp out horizontal replicas of it, and load-balance between them.
This is why so many phones and other low-power devices are shipping with "AI/ML accelerators": putting ASICs in, tuned to run specific ML-model architectures faster, can vastly lower the power budget for the device as a whole, as running (concurrent) models on the device would otherwise have to ramp up the CPU to its highest frequency and pin it at 100% to do one of these predictions.
Shared RAM across all cores can be a reason to use fewer larger machines rather than more smaller ones in such a case. Postgres does give you options either way though.