Specifically:
- What hardware caters to 32x disks (24+8)? (I'm picturing enterprise gear)
- What software are you using to coordinate it? TrueNAS?
Specifically:
- What hardware caters to 32x disks (24+8)? (I'm picturing enterprise gear)
- What software are you using to coordinate it? TrueNAS?
It's a fairly typical 1U HP Proliant server, which counts as "enterprise gear". I don't remember the model but it's somewhat unremarkable. It only has like six bays in the front, and I use one for a NixOS install, and two of them with SSDs for read and write cache. It has an old 10 gigabit ethernet card for connectivity. The server I bought used for around $150, the ethernet card I got for about $35 used on ebay.
The actual disks live in 8 bay external chassis, and connected via USB3.0. This may sound horrifying, but USB3 is rated for like 5 gigabits, way more than you can easily get from a home RAID with spinning disks. I have a PCIe 3.0 x16 USB3 card in there, with four plugs, each with a dedicated controller, meaning that in theory each port could get 5 gigabits. The reason I could fairly easily get eight more in there is that I'm only using three of the four ports, so without much effort I could buy 8 more drives (of any size I suppose), put them into another 8 bay chassis, and just add vdev to my RAID. This particular PCIe USB card was $41 on eBay.
I use ZFS for a software RAID.
This sort of happened accidentally; I didn't used to have a rack mount server, I was using a bunch of Nvidia Jetson Nanos, and as such there wasn't a real way directly plug hard drives in, so I ended up having to use the external chassis. Eventually I bought three of those cases, had all the drives in there, and when I decided to buy a "real" server it was substantially easier and cheaper to just add proper USB support to a server than it was to find a decent 24-bay server at a decent price, particularly since I already had the cases.
In hindsight I probably should have got chassis that have eSATA support, but they've been getting decent speed and it's an extremely convenient system. It's also come in handy once when I accidentally broke my server, and I needed to get files off the RAID; all I had to do was plug them into my laptop and mount the ZFS RAID there.
ETA:
Forgot to answer what I use to manage it. I installed NixOS on my server with a pretty vanilla ZFS on Linux install and an NFS mount. I do the tmpfs-on-root trick, so the root partition gets nuked on every reboot. https://elis.nu/blog/2020/05/nixos-tmpfs-as-root/
I don't use any fancy GUI tools or anything; NixOS does a good enough automating away the un-fun parts of maintaining a server, and if I break my server I can very easily roll back to a previous state, and if I really have to, reinstalling it from scratch takes like 20 minutes since I can just copy over my configuration file (which I back up on Gitlab) and have everything automatically set up again. NixOS is pretty cool.
ETAA:
According to `dmidecode` and `lscpu`, server model is `ProLiant DL380 G7`, 24 cores, 128GB of RAM. I think it's actually 12 cores, and it says 24 because of hyperthreading.
It's pretty unremarkable. I like it a lot, but it's a fairly typical and bog-standard used server.
Can you talk more about this? I personally have a Dell R720 and about 90TB available for use, and use 40TB of it for my media server. I'm wondering if that's some use case I want to look into.
I wanted to play with paper trading strategies. I listen on websocket streams of cryptocurrency and stock ticker stuff [1], and the service listening on that socket just plops it into Kafka, with the partition key being the ticker name.
The reason I bother adding Kafka largely comes down to the ability to use the Kafka Streams API. Kafka Streams allows me to do time windows of different trades, or lets me do a sql-style join across two different streams, or lets me filter out data that I don't think is relevant.
Kafka Streams gives you most of the fun of a map-reduce framework, but without having to administer a map-reduce server; it's even smart enough to handle internal state by creating intermediate topics and/or creating local RocksDB instances. It's pretty cool.
There's also has the advantage of allowing me to set retention, so I don't have to worry about doing any kind of manual cleanup; I only have to worry about having N days of stuff on the server, and it's trivial to change that.
Also, I have a number of topics, each with 32 partitions, meaning that if I do find any kind of strategy that works, I can very easily scale it up without changing any code.
I'm basically using Kafka as a streaming-only database. I think it's neat but I haven't actually found any strategies that make or lose money. It's just been a fun way to play around with different bits of server stuff.
[1] https://docs.alpaca.markets/docs/real-time-stock-pricing-dat...
I can read about all the fun theory behind trading math and for distributed systems, and that has some value, but I will understand stuff a lot better if I give myself a real project to do stuff with. I guess I am more of an engineer than a mathematician.