Can You Solve This? 1B Records per Second Data Streaming Challenge
nanosai.substack.com
nanosai.substack.com
Can I know (or bound) the number of orders or products in advance to preallocate? Can I design the dataset myself with certain assumptions (e.g. sorted with respect to time)? Can I bound certain aspects of the dataset (e.g. orders must not contain more than 255 products, orders always contain the prices of everything, etc.)?
Latency isn’t apparently a factor - so if I’m processing 1B records, do we care how quickly it gets done? If not, I’ll just stream the data off to a GPU and get the results later?
The answer might be different on different types of hardware, and with different types of data sets, and with different types of data set sculpting. Yes, it is okay to have one benchmark where there are no more than e.g. 255 products, or 255 customers, but then we should probably also benchmark with e.g. up to 65.536 products and 65.536 customers, and up. Part of achieving high performance data streaming is the ability to make your data small.
It would also be okay to use a GPU - although we have not (yet) plans about doing that. Still, it would be very interesting to see what kind of results you could get with that design.
We just have the requirement, that the data streaming engine must not be exclusively designed for this challenge. It must be a reasonably functional general purpose data streaming engine.
By the way, we hope to reach the 1 BRS milestone on a single server, i7-6700 Quad-Core Skylake CPU, with 2 NVME SSDs mounted in RAID 1. 1 GB of memory to run the benchmark app should be enough, but the server will probably have 64 GB by default.
For instance, simply writing a program that loads 1 billion bytes into memory, iterates them and sums them, would not count as a "general purpose data streaming engine". But you don't have to use Spark, Kafka or something like that. You can write your own.
Also, kind of funny that C/C++ are not listed in the application form.
The challenge is not for others to write a data streaming engine to give to us. The challenge is for us to write a data streaming engine and give to you! ... but if you want to try beating the 1BRS milestone too - that would be fun too :-)
Contest closes ---- Nanosai advertises that their platforms now supports 1B requests per sec
I also love that you include NO INCENTIVE with this 'challenge'. Why would anyone submit code to you guys for free?
All smells fishy to me.
Also, yes the streaming engine will be useful to some use cases for Kahler. However, it is open source and so anyone will be able to access the same underlying engine.
By the way, we already have several people signed to the challenge including people working for notable tech companies!:)
WE (at Nanosai.com) will attempt to build a data streaming engine that can process 1 billion records per second, and release it as open source. If YOU want to try the challenge too, that's fine (e.g. someone already working on data streaming engine tech). We did not mean for YOU to solve this problem for US.
Our initial measurements and calculations show that it should be possible to reach 1BRS, although the records would have to be small. Still, a data streaming engine will always have some record iteration overhead, so it would take some tuning to get that overhead small enough to reach 1BRS even with 1 byte records.
Ignoring use of a GPU, how many IPS is a quad core i7 (mentioned by jjenkov)?
And how many instructions might it take to do something useful to a record? Say read 16 bytes and do some compares.
Or would the SSDs be the bottleneck? Also, RAID1, not RAID0, so effectively just 1 SSD.
Or NVMe? Wow, a google search says vendors are pushing for 32 GBps. (My SATA setup is obsolete.)
Maybe it's purely a software problem.
Also, it would probably help if it was written in a programming language appropriate for the purpose, such as C or C++.