We chose them because they were fairly straightforward to implement and different enough from one another that we could ensure our interface generalized well.
Re benchmarking - at this point we're looking to show directionality, not necessarily blinding speed. We intend to get the tooling feeling right, then work to optimize perf.
Right now, training data comes from the local disk, InfluxDB, or can be piped in from your application via our API. We're looking to build out a set of community-driven components for streaming and processing data. You can learn more about that here - https://github.com/spiceai/data-components-contrib
We'd love for you to contribute!