I may write something like you have described. It would use OpenCL rather than CUDA. Though I would like it to be scalable, the starting goal will be simplicity and ease of deployment on a range of devices.
Do you have any suggestions for the API?
If you wrote some simple code that would work if there was a functioning learning module, what would it look like (if you feel like contributing to this)?