The post says that 1000 neurons will give you 130 LLM responses - but of what length?
(LLMs are generally priced by input and output tokens. The longer the tokens the longer the compute time. Without an idea of what you mean by a response it's hard to understand.)
Likewise: 1,250 embeddings – How big is the text size in the example?
I'm VERY excited to see you doing this and understand it's early stages, but I wan't wrap my head around the pricing without context.
“neuron” is a cute name but there’s too much conceptual overlap with floating point ops, layers, model parameters etc which are time independent. Should just call them inference credits or something. When some large model runs on multiple GPUs it’s even more confusing what neurons / dollars per second might be.