4,016 karma · joined September 30, 2014
Additionally, my experience with Bedrock hasn't made me a huge fan. If anything its pushed me towards OpenRouter. Way too many 500 errors when we're well below our service quotas.
I know they claim they work, but that's only on their happy path with their very specific AMI's and the nightmare that is the neuron SDK. You try to do any real work with them and use your own dependencies and things tend to fall apart immediately.
It was just in the past couple years that it really became worthwhile to use TPU's if you're on GCP and that's only with the huge investment on Google's part into software support. I'm not going to sink hours and hours into beta testing AWS's software just to use their chips.
This has been an oddly difficult benchmark for Gemini's NB models. Googles images models have always been pretty bad at the studio ghibli prompt, but I'm shocked at how poorly it performs at this task still.
-Linear (transformer) complexity at training time
-Linear scaling with number of tokens
-Online learning(!!!)
The main point that made me cautiously optimistic:
-Empirical results on par with GPT-2
I think this is one of those ideas that needs to be tested with scaled up experiments sooner rather than later, but someone with budget needs to commit. Would love to see HuggingFace do a collab and throw a bit of $$$ at it with a hardware sponsor like Nvidia.