Opinions are our own and not of Google DeepMind.
Opinions are our own and not of Google DeepMind.
I don't know who writes Google's documentation or does the copyediting for their console, but it is hard to adapt. I have spent hours troubleshooting, only to find out it's because the documentation is referring to the same thing by two different names. It's 2024 also, I shouldn't be seeing print statements without parentheses.
Each thing seems to have a bunch of clicks to setup that startup LLM providers don't hassle people with. They're more likely to just let you sign in with some generic third party oAuth, slap on Stripe billing, let you generate keys, show you some usage stats, getting started docs, with example queries and a prompt playground etc.
What about the Vertex models though? Are they all actually available via Google AI Studio?
Billing for the Gemini models (on Vertex AI, the Generative Language AI variant still charges by tokens) I would argue is simpler than every other provider, simply because you're charged by characters/image/video-second/audio-second and don't need to run a tokenizer (if it's even available cough Claude 3 and Gemini) and having to figure out what the chat template is to calculate the token cost per message [2] or figure out how to calculate tokens for an image [3] to get cost estimates before actually submitting the request and getting usage info back.
[1]: https://cloud.google.com/vertex-ai/generative-ai/docs/multim...
[2]: https://platform.openai.com/docs/guides/text-generation/mana...
[3]: https://platform.openai.com/docs/guides/vision/calculating-c...
I'm also not sure if I understand your problem with pricing? Depending on what you do with it, it's not just an LLM. It actually started before llms.
Pricing for image classification and other features are completely different products like an LLM.
I use gcp professional every day and always found it quite intuitive.
Did plenty of image classification with vertex ai too
It's a shitty solution to a stupid problem ;)
But I did mention that vertex AI is more than just hosting llms though
You can try 27b at www.aistudio,google.com. Send in your favorite prompts, and we hope you like the responses.
The Google API models support 1M+ tokens, but these are just 8K. Is there a fundamental architecture difference, training set, something else?
How much is pre-training dataset changes, how much is tuning?
How do you think about this problem, how do you solve it?
Seems tricky to me.
Literature has identified self-proliferation as dangerous capability of models, and details about how to define it and example of form it can take have been openly discussed by GDM (https://arxiv.org/pdf/2403.13793).
Current Gemma 2 models' success rate to end-to-end challenges is null (0 out 10), so the capabilities to perform such tasks are currently limited.
Or does the model have to later be finetuned, to not be good at certain tasks?
Or are we not at that stage yet?
Is something like tree-of-thought used, to get the best of the models for these tasks?
Is this a contradiction or am I misunderstanding something?
Btw overall very impressive work great job.
However, I wouldn't draw conclusions about different model families, like Llama and Gemma, based on their token count alone. There are many other variables at play - the quality of those tokens, number of epochs, model architecture, hyperparameters, distillation, etc. that will have an influence on training efficiency.
Still no 27B 4-bit GGUF quants on HF yet!
I'm monitoring this search: https://huggingface.co/models?library=gguf&sort=trending&sea...