Localllm lets you develop gen AI apps on local CPUs
cloud.google.com
cloud.google.com
Technically a wrapper of a wrapper, though llama-cpp-python is quite excellent... TBH I would just recommend using that and all its under-the-radar features (like grammar, and built in function calling support with Functionary).
https://github.com/Mozilla-Ocho/llamafile
You can download any .gguf model (not just the ones in their examples) and run it locally (as long as you have the ram for it). I was running 7B models with ease on an old FX8350 and now 13B models on a 5600X (32GB RAM on both machines).
This wrapper spins up a local web server that runs a simple web frontend to use immediately with no code, but also exposes an OpenAI compatible API for dev work and alt frontends (like SillyTavern).
I see llama.cpp in there, but I didn't go deeper than that. Ollama is not the same thing.
The need to fill more seats on the AI hype train.
(Someone's probably got "release AI software as Google open source" as an OKR)
Functionary is a series of function calling llama architecture finetuned models, not a new base model, if thats what you mean.
Hence, it can be turned into a gguf like any other llama finetune.
It just has a very esoteric prompting format, but llama-cpp-python is specifically set up to leverage it.
> TBH I would just recommend using that and all its under-the-radar features (like grammar, and built in function calling support with Functionary).
I never heard about Functionary, this comment seems to say “I would use llama.cpp plus Functionary”. Now, since llama.cpp allows to plug in several models, I was wondering if Functionary works with several models too, ie is Functionary a library that allows to add function calling to models?
The following answer to my comment helped clarifying that Functionary is not a lib, but a specific model, or that is what I understand from the answer. So the answer is “no, you can’t use functionary with many models (eg it’s not like llama.cpp that works with multiple models) because Functionary is a model itself”
:-)
Very interesting question, as I also read top-level post as saying it's a llama-cpp-python feature, and therefore I could shove some random dolphin-mixtral or whatever I have on my hard drive at it, and it will work.
> run LLMs locally on CPU and memory, right within the Google Cloud Workstation
And in the code https://github.com/GoogleCloudPlatform/localllm/blob/main/qu... there's just a shebang halfway down?
import sys
def main(argv):
pass
if __name__ == '__main__':
main(sys.argv)
#!/usr/bin/env python3
This is awful. Did bard write this?- First, they "steal" @simonw's 'llm' keyword for their own CLI tool (https://news.ycombinator.com/item?id=39296567)
- Then they create a wrapper around llama-cpp-python which is a wrapper around llama-cpp, a project they don't even credit nor support.
- Then they steal @mcapodici's article title to publish it on their blog (https://news.ycombinator.com/item?id=39297563).
- And the ironic part is that their Gemini models are closed-source, so no gguf for those.
Plus I like that it's three letters, since it's designed for CLI usage.
At least it’s not called “ai”, we have that already, e.g. https://github.com/yufeikang/ai-cli
I mean, let’s say you picked say, “ghost pony”; ok, sure, it’s obviously a name.
But LLM? I mean, it’s not like you can get a lot more generic than that. If your tools name is “tool” and some else makes a tool called “tool”, you can’t really complain that much…
I want to know: are they contributing back to llamacpp? Are they paying GG?
Average price per month $966.56
If someone has a microcenter nearby, they could build an 8 core ddr4 system with 128gb of ram for maybe $800.
$3200+ could build a 64 core/128gb ddr5 rig.
https://www.newegg.com/tyan-s8050gm2ne-9534-amd-epyc-9534-2-...
Non-ECC issues are statistically rare. You might never see them if you aren’t running a lot of systems at heavy load at scale.
Also, is this localllm kind of like a Google competitor to some aspects of ollama?
It doesn't support everything other backends need (like certain quantization schemes). If it somehow does, then its not universally compatible and that's just going to create more confusion.
> “…within the Google Cloud Workstation”
Look at this positioning Google use in the blog post: "localllm combined with Cloud Workstations revolutionizes AI-driven application development by letting you use LLMs locally on CPU and memory within the Google Cloud environment." - implying a local setup that's actually reliant on cloud resources is a contradiction.
Maybe they could rename it with a single name that has no conflicts or misleading terms that makes it clear it's to help setup and run inference without confusing the matter about how local things might sometimes be. It won't be nice to read the scathing critiques in some of the other comments! Maybe they could fix up the mess in querylocal.py!
Being able to run the same code locally as in a cloud platform is a nice development pattern but optimal inference setup for many ML situations is a tricky problem, are you really making good use of the CPU / GPU / TPU?
The underlying server tech: llama-cpp-python recently added support for draft models that could make inference much faster. How easy would using that from Localllm be? I'd be concerned `Localllm > llama-cpp-python > llama.cpp` might be too many layers of systems to optimise or code with. Milliseconds often matter.
Having a fast, well priced solution is not easy, people want as many tokens per second as they can get, ideally with no warm-up. How would they solve that?
People looking at fast inference options might like to see Banana's parting words, as they wind down their serverless GPU cloud, in their swansong, they suggest a few hosted inference options people might want to consider: https://news.ycombinator.com/item?id=39288915
I say all that to encourage people to try it out with different models. It's not production quality but totally usable for testing things out even on older machines.
Here's a whole list of indie, OSS AI frameworks you can use: https://github.com/janhq/awesome-local-ai
Disclaimer: I'm one of the core devs on Jan.
Why is that required?
Its actually great. Most backends and frontends "just work" with each other because they all talk with openai (albeit with some caveats).