Forget ChatGPT: why researchers now run small AIs on their laptops
nature.com
nature.com
https://future.mozilla.org/builders/news_insights/introducin...
https://github.com/Mozilla-Ocho/llamafile
They even have whisperfiles now, which is the same thing but for whisper.cpp, aka real-time voice transcription.
You can also take this a step further and use this exact setup for a local-only co-pilot style code autocomplete and chat using Twinny. I use this every day. It's free, private, and offline.
https://github.com/twinnydotdev/twinny
Local LLMs are the only future worth living in.
It's been pretty flawless, and honestly pretty darn useful here and there. The big guns go faster and do more, but I'd prefer not having every interaction logged etc.
6core 8th gen i7 I think, with a 1050ti. Old stuff. And it's quick enough on the smaller 7/8b models for sure.
Sorry for the blatant ad, though I do hope it's useful for some ppl reading this thread: https://getfluid.app
Roadmap is following:
- October - private remote AI (when you need smarter AI than your machine can handle, but don't want your data to be logged or stored anywhere)
- November - Web search capabilities (so the AI will be capable of doing websearch out of the box)
- December - PDF, docs, code embedding. 2025 - tighter MacOS integration with context awareness.
What are some recommendations for running models locally, on decent CPUs and getting good valuable output from them? Is that llama stuff portable across CPUs and hardware vendors? And what do people use it for?
https://github.com/jart/cosmopolitan
"Cosmopolitan Libc makes C a build-once run-anywhere language, like Java, except it doesn't need an interpreter or virtual machine. Instead, it reconfigures stock GCC and Clang to output a POSIX-approved polyglot format that runs natively on Linux + Mac + Windows + FreeBSD + OpenBSD + NetBSD + BIOS with the best possible performance and the tiniest footprint imaginable."
I use it just fine on a Mac M1. The only bottleneck is how much RAM you have.
I use whisper for podcast transcription. I use llama for code complete and general q&a and code assistance. You can use the llava models to ingest images and describe them.
> … by combining llama.cpp with Cosmopolitan Libc into one framework that collapses all the complexity of LLMs down to a single-file executable (called a "llamafile") that runs locally on most computers, with no installation.
Low cost to experiment IMO. I am personally using MacOS with an M1 chip and 64gb memory and it works perfectly, but the idea behind this project is to democratize access to generative AI and so it is at least possible that you will be able to use it.For tooling to train models it's a bit more difficult but inference works great on AMD.
My CPU is an AMD Ryzen and the OS Linux. No problem.
I use OpenWebUI as frontend and it's great. I use it for everything that people use GPT for.
Ollama works on anything: Windows, Linux, Mac and Nvidia or AMD. I don't know if other cards like Arc are supported by anything yet, bit of it supports the open Vulkan API (like AMD) then it should work.
Every inference server out there supports running from CPU, but realize that it's much slower than running on a GPU - that's why this revolution didn't begin until GPUs became powerful and affordable.
As far as being clear to setup, Ollama is trivial: it's a single command line that only asks what model you want and they provide you with a list on their website. They even have a Docker container if you don't want to worry about installing any dependencies. I don't know what could be easier than that.
Most other tools like LM Studio or Jan are just a fancy UI running llama.cpp as their server and using HuggingFace to download the models. They don't even offer anything beyond simple inference, such as RAG or agents.
I've yet to see anything more than a simple RAG that's available to use out of the box for local use. The only full service tools are online services like Microsoft Copilot or ChatGPT. Anyone else who wants to do that more advanced kind of system ends up writing their own code. It's not hard if you know Python - there are lots of libraries available like HuggingFace, LangChain, and Llama-Index, as well as millions of tutorials (every blog has one).
Maybe that's a sign that there's room for an open source platform for this kind of thing, but given that it's a young field and everyone is rushing to become the next big online service or toolkit, there might not be as much interest from developers to build an open source version of a high quality online service.
Which is an artificial restriction from MS that's really easily bypassed.
Personally I don't care whether the telemetry is identifiable. I just don't want it.
> extensions may be collecting their own usage data and are not controlled by the telemetry.telemetryLevel setting. Consult the specific extension's documentation to learn about its telemetry reporting and whether it can be disabled.
They must have reintroduced the telemetry setting. I can't remember if I deleted the old one, but my setting on that new value was set to "all" by default.
But I don't want it. I want my software to work for me, not against me.
>and some of the most popular extensions are only available with VSCode, not with Codium.
I'll manage without them. What's especially annoying is that this restriction is completely artificial.
Having said that, MS did a great job with VsCode and I applaud them for that. I guess nothing is perfect, and I bet these decisions were made by suits against engineer wishes.
How is said software working "against" you by collecint non-personal telemetry while purpose of that telemetry usually is making the software better for most users?
that usually hasn't been the case since at least a decade. it's truly bewildering that someone especially on hackernews would voluntarily give big tech there finger and not expect to get bitten.
Point in case: Software actually got worse.
Second point in case: Great software and editors have been built without telemetry for decades.
"How is that chair working 'against' you by collecting 'non-personal' sitting patterns tagged with timestamps and information about the chair and house that it's in while the purpose of that data collection 'usually' is making the chair better for other people?"
When I use a product, I'm not implicitly inviting the makers of that product to perpetually monitor my usage of the product so that they can make more money based on my data. In any other part of life other than software, this would be an obscene assumption for a product maker to make. But in software, people give it a pass.
No.
This type of data collection is obscene when informed consent is not clearly and authoritatively acquired in advance.
We know from the body of work in deobfuscation that there's no such thing as "strictly anonymous metrics".
I don't care, I don't want my text editor to send _any_ telemetry, _especially_ without my explicit consent.
> some of the most popular extensions are only available with VSCode
This has never been an issue for me, fortunately. The only issue is Microsoft's proprietary extensions, which I have no interest in using either. If I wanted a proprietary editor I'd use something better.
So yeah, I'll use Excel to interoperate with fancy spreadsheets, but if LibreOffice will do the job, I'll use it instead. I tried out several of the fancy proprietary editors at various times (SublimeText, VSCode, even Jetbrains), but IMO they were not better _enough_ to justify switching away from something like vim, which is both ubiquitously available and FOSS.
The biggest benefits for me are the uncensored models. I'm pretty kinky so the regular models tend to shut me out way too much, they all enforce this prudish victorian mentality that seems to be prevalent in the US but not where I live. Censored models are just unusable to me which includes all the hosted models. It's just so annoying. And of course the privacy.
It should really be possible for the user to decide what kind of restrictions they want, not the vendor. I understand they don't want to offer violent stuff but 18+ topics should be squarely up to me.
Lately I've been using grimjim's uncensored llama3.1 which works pretty well.
Not all uncensored models are great. Some return very sparse data or don't return the end tags sometimes so they keep hallucinating and never finish.
If you import grimjim's model, make sure you use the complete modelfile from vanilla lama3.1, not just an empty modelfile. Because he doesn't provide one. This really helps setting the correct parameters so the above doesn't happen so much.
But I have seen it happen with some official ollama models like wizard-vicuna and dolphin-llama. They come with modelfiles so they should be correct.
It's just that the LLMs trigger immediately on minor words and shut down completely.
User: Hey, how are you?
Llama: [object Object]
It's funny but I don't think I did anything wrong?2010: Javascript is webservers.
2020: Javascript is desktop applications.
2024: Javascript is AI.
Trying to chat to an INSTRUCT model will be disappointing, much as you describe.
I don't know if it was some sort of error on the UI or what.
Trying to interrogate it about the first message yielded no results. It just repeated back my question, verbatim, unlike the rest of the chat which was more or less chat-like :shrugh:
So do you think government doesn't use networks?
[1] https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
It lets you use local llama.cpp without setup, chat with PDF offline and provides chat history / nested folders chat organization, and can handle thousands of conversations. In addition you can import your ChatGPT history and continue chats with local AI.
this was with 4bit quantization and offload of as many layers as possible to the gpu.
in any event it was fun to try out, but still didn't seem anywhere near how well the hosted models work. a heavy duty workstation with a bunch of gpus/vram would probably be a different story though.
You could try a model that fits entirely into in VRAM. It"s a trade of precision for a decent bit of performance. 16GB is plenty to work with as i've seen acceptable enough results with 7B models on my 8GB GPU.
If you use their API instead of their sub-based offers, the most popular models are cheap to use and with BYOK tools, switching model is as easy as entering another string in a form.
For instance I put $15 on my OpenAI account in August 2023, since then I used Dall-E weekly and I still got more than $5 credit left!
I've found that you can get faster performance by choosing a smaller model and/or by using a smaller quantization. You can use other models with llamafile as well. They have some prebuilt ones:
https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file...
You can also search for other llamafiles for other models on HuggingFace by using the llamafile tag.
https://huggingface.co/models?library=llamafile&sort=trendin...
And you can download model weights directly and use them by providing an -m flag to llamafile but that's getting a bit less straightforward.
https://github.com/Mozilla-Ocho/llamafile?tab=readme-ov-file...
I have now learned that my laptop is capable of a whopping 0.37 tokens per second.
11th Gen Intel® Core™ i7-1185G7 @ 3.00GHz × 8
When the article says that researchers are using their laptops those researchers are either using very small models on a gaming laptop or they have a fairly modern MacBook with a lot of ram.
There are also options for running open LLMs in the cloud. Groq (not to be confused with Grok) runs Llama, Mixtral and Gemma models really cheaply: https://groq.com/pricing/
Groq looks interesting and might be a better option for me. Thank you.
If anyone is reading this and had trouble with a larger model, that might be the one to try next.
> Subject to the Agreement, Company grants you a limited license to reproduce portions of Company Properties for the sole purpose of using the Services for your personal, non-commercial purposes.
[1] It feels super strange to talk to yourself, but luckily I'm out early enough that I'm often alone. Worst case, I pretend I'm talking to someone on the phone.
Still very reasonable on modern hardware.
I'm on a Mac and I found the easiest way to run & use local models is Ollama as it has a rest interface: https://github.com/ollama/ollama/blob/main/docs/api.md
I just have a local script that pulls the audio file from Voice Memos (after it syncs from my iPhone), runs it through openai's whisper (really the best at voice to speech; excellent results) and then makes sense of it all with a prompt that asks for organized summary notes and todos in GH flavored markdown. That final output goes into my Obsidian vault. The model I use is llama3.1 but haven't spent much time testing others. I find you don't really need the largest models since the task is to organize text rather than augment it with a lot of external knowledge.
Humorously the harder part of the process was finding where the hell Voice Memos actually stores these audio files. I wish you could set the location yourself! They live deep inside ~/Library/Containers. Voice Memos has no export feature, but I found you can drag any audio recording out of the left sidebar to the desktop or a folder. So I just drag the voice memo into a folder my script watches and then it runs the automation.
If anyone has another, better option for recording your voice on an iPhone, let me know! The nice thing about all this is you don't even have to start / stop the recording ever on your walk... just leave it going. Dead space and side conversations and commands to your dog are all well handled and never seem to pollute my notes.
Also what kind of local machine do you need? I have an imac pro, wondering if this will run the models or if I ought to be on an apple silicon machine? I have an M1 macbook air as well.
I love it and I sincerely hope that "Apple Intelligence" won't kill the button and replace it with a sub-viable conversational model, but I probably ought to figure out local whisper sooner rather than later because it's probably inevitable.
> Finding the Truth – Surprisingly, my iZYREC revealed more than I anticipated. I had placed it in my husband's car, aiming to capture some fun moments, but it instead recorded intimate encounters between my husband and my close friend. Heartbreaking yet crucial, it unveiled a hidden truth, helping me confront reality.
> A Voice for the Voiceless – We suspected that a relative's child was living in an abusive home. I slipped the device into the child's backpack, and it recorded the entire day. The sound quality was excellent, and unfortunately, the results confirmed our suspicions. Thanks iZYREC, giving a voice to those who need it most.
Is this a physical button or on-screen? I’ve been rewatching Twin Peaks recently and would love a high-tech implementation of Cooper’s tape recorder.
Check out the new voice transcription feature in iOS 18. On my SE 2022 (very much not high-end), I have a good-size record/pause button on screen.
After you’re done with the voice recording, it gives you a transcription of what you spoke.
I remember the first lecture in the Theory of Communication class where the professor introduced the idea that communication by definition requires at least two different participants. We objected by saying that it can perfectly be just one and the same participant (communication is not just about space but also time), and what you say is a perfect example of that.
I put together a script that takes any audio file (mp3, wav), normalizes it, runs it through ggerganov's whisper, and then cleans it up using a local LLM. This has saved me a tremendous amount of time. Even modestly sized 7b parameter models can handle syntactical/grammatical work relatively easily.
Here's the gist:
https://gist.github.com/scpedicini/455409fe7656d3cca8959c123...
EDIT: I've always talked out loud through problems anyway, throw a BT earbud on and you'll look slightly less deranged.
Which local LLM do you use?
Edit:
And self talk is quite a healthy and useful thing in itself, but avoiding it in public is indeed kind of necessary, because of the stigma
I do a lot of stargazing and have experimented with voice memos for recording my observations. The problem of course is later going back and listening to the voice memo and getting organized information out of what essentially turns into me rambling to myself.
I'm going to try to use whisper + AI to transcribe my voice memos into structured notes.
That's how I'm writing this message to you.
Learning to use these speech-to-text systems will be a new kind of literacy.
I think pushing the transcription through language models is a fantastic way to deal with the complexity and frankly, disorganization of directly going from speech to text.
By doing this we can all basically type at 150-200 words a minute.
Neat. Can you explain your setup a little? How do you go from voice to whisper to writing in this reply input form on a webpage?
If it matters, I managed to install Jan AI on my Linux Mint and able to use the Mistral model. I use an Android phone if it helps. Thanks.
As a daily user of all the latest OpenAI/Claude models for the last few years, I'm amazed at how good llama3.1 is - all running locally, privately, with zero connection to the web. How did I not know about this?!
- https://videocardz.com/newz/amd-ryzen-ai-max-395-to-feature-...
It seems that the thermal design power for Strix Halo can be configured between 55 W and 120 W, which is similar to the power used now by a combo laptop CPU + discrete GPU.
PC Part Picker, DDR5-8400 48 GB (2x24GB) is... $340 right now.
For $680 you can get 96 GB of very fast RAM.
How about someone make an NVidia GPU with 96 GB of RAM at a reasonable price? Please?
The one you listed does around 50Gbps. A really good gpu does almost 450Gbps. Prices as you know also don’t scale linearly. For something twice as good sometimes you pay 4x the price and so on.
Large.
Cheap.
You may only pick two.
Can Apple Silicon manage this? Would it be feasible to do with some quantization perhaps?
You can get better performance using a good CPU + 4090 + offloading layers to GPU. However one is a laptop and the other is a desktop...
I'm not sure what the impact is on a 70b model but it seems there's a lot of exaggeration going on in this space by Mac fans.
The results for Llama 2 70B Q4_0 (39GB) was 8.5 tok/s for text generation (you'd expect a theoretical max of a bit over 10 tok/s based on theoretical MBW) and a prompt processing of 19 tok/s. On a 4K context conversation, that means you would be waiting about 3.5min between turns before tokens started outputting.
Sadly, I doubt that Strix Halo will perform much better. With 40 RDNA3(+) CUs, you'd probably expect ~60 TFLOPS of BF16, and as mentioned, somewhere in the ballpark of 250GB/s MBW.
Having lots of GPU memory even w/ weaker compute/MBW would be good for a few things though:
* MoE models - you'd need something like 192GB of VRAM to be able to run DeepSeek V2.5 (21B active, but 236B in weights) at a decent quant - a Q4_0 would be about 134GB to load the weights, but w/ far fewer activations, you would still be able to inference at ~20 tok/s). Still, even with "just" 96GB you should be able to just fit a Mixtral 8x22B, or easily fit one of the new MS (GRIN/Phi MoEs).
* Long context - even with kvcache quantization, you need lots of memory for these new big context windows, so having extra memory for much smaller models is still pretty necessary. Especially if you want to do any of the new CoT/reasoning techniques, you will need all the tokens you can get.
* Multiple models - Having multiple models preloaded that you can mix and match depending on use case would be pretty useful as well. Even some of the smaller Qwen2.5 models looks like they might do code as well as some much bigger models, you might want a model that's specifically tuned for function calling, a VLM, SRT/TTS, etc. While you might be able to swap adapters for some of this stuff eventually, for now, being able to have multiple models pre-loaded locally would still be pretty convenient.
* Batched/offline inference - being able to load up big models would still be really useful if you have any tasks that you could queue up/process overnight. I think these types of tools are actually relatively underexplored atm, but has as many use cases/utility as real-time inferencing.
One other thing to note is that on the Mac side, you're mainly relegated to llama.cpp and MLX. With ROCm, while there are a few CUDA-specific libs missing, you still have more options - Triton, PyTorch, ExLlamaV2, vLLM, etc.
Wouldn't the time be negligible with interturn kv caching? Many inference providers already do this.
- "Running Qwen 2.5 Math 72B distributed across 2 MacBooks. Uses @exolabs_ with the MLX backend." https://x.com/ac_crypto/status/1836558930585034961
If you require an x86-64 based mobile solution with CUDA support, the maximum VRAM available is 16GB. The Strix HALO is positioned as a competitor to the RTX 4070M.
"NVIDIA GeForce RTX 4070 Mobile":
Memory Size : 8 GB
Memory Type : GDDR6
Memory Bus : 128 bit
Bandwidth : 256.0 GB/s
"NVIDIA GeForce RTX 4090 Mobile" Memory Size : 16 GB
Memory Type : GDDR6
Memory Bus : 256 bit
Bandwidth : 576.0 GB/sYou realize it'll still be much faster than trying to run larger models on system RAM?
So who will be interested in a shitty assistant next year when you can have an amazing one, is what I wonder? Is this just the biggest cup of wishful thinking that we have ever seen?
It is also evident in the moderation that your usage is subject to human review and I don't think that should even be possible.
There are different use cases and computers are already pretty powerful. Maybe your local model won't be able to produce tests that check all the corner cases of the class you just wrote for work in your massive code base.
But the small model is perfectly capable of summarizing the weather from an API call and maybe tack on a joke that can be read out to you on your speakers in the morning.
They want compliant Linux drivers?
If I’ve raised $1B to buy GPUs and train a “bigger model”, a major part of my competitive advantage is having $1B to spend on sufficient GPUs to train a bigger model.
If, after having raised that money it becomes apparent that consumer hardware can run smaller models that are optimized and perform as well without all that money going into training them, how am I going to pivot my business to something that works, given these smaller models are released this way on purpose to undermine my efforts?
It seems there are two major possibilities: one, people raising billions find a new and expensive intelligence step function that at least time-locally separates them from the pack, or two (and significantly more likely in my view) they don’t, and the improvements come from layering on different systems such as do not require acres of GPUs, while the “more data more GPUs” crowd is found to have hit a nonlinearity that in practical terms means they are generations of technology away from the next tier.
edit: I guess to your point if it is not knowingly then the electricity costs are not a factor either.
Only with memcoins.
Scepticism is fine, if it's plausible. If not it's conspiratorial.
1) optimizing the model training
2) optimizing the model operation
The $1B-spend holy grail is that it costs a lot of money to train, and almost nothing to operate, a proprietary model that benchmarks and chats better than anyone else’s.
OpenAI’s optimizations fall into the latter category. The risk to the business model is in the former — if someone can train a world-beating model without lots of money, it’s a tough day for the big players.
To see what optimizing model operation looks like, groq is a good example. OpenAI isn’t (yet) obviously in that kind of optimization, though I’m sure they’re working on it internally.
I would roll data acquisition/cleaning processes into training costs for purposes of this because what else is the data for if not training?
If 4o wasn’t an optimization for model operation costs what was it?
Leave the problems that require competent reasoning ability to the larger models.
yi: previously non-commercial but Apache 2.0 now?
deepseek: usage policy
larger gemma models: usage policy
databricks: similar to llama: no more than 700 million MAUs, usage policy, additional restrictions on using outputs to train other models
qwen: no more than 100 million MAUs
mistral: non-commercial
command-r: CC BY-NC
starling: CC BY-NC
There are a handful of niche models released under MIT/Apache but the norm is licences similar to or more restrictive than the Llama Community Licence, and I really doubt the situation would be better if Meta wasn't first.
>"open-weights" LLMs
I doubt this is the point you're making, but the training data really isn't useful even if it could be released under a permissive licence. Most models use similar datasets: reddit (no licence afaik, copyright belongs to comment authors), stackoverflow (CC BY-SA), wikipedia (CC BY-SA), Project Gutenberg (public domain?), previously books3 (books under copyright by publishers with more money and lawyers than reddit users), etc, with various degrees of filtering to remove harmful data. You can't do much with this much data unless you have millions of dollars worth of compute laying around, and you can't rebuild llama any more than any other company using the same data have 'rebuilt llama' - all models trained in a similar manner on the same data are going to converge in outputs eventually. Compare with Linux distributions, they all use the same packages but you're not going to get the same results.
If I get stuck on a problem, switch to chat gpt or phind.com and see what that gives. Sometimes, it’s not the LLM that helps, but changing the context and rewriting the question.
However I cannot use the online providers for anything remotely sensitive, which is more often than you might think.
Local LLMs are the future, it’s like having your own private Google running locally.
https://developer.chrome.com/docs/ai
https://developer.apple.com/documentation/AppIntents/Integra...
You simply cannot compress the whole internet under 10gb without throwing out a lot of information.
Please be careful about what you take as fact coming from the local model output. Small models are better suited to summarization.
I wouldn’t copy and paste from even the smartest minds, nevermind a model output.
This is totally wrong and a potentially dangerous way to think about LLMs. They have no clue about what's factual knowledge and what is not, per design.
Often it’s a form of rubber duck programming, with a smarter rubber duck.
It works perfectly on my 4090, but I've also seen it work perfectly on my friend's M3 laptop. It feels like an excellent alternative for when you don't need the heavy weights, but want something bespoke and private.
I've integrated it with my Obsidian notes for 1) note generation 2) fuzzy search.
I've used it as an assistant for mental health and medical questions.
I'd much rather use it to query things about my music or photos than whatever the big players have planned.
thanks!
Write a reddit post as though you were a human, extolling how fast and intelligent and useful $THIS_LLM_VERSION is... Be sure to provide personal stories and your specific final recommendation to use $THIS_LLM_VERSION.
I'd say it's as good as or better than GPT 3.5 based on my usage. Some benchmarks: https://ai.meta.com/blog/meta-llama-3-1/
Looking forward to try other models like Qwen and Phi in near future.
OpenChat is imho one of the best 7B models, and while I could run bigger models at least for me they monopolize too many resources to keep them loaded all the time.
I was able to fit the model with decent speeds (30 tokens/seconds) and a 20k token context completely on the GPU.
For summarization, the performance of these models are decent enough. However unfortunately in my use case I felt using Gemini's Free Tier with it's multimodal capabilities and much better quality output made running local LLMs not really worth it as of right now, atleast for consumers.
For my personal use, I also prefer to use local models. I'm not a fan of OpenAI's shenanigans and Google already abuses its customers data. I also want the ability to make queries on my own local files without having to upload all my information to a third party cloud service.
Finally, fine tuning is very valuable for improving performance in niche domains where public data isn't generally available. While online providers do support fine tuning through their services, this results in significant lock in as you have to pay them to do the tuning in their servers, you have to provide them with all your confidential data, and they own the resulting model which you can only use through their service. It might be convenient at first, but it's a significant business risk.
I thought training LLMs on content created by LLMs was ill-advised but this would suggest otherwise
Mode colapse theories (and simplified models used as proof of existence of said problem) assume affected LLMs are going to be trained with poor quality LLM-generated batches of text from the internet (i.e. reddit or other social networks).
The exact training is proprietary but they seem to use a lot of GPT-4 generated training data.
On that note... I've often wondered if broad memorization of trivia is really a sensible use of precious neurons. It seems like a system trained on a narrower range of high quality inputs would be much more useful (to me) than one that memorized billions of things I have no interest in.
At least at the small model scale, the general knowledge aspect seems to be very unreliable anyways -- so why not throw it out entirely?
You can get diverse low quality data from the web, but for diverse high quality data the organic content is exhausted. The only way is to generate it, and you can maintain a good distribution by structured randomness. For example just sample 5 random words from the dictionary and ask the model to compose a piece of text from them. It will be more diverse than web text.
I agree if we are talking about maxing raw reasoning and logical onference abilities, but the problem is that the ship has sailed and people expect llms to have domain knowledge (even more than expert users are clamoring for LLMs to have better logic).
I bet a model with actual human “intelligence” but no Google-scale encyclopedic knowledge of the world it lives in would be scored less preferentially by the masses than what we have now.
Phi is another really good example but that's already covered from the article.
[0] - https://www.latent.space/i/146879553/synthetic-data-is-all-y...
Millions? Where are they? Where are they used?
What I think will be interesting is seeing which of the open models stick around and for how long when we have super easy ‘good enough’ models that provide quality integration. My bet is not many, sadly. I’m sure Llama will continue to be developed, and perhaps Mistral will get additional European government support, and we’ll have at least one offering from China like Qwen, and Bytedance and Tencent will continue to compete a-la Google and co. But, I don’t know if there’s a market for ten separately trained open foundation models long term.
I’d like to make sure there’s some diversity in research and implementation of these in the open access space. It’s a critical tool for humans, and it seems possible to me that leaders will be able to keep extending the gap for a while; when you’re using that gap not just to build faster AI, but do other things, the future feels pretty high volatility right now. Which is interesting! But, I’d prefer we come out of it with people all over the world having access to these.
Only iPhone 15 Pro or later will get Apple Intelligence, so the number will be wayyy smaller.
When people describe it as a "critical tool" i feel like I'm missing basic information about how people use computers and interact with the world. In what way is it critical for anything? It's still just a toy at this point.
Covers 90% of OCR needs with 10% of the effort. No API keys, scripting, or network required.
https://www.reddit.com/r/LocalLLaMA/comments/1ej9uzh/local_l...
Admittedly it's slow (3.5 token/sec)
You should expect somewhere around 30t/s for a single response, if running the FP8 rowwise quant that would typically be used on such a node, with TensorRT-LLM. Massively more in total with batching.
That quant is twice the size as the 4.5bpw one used on the Mac though. A lower quality one would be faster.
It is about the only thing I can do on my M1 Pro to spin up the fans and make the bottom of the case hot.
Llama3.1, Deepseek Coder v2, and some of the Mistral models are good.
ChatGPT and Claude top tier models are still better for very hard stuff.
Also is it sensible to wait for newer mac, amd, nvidia hardware releasing soon?
The last one I used was Llama 3.1 8B which was pretty good (I have an old laptop).
Has there been any major development since then?
Llama8b is the new mistral.
There are same issues with GPT API,
1. non reproducible is there in the API
2. even after we ensure we do a moderation check on the input prompt, soemtimes GPT will produce "unsafe" output and accuse itself of "unsafe" stuff and we get an error but we pay for GPT "un-safeness" IMO if the GPT is producing unsafe stuff then I should not pay for it's problems.
3. dalle gives no seed so no reproducible, and no option to opt out on their GPT modifying the prompt , so images are sometimes absurdly enhanced with extreme amount of details or extreme diversity, so you need to fight against their GPT enhancing.
What extra option we have with the APIs that is useful ?
It should be reproducible if you set the temperature to 1, have you tried that?
With regard to DALLE - that's a fair complaint I didn't realize they don't have a seed for their API. You should really try switching to an open model if you can. You'll have complete control. I recommend flux-schnell or flux-dev.
"2 MacBooks is all you need. Llama 3.1 405B running distributed across 2 MacBooks using @exolabs_ home AI cluster" https://x.com/AIatMeta/status/1834633042339741961
And not having tried it, I’m guessing it will probably run at 1-2 tokens per second or less since the 70b model on one of these runs at 3-4, and now we are distributing the process over the network, which is best case maybe 40-80Gb/s
It is possible, and that’s about the most you can say about it.
I have a few thousand book, papers and articles collected over the last decade. And while I have meticulously categorised them for fast lookup, it's getting harder and harder to search for the desired info, especially in categories which I might not have explored recently.
I do have a 4070 (12 GB VRAM), so I thought that LLMs might be a solutions. But trying to figure out the whats and hows hase proven to be extremely complicated, what with deluge of techniques (fine-tuning, RAG, quantisation) that might not might not be obsolete, too many grifters hawking their own startups with thin wrappers, and a general sense that the "new shiny object" is prioritised more than actual stable solutions to real problems.
Segment the texts into chunks that make sense (i.e. into the lengths of text you'll want to find, whether this means chapters, sub-chapters, paragraphs, etc), create embeddings of each chunk, and store the resultant vectors in a vector database. Your search workflow will then be to create an embedding of your query, and perform a distance comparison (e.g. cosine similarity) which returns ranked results. This way you can now semantically search your texts.
Everything I've mentioned above is fairly easily doable with existing LLM libraries like langchain or llamaindex. For reference, this is an RAG workflow.
And this: https://microsoft.github.io/graphrag/
The Continue extension (Jetbrains, VSCodium) lets you set up assistant and autocompletion independently with different API keys.
Running locally avoids the concern of sending IP or PII to third parties.
Generate documentation in Rust format for use in code.
/// An enum representing possible errors during CRUD operations.
#[derive(Error, Debug)]
enum CrudError {
/// Indicates an error occurred while interacting with the SQLite database.
///
/// This variant wraps a `rusqlite::Error`, providing details about the specific SQLite error.
#[error("SQLite error: {0}")]
SqliteError(rusqlite::Error),
/// Indicates an error related to input/output operations.
///
/// This variant wraps a `std::io::Error`, giving information about the specific IO problem encountered.
#[error("IO error: {0}")]
IoError(std::io::Error),
}1. Local LLM AI models with GUI and command line
2. > Local LLM-based coding tools do exist (such as Google DeepMind’s CodeGemma and one from California-based developers Continue)
In comments sections for repost, I did mentioned that big data is not dead but just having its AI winter moment, and big data is not only about volume (of storage) but also on the memory requirements (of RAM) [3].
Fast forward a few months, it seems the AI LLM is reviving the big data scenario and is seems that big data is far from dead but it's very much healthy in the era of LLMs, cloud or local. Wait until IoT and machine-to-machine (M2M) based systems are in full effects and big data will not just surviving but will be thriving.
[1] Big data is dead (2023) - original
https://news.ycombinator.com/item?id=34694926
[2] Big data is dead (2023) - repost
https://news.ycombinator.com/item?id=40488844
[3] Comments on: Big data is dead (2023):
This is why I’m putting my money on Google in the long run. They have the reach to make it useful and the monetization behemoth to make it profitable.
It's fizzling out.
The current incumbents are sitting on multi-billion dollar valuations and juicy funding rounds. This buys runtime for a good couple of years, but it won't last forever. There's a limit to what can be achieved with scraped datasets and deep Markov chains.
Over time, it will become difficult to judge what makes one general-purpose LLM be any better than another general-purpose LLM. A new release isn't necessarily performing better or producing better quality results, and it may even regress for many use-cases (we're already seeing this with OpenAI's latest releases).
Competitors will have caught up to eachother, and there shouldn't be any major differences between Claude, ChatGPT, Gemini, etc - after-all, they should all produce near-identical answers, given identical scenarios. Pace of innovation flattens out.
Eventually, the technology will become wide-spread, cheap and ubiquitous. Building a (basic, but functional) LLM will be condensed down to a course you take at university (the same way people build basic operating systems and basic compilers in school).
The search for AGI will continue, until the next big hype cycle comes up in 5-10 years, rinse and repeat.
You'll have products geared at lawyers, office workers, creatives, virtual assistants, support departments, etc. We're already there, and it's working great for many use-cases - but it just becomes one more tool in the toolbox, the way Visual Studio, Blender and Photoshop are.
The big money is in the datasets used to build, train and evaluate the LLMs. LLMs today are only as good as the data they were trained on. The competition on good, high-quality, up-to-date and clean data will accelerate. With time, it will become more difficult, expensive (and perhaps illegal) to obtain world-scale data, clean it up, and use it to train and evaluate new models. This is the real goldmine, and the only moat such companies can really have.
I even considered blocking HN.
Two things to remember:
1. OpenAI can analyze which "wrappers" or "apps" are most successful, and make better purchasing decisions that way. This is information which isn't available outside of OpenAI.
2. OpenAI can in theory analyze the actual queries and interactions in an organization, record them, analyze, etc - in an attempt to get a hold of the organization's internal data. Unclear on the legality of this, but could perhaps be enforced through a draconic license.
Gmail
Docs
Android
Chrome (browser and Chromebooks)
I don't use any Meta properties at all, but at least a dozen alphabet ones. My wife uses Facebook, but that's about it. I can see it being handy for insta filters.
YMMV of course, but I suspect alphabet has much deeper reach, even if the actual overall number of people is similar.
I know it sounds crazy but that is what they actually believe and is a regular theme of conversations in SF. They also think it is a flywheel and whoever wins the race in the next few years will be so far ahead in terms of iteration capability/synthetic data that they will be the runaway winner.
PS: for reading web pages I know there's voices integrated in the browser/OS but those are horrible
---
Open WebUI has a voice chat but the voices are not great. I'm sure they'd love a PR that integrates StyleTTS2.
You can give it a Serper API Key and it will search the web to use as context. It connects to ollama running on a linux box with a $300 RTX 3060 with 12GB of VRAM. The 4bit quant of Llama 3.1 8B takes up a bit more than 6GB of VRAM which means it can run embedding models and STT on the card at the same time.
12GB is the minimum I'd recommend for running quantized models. The RTX 4070 Ti Super is 3x the cost but 7 times "faster" on matmuls.
The AMD cards do inference OK but they are a constant source of frustration when trying to do anything else. I bought one and tried for 3 months before selling it. It's not worth the effort.
I don't have any interest in allowing it to run shortcuts. Open WebUI has pipelines for integrating function calling. HomeAssistant has some integrations if that's the kind of thing you are thinking about.
Nexa AI local model hub: https://nexaai.com/ Toolkit: https://github.com/NexaAI/nexa-sdk
It also comes with a built-in local UI to get started with local models easily and OpenAI-compatible API (with JSON schema for function calling and streaming) for starting local development easily.
You can run the Nexa SDK on any device with a Python environment—and GPU acceleration is supported!
Local LLMs, and especially multimodal local models are the future. It is the only way to make AI accessible (cost-efficient) and safe.
2) Privacy
3) Removing safety filters — there are some great “abliterated” models out there that have had their refusal behavior removed. Running these locally and never having your request refused due to corporate risk aversion is a very different experience to calling a safety-neutered API.
Depending on your use case some, all, or none of these will be relevant, but they are undeniable benefits that are very much within reach using a laptop and the current crop of models.
Please correct me if I'm wrong. It would make my life slightly more comfortable
Right now the small models (llama 8B) can't handle this type of task, although they could if they were trained for bi lingual data.
Don’t want to pay for Claude API and Cursor sub gets past the limits quick. Gonna get a M4 Mac with maxed out ram when they come out next month
but truthfully the 8B just aren't that great yet, they can provide some decent info if you're just investigating things but a google search is still faster
Local LLM community has been using Apple Silicon Mac GPUs to do inference.
I’m sure Apple Intelligence uses the NPU and maybe the GPU sometimes.
Man, imagine being OpenAI and flushing your brand down the toilet with an explicit customer noncompete rule which totally backfires and inspires 100x more competition than it prevents
"Llama 3.1 materials or outputs cannot be used to improve or train any other large language models outside of the Llama family."
https://ai.meta.com/llama/license/
Section 1.b.iv
The official llama 3 repo still says this, which is a different phrasing but effectively equal in meaning to what the commenter above said.
every single query to ChatGPT/Claude/Gemini/etc will be used for any purpose, by any party, at any time. shamelessly so, because this is the new normal. Welcome to 2024. I own nothing, have no privacy, and life has never been better.
>(and less useful, even illegal) things
the same illegal things you can do with Notepad, or a pencil and a piece of paper.