4,108 karma · joined December 12, 2007
I do like the sound of Verscheissification.
As Picard said: Enshitify! (Apologies to my Trekkie friends.)
You can chat with any model in the built-in WebUI, connect other apps (coding agents, chat UIs, editors), or use the API directly. Models load when requested and unload when idle, so they don't take up memory when not in use.
> Features
- 100% local — Models run on your Mac; no data ever leaves it
- Small footprint — 4 MB native macOS app
- Zero configuration — models are auto-configured with optimal settings for your Mac
- Model recommendations — a built-in list of models your Mac can run, installable in one click
- Standard storage — models live in the Hugging Face cache, shared with llama.cpp and other tools
- Built on llama.cpp — from the GGML org, developed alongside llama.cpp
> Installation
To install llama.cpp , run:
brew install llama.cpp
To install llama-app, run: brew install --cask llama-appOfficial repo, also has documentation how to configure server parameters:
https://github.com/ggml-org/Llama-macOS
Small tip, install llama.cpp with brew before llama.app, which will pick up the existing llama.cpp. That way it's easier to stay up to date with llama.cpp, since llama.app is on a slower release cadence.
Also, models installed with the hugging face CLI (hf) are picked up by llama.app automatically. The CLI will keep the model cache updated, e.g. when models get updated.
Llama.cpp became part of Huggingface recently.
https://luxurylaunches.com/transport/gabe-newell-explorer-ve...
https://www.forbes.com/sites/deajusufi/2026/06/13/gabe-newel...
High-Performance AI on a Budget: Optimizing llama.cpp for Qwen3.5 Inference on a Dual-GPU HP Z440
But this is of huge interest to carriers, since it allows them to skip the PSTN/peering cost when the callee endpoint is an IP phone.
There is private ENUM for carrier use I recall, not sure what the current status is, with LTE/VoLTE, RCS etc.pp.
http://dam3d3.free.fr/PFE/Pathfinder/GSMA_PathFinder_WebSite...
Here the list of countries that have ENUM delegated for their country code.
https://www.itu.int/en/ITU-T/inr/enum/Pages/delegations.aspx
Decentralized and under user control, no shitty silos like FaceTime, WhatsApp.
ENUM stands for “Telephone Number Mapping.” It is essentially a bridge between the world of telecommunications and the Internet. With a single ENUM domain, you can combine all your contact options under your familiar phone number:
It would have provided geographical information based on a domain encoded grid, not for human but machine consumption (e.g. acme.2e5n.10e30n.geo).
https://en.wikipedia.org/wiki/.geo
In a similar vein there is the 'e164.arpa' domain for mapping telephone numbers.
https://huggingface.co/collections/unsloth/gemma-4
Edit: Sorry, I'm not sure if this is a quant, but it says 'finetuned' from the Google Gemma 4 parent snapshot. It's the same size as the UD 8-bit quant though.
https://unsloth.ai/docs/basics/unsloth-dynamic-2.0-ggufs
For the best quality reply, I used the Gemma-4 31B UD-Q8_K_XL quant with Unsloth Studio to summarize the URL with web search. It produced 4.9 tok/s (including web search) on an MacBook Pro M1 Max with 64GB.
Here an excerpt of it's own words:
Unsloth Dynamic 2.0 Quantization
Dynamic 2.0 is not just a "bit-reduction" but an intelligent, per-layer optimization strategy.
- Selective Layer Quantization: Instead of making every layer 4-bit, Dynamic 2.0 analyzes every single layer and selectively adjusts the quantization type. Some critical layers may be kept at higher precision, while less critical layers are compressed more.
- Model-Specific Tailoring: The quantization scheme is custom-built for each model. For example, the layers selected for quantization in Gemma 3 are completely different from those in Llama 4.
- High-Quality Calibration: They use a hand-curated calibration dataset of >1.5M tokens specifically designed to enhance conversational chat performance, rather than just optimizing for Wikipedia-style text.
- Architecture Agnostic: While previous versions were mostly effective for MoE (Mixture of Experts) models, Dynamic 2.0 works for all architectures (both MoE and non-MoE).
And then this blew my mind:
Quite the underground scene:
As I recall, there were tons of books about GEM for the Atari ST, at least in Europe.
Thankfully Atari licensed GEM for their 68000 machines before the lawsuit, and wasn't affected by these changes. The Atari ST (Sixteen/Thirtytwo) was very Mac like at the time. It even ran the Mac OS from Apple ROMs (Spectre 128 and Aladin) on its much cheaper hardware.
SHATTER:
https://imgur.com/gallery/shatter-1984-was-first-commerciall...
Robot Empire:
https://www.reddit.com/r/atarist/comments/xgs4rh/comicbook_c...
Works fine on MacOS now (chat only).
On Ubuntu 24.04 with two GPU's (3090+3070), it appears that Llama.cpp sometimes uses the CPU and not GPU. This is judging from the tk/s and CPU load for identical models run with US-studio vs. just Llama.cpp (bleeding edge).
(base) unsloth git:(main) unsloth studio setup
╔══════════════════════════════════════╗
║ Unsloth Studio Setup Script ║
╚══════════════════════════════════════╝
Node v25.8.1 and npm 11.11.0 already meet requirements. Skipping nvm install.
Node v25.8.1 | npm 11.11.0
npm run build failed (exit code 2):
> unsloth-theme@0.0.0 build
> tsc -b && vite build
src/features/chat/shared-composer.tsx(366,17): error TS6133: 'status' is declared but its value is never read.