We are both late and early.
730 karma · joined July 31, 2014
We are both late and early.
For our code assistant use cases the local inference on Macs will tend to favor workflows where there is a lot of generation and little reading and this is the opposite of how many of use use Claude Code.
Source: I started getting Mac Studios with max ram as soon as the first llama model was released.
nix run --impure 'git+https://codeberg.org/ingenieroariel/arcan?ref=nix-flake-buil...'
An article called "A Spreadsheet and a Debugger walk into a Shell" [0] by Bjorn (letoram) is a good showcase of an alternative to cells in a Jupyter notebook (Excel like cells!). Another alternative a bit more similar to Jupyter that also runs on Arcan is Pipeworld.
[0] https://arcan-fe.com/2024/09/16/a-spreadsheet-and-a-debugger... [1] https://arcan-fe.com/2021/04/12/introducing-pipeworld/
PS: I hang out at Arcan's Discord Server, you are welcome to join https://discord.com/invite/sdNzrgXMn7
Chris Randall is pretty awesome too: https://www.instagram.com/chris.randall.art/
Since it is a complete toolkit, you can have detachable applications where you send both code and state to a server and retrieve it from another device (like Apple's continuity).
In the end it is just a bunch of lua scripts talking to other components via /dev/shm and to other computers using a new protocol called a12://
nix run nixpkgs#pyspread [0/1 built, 3/113/132 copied (1311.8/1721.6 MiB), 280.4/300.7 MiB DL] fetching llvm-16.0.6 from https://cache.nixos.org
There is a library by the same author called lonboard that provides the JS bits inside JupyterLab. https://github.com/developmentseed/lonboard
<speculation>I think it is based on the Kepler.gl / Deck.gl data loaders that go straight to GPU from network.</speculation>
I kept getting emails about small changes and the bills got bigger all over the place including BigQuery and how they dealt with queries on public datasets. Bill got higher.
There is a non zero chance I conflated things. But from my point of view: I created a system and let it running for years - afterwards bills got higher out of the blue and I moved out.
{ pkgs, ... }: {
services.postgres = {
enable = true;
package = pkgs.postgresql_15;
initialDatabases = [{ name = "mydb"; }];
extensions = extensions: [
extensions.postgis
extensions.timescaledb
];
settings.shared_preload_libraries = "timescaledb";
initialScript = "CREATE EXTENSION IF NOT EXISTS timescaledb;";
};
}A project funded by the EU to bring Nix to Windows.
(edit: typo and clarity)
System is a Mac Studio with 128GB + Asahi Linux + mmapped parquet files and DuckDB, it also runs airflow for us and with Nix can be used to accelerate developer builds and run the airflow tasks for the data team.
GCP is nice when it is free/cheap but they keep tabs on what you are doing and may surprise you at any point in time with ever higher bills without higher usage.
In my use case, the niche is not having WSL/Docker available and letting end users repeat studies or re-run configuration scripts.
Adding it here since a few people were wondering about it in the comments, but feel free to check the original article for the 2024 update:
For other architectures like Power/armv7/i686 the software can run using the Blink project [0] [0] https://github.com/jart/blink
As Cosmopolitan Libc has evolved, it has been possible to compile more software without modifications, and that includes latest Python through a project called superconfigure[1].
Last person who tried to reproduce it from scratch did it last week (granted it too them a few days of solid work) but in the end they ended with a portable binary with Python 3.11.9, brotli, ssl and asyncio for their work related project.[2]
[0] https://github.com/jart/cosmopolitan/tree/master/third_party... [1] https://github.com/ahgamut/superconfigure/ [2] https://github.com/croqaz/cpython/
- Worst case: as good as 3.5 - Common case: way better than 3.5 - Best case: as good as 4.0
It is as fast as 34B model, but uses as much memory as a 132B model. A mixture of 16 experts, activates 4 at a time, so has more chances to get the combo just right than Mixtral (8 with 2 active).
For my personal use case (a top of the line Mac Studio) it looks like the perfect size to replace GPT-4 turbo for programming tasks. What we should look out for is people using them for real world programming tasks (instead of benchmarks) and reporting back.
./mixtral-8x7b-instruct-v0.1.Q8_0.llamafile --cli -t 16 -n 200 -p "In terms of Lasso"
I got 15 tokens per second for prompt evaluation and 8 tokens per second for regular eval.
The same hardware can run things much faster on OSX, or if you use more quantization but I prefer to run things at Q8 or f16 even if they are slow. In the future I how to use GPU, ANE and the crazy 1.58 or 0.68 bit quantization but for now this does the trick handsomely.
On a Mac Studio with NixOS based Asahi Linux and 128Gb of RAM, mixtral 8x7b uses 49GB of RAM. At the same time I load airflow tasks that deal with world wide datasets (using ~60GB on 16 parallel streams with the performance cores) format is parquet and also mmaped.
Computer still has 8 efficiency cores and the whole GPU for visualizing the maps using lonboard / browsing / etc.
The computer uses 8-10W when idle, ~100W when running jobs or actively using the LLM and around ~200W when really using the GPU.
This makes it very efficient energy wise in my book compared to the beast of keeping a modern CPU and nvidia GPU on when idle. My electricity bill is unaffected.
Small, pure, functional, content-addressable and network-first sounds a lot like a mini Nix+ca-derivations [1]
Congrats, this looks very useful and awesome.
dependencies = [
# cli
"click>=8.0,<9",
# python 3.8 compatibility
"importlib_resources>=5.10.2; python_version < \"3.9\"",
# code completion
"jedi>=0.18.0",
# compile markdown to html
"markdown>=3.4,<4",
# add features to markdown
"pymdown-extensions>=9.0,<11",
# syntax highlighting of code in markdown
"pygments>=2.13,<3",
# for reading, writing configs
"tomlkit>= 0.12.0",
# web server
"tornado>=6.1,<7",
# python <=3.9 compatibility
"typing_extensions>=4.4.0; python_version < \"3.10\"",
# for cell formatting; if user version is not compatible, no-op
"black",
]That said, it does not seem yet ready to be a daily driver, small rough edges in the terminal behavior or the editor when using Vim mode were too distracting to do a long programming session.
Kudos to the authors, looking forward to giving it a spin again the future.