Why this exists: Getting historical decoded event data usually means running your own archive node or paying per-RPC-call. We had the infrastructure from another project, so we packaged the output as a flat file.
35 karma · joined July 27, 2025
Why this exists: Getting historical decoded event data usually means running your own archive node or paying per-RPC-call. We had the infrastructure from another project, so we packaged the output as a flat file.
A record is one unit of structured input handled by the system, such as a JSON document, an NDJSON entry, or an API request body. An obligation is a single requirement evaluated against that record, for example a field must be present, a value must satisfy a constraint, or a record must be accepted or rejected based on its contents.
A boundary execution model refers to evaluating those obligations at the system boundary where data enters, rather than deep inside downstream application logic. The intent is to make decisions as part of intake handling, before the data is handed off to other components.
:)
Requests that would normally be fully parsed, tokenized, embedded, and sent to a model are often decided early and dropped… before any of that expensive work happens.
That’s fewer tokens generated, fewer CPU cycles burned, and fewer dollars spent at scale.
The demo focuses on behavior, not throughput tuning. Startup cost scales; runtime does not.
Details available under NDA. I’m reachable at: michaelallenkuykendall [at] gmail [dot] com
Yippee!
# Download a model huggingface-cli download MikeKuykendall/phi-3.5-moe-q4-k-m-cpu-offload-gguf
# Run with MoE offloading ./shimmy serve --cpu-moe --model-path phi-3.5-moe-q4-k-m.gguf Standard OpenAI-compatible API, so existing code works unchanged. Why this matters This democratizes access to state-of-the-art models. Instead of needing a $10,000 GPU or cloud spending, you can run expert models on gaming laptops or modest server hardware. It's not just about making models "work" - it's about sustainable AI deployment where organizations can experiment with cutting-edge architectures without massive infrastructure investments. The technique itself isn't novel (llama.cpp had MoE support), but the Rust bindings, production packaging, and curated model collection make it accessible to developers who just want to run large models locally. Release: https://github.com/Michael-A-Kuykendall/shimmy/releases/tag/... Models: https://huggingface.co/MikeKuykendall Happy to answer questions about the implementation or performance characteristics.
Rust library: Absolutely! Add shimmy = { version = "0.1.0", features = ["llama"] } to Cargo.toml. Use the inference engine directly:
let engine = shimmy::engine::llama::LlamaEngine::new(); let model = engine.load(&spec).await?; let response = model.generate("prompt", opts, None).await?;
No need to spawn processes - just import and use the components directly in your Rust code.
1. Install Shimmy:
cargo install shimmy
2. Get GGUF models (same models you'd use with Ollama):
# Download to ./models/ directory
huggingface-cli download microsoft/Phi-3-mini-4k-instruct-gguf --local-dir
./models/
# Or use existing Ollama models from ~/.ollama/models/
3. Start serving:
./shimmy serve
4. Use with any OpenAI-compatible client at http://localhost:11435 Key differences:
- Architecture: llama-swap = proxy + multiple servers, Shimmy = single server
- Resource usage: llama-swap runs multiple processes, Shimmy = one 50MB process
- Use case: llama-swap for managing many models, Shimmy for simplicityQuick demo - working VSCode + local AI in 30 seconds: curl -L https://github.com/Michael-A-Kuykendall/shimmy/releases/late... ./shimmy serve # Point VSCode/Cursor to localhost:11435
The technical achievement: Got it down to 5.1MB by stripping everything except pure inference. Written in Rust, uses llama.cpp's engine.
One feature I'm excited about: You can use LoRA adapters directly without converting them. Just point to your .gguf base model and .gguf LoRA - it handles the merge at runtime. Makes iterating on fine-tuned models much faster since there's no conversion step.
Your data never leaves your machine. No telemetry. No accounts. Just a tiny binary that makes GGUF models work with your AI coding tools.
Would love feedback on the auto-discovery feature - it finds your models automatically so you don't need any configuration.
What's your local LLM setup? Are you using LoRA adapters for anything specific?
ContextLite uses SMT (Satisfiability Modulo Theories) solvers + BM25 heuristics to mathematically prove optimal context matches instead of guessing with similarity scores. We process 2,406 files/second with formal verification - understanding imports, dependencies, and code relationships that vector embeddings completely miss.
No GPU required, no embedding models, no vector databases. Just blazing fast context that actually reasons about your codebase structure using constraint satisfaction and theorem proving.
We're live in production with npm, PyPI, VS Code marketplace, and 8 other package managers. 14-day SMT trial with full formal reasoning, then $99 lifetime license. Enterprise teams get advanced analytics, multi-repo support, and custom deployment options.
The future of AI context isn't more computation – it's mathematical precision. SMT solvers can prove correctness; vector databases can only guess similarity.
Try the math: contextlite.com/downloads