HNHacker News
TopNewBestAskShowJobs

MKuykendall

35 karma · joined July 27, 2025

coder
submissionscomments
MKuykendall··on Show HN: We decoded 286M Ethereum events and packaged them as a dataset
We decoded 286 million raw Ethereum events from Sepolia's early blocks (genesis through ~3M) into clean, query-ready Parquet files. The focus is on ERC-20/ERC-721 transfers and approvals — the kind of data you'd need for ML training, compliance analysis, or building indexers.

Why this exists: Getting historical decoded event data usually means running your own archive node or paying per-RPC-call. We had the infrastructure from another project, so we packaged the output as a flat file.

MKuykendall··on Show HN: Wanderlust-Rust daemon nukes Windows PATH bloat every 30 min
Nothing was doing it that often, but nothing WILL ever taint my PATH again.
MKuykendall··on Show HN: Shimmytok – Pure Rust GGUF tokenizer (no C++, no extra files)
Thanks! Rust needed it
MKuykendall··on Runtime invariant to rule count in a single-pass boundary execution model
I did not expect a reply sorry!

A record is one unit of structured input handled by the system, such as a JSON document, an NDJSON entry, or an API request body. An obligation is a single requirement evaluated against that record, for example a field must be present, a value must satisfy a constraint, or a record must be accepted or rejected based on its contents.

A boundary execution model refers to evaluating those obligations at the system boundary where data enters, rather than deep inside downstream application logic. The intent is to make decisions as part of intake handling, before the data is handed off to other components.

:)

MKuykendall··on Runtime invariant to rule count in a single-pass boundary execution model
This probably went right over everyone’s head. What it actually means is cheaper inference compute and faster, cheaper processing of JSON (or any structured data).

Requests that would normally be fully parsed, tokenized, embedded, and sent to a model are often decided early and dropped… before any of that expensive work happens.

That’s fewer tokens generated, fewer CPU cycles burned, and fewer dollars spent at scale.

MKuykendall··on Show HN: Auxide- a Real-Time Audio Graph Library for Rust
MIDI!

https://github.com/Michael-A-Kuykendall/auxide-midi

MKuykendall··on Show HN: Auxide- a Real-Time Audio Graph Library for Rust
DSP!

https://github.com/Michael-A-Kuykendall/auxide-dsp

MKuykendall··on Show HN: Auxide- a Real-Time Audio Graph Library for Rust
IO!

https://github.com/Michael-A-Kuykendall/auxide-io

MKuykendall··on Runtime invariant to rule count in a single-pass boundary execution model
https://github.com/Michael-A-Kuykendall/dzero
MKuykendall··on Runtime invariant to rule count in a single-pass boundary execution model
This is a small experiment showing per-record runtime independent of the number of compiled obligations in a single-pass boundary execution model.

The demo focuses on behavior, not throughput tuning. Startup cost scales; runtime does not.

Details available under NDA. I’m reachable at: michaelallenkuykendall [at] gmail [dot] com

MKuykendall··on Show HN: Muxide – Zero-dep pure Rust MP4 muxer (H.264/H.265/AV1, no FFmpeg)
I built this for my Crabcamera Tauri camera plugin; I needed MP4 recording without shipping FFmpeg binaries. Takes encoded frames + timestamps, outputs a playable MP4.

Yippee!

MKuykendall··on Shimmy v1.7.0: Running 42B Moe Models on Consumer GPUs with 99.9% VRAM Reduction
I just released Shimmy v1.7.0 with MoE (Mixture of Experts) CPU offloading support, and the results are pretty exciting for anyone who's hit GPU memory walls. What this solves If you've tried running large language models locally, you know the pain: a 42B parameter model typically needs 80GB+ of VRAM, putting it out of reach for most developers. Even "smaller" 20B models often require 40GB+. The breakthrough MoE CPU offloading intelligently moves expert layers to CPU while keeping active computation on GPU. In practice: Phi-3.5-MoE 42B: Runs on 8GB consumer GPUs (was impossible before) GPT-OSS 20B: 71.5% VRAM reduction (15GB → 4.3GB, measured) DeepSeek-MoE 16B: Down to 800MB VRAM with Q2 quantization The tradeoff is 2-7x slower inference, but you can actually run these models instead of not running them at all. Technical implementation Built on enhanced llama.cpp bindings with new with_cpu_moe() and with_n_cpu_moe(n) methods Two CLI flags: --cpu-moe (automatic) and --n-cpu-moe N (manual control) Cross-platform: Windows MSVC CUDA, macOS Metal, Linux x86_64/ARM64 Still sub-5MB binary with zero Python dependencies Ready-to-use models I've uploaded 9 quantized models to HuggingFace specifically optimized for this: Phi-3.5-MoE variants (Q8.0, Q4 K-M, Q2 K) DeepSeek-MoE variants GPT-OSS 20B baseline Getting started # Install cargo install shimmy

# Download a model huggingface-cli download MikeKuykendall/phi-3.5-moe-q4-k-m-cpu-offload-gguf

# Run with MoE offloading ./shimmy serve --cpu-moe --model-path phi-3.5-moe-q4-k-m.gguf Standard OpenAI-compatible API, so existing code works unchanged. Why this matters This democratizes access to state-of-the-art models. Instead of needing a $10,000 GPU or cloud spending, you can run expert models on gaming laptops or modest server hardware. It's not just about making models "work" - it's about sustainable AI deployment where organizations can experiment with cutting-edge architectures without massive infrastructure investments. The technique itself isn't novel (llama.cpp had MoE support), but the Rust bindings, production packaging, and curated model collection make it accessible to developers who just want to run large models locally. Release: https://github.com/Michael-A-Kuykendall/shimmy/releases/tag/... Models: https://huggingface.co/MikeKuykendall Happy to answer questions about the implementation or performance characteristics.

MKuykendall··on Show HN: Rustchain – Rust toolchain AI agent framework universal transpilation
Side note I built a 2500 star and climbing rust application from stem to stern using it

https://github.com/Michael-A-Kuykendall/shimmy

MKuykendall··on Show HN: CrabCamera – Cross-platform camera plugin for Tauri desktop apps
Great alternatives list! Each serves different use cases: QtMultimedia: Excellent for C++/Qt developers, but requires Qt framework react-native-vision-camera: Perfect for mobile, but CrabCamera targets desktop OpenCV: Great for computer vision, but heavy for simple camera access CrabCamera's niche: Rust developers building Tauri desktop apps who want: Zero Qt dependencies Native Rust integration Minimal bundle size Cross-platform camera control Different tools for different ecosystems! Currently powering our Budsy plant identification app.
MKuykendall··on Show HN: CrabCamera – Cross-platform camera plugin for Tauri desktop apps
Good point about explaining Tauri better! For context: Tauri = Rust + Web frontend (like Electron but smaller/faster) Problem: Desktop apps need camera access, but web APIs are limited CrabCamera: Provides native camera control for Tauri desktop apps Real example: Our Budsy plant identification app uses CrabCamera to capture photos for botanical analysis - something web camera APIs can't do effectively. Thanks for the feedback on clarity!
MKuykendall··on Show HN: CrabCamera – Cross-platform camera plugin for Tauri desktop apps
ou were absolutely right about WebRTC complexity! Since that feedback, we've refocused CrabCamera on its core mission - desktop camera access for Tauri apps. Changes made based on your feedback: Removed WebRTC from core scope Focused on clean camera capture API Left streaming protocols to dedicated libraries Current CrabCamera v0.3.0: 45/45 tests passing Production-ready in Budsy plant identification app Clean separation of concerns Thanks for steering us toward better architecture!
MKuykendall··on Show HN: CrabCamera – Cross-platform camera plugin for Tauri desktop apps
This is a crate I am using for another application, I thought it was neato
MKuykendall··on Show HN: CrabCamera – Cross-platform camera plugin for Tauri desktop apps
Thx good call!
MKuykendall··on Show HN: 5MB Rust binary that runs HuggingFace models (no Python)
GPU/CUDA: Yes, but disabled by default for faster builds. To enable: remove LLAMA_CUDA = "OFF" from config.toml and rebuild with CUDA toolkit installed.

Rust library: Absolutely! Add shimmy = { version = "0.1.0", features = ["llama"] } to Cargo.toml. Use the inference engine directly:

let engine = shimmy::engine::llama::LlamaEngine::new(); let model = engine.load(&spec).await?; let response = model.generate("prompt", opts, None).await?;

No need to spawn processes - just import and use the components directly in your Rust code.

MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
I didn't have that path set to autodiscover; pull the newest version this is fixed now!!
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
To use Shimmy (instead of Ollama):

  1. Install Shimmy:
  cargo install shimmy
  2. Get GGUF models (same models you'd use with Ollama):
  # Download to ./models/ directory
  huggingface-cli download microsoft/Phi-3-mini-4k-instruct-gguf --local-dir
   ./models/
  # Or use existing Ollama models from ~/.ollama/models/
  3. Start serving:
  ./shimmy serve
  4. Use with any OpenAI-compatible client at http://localhost:11435
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
Try cargo install or intentionally exclude, unsigned Rust binaries will do this.
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
This should be fixed now!
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
Shimmy is for when you want the absolute minimum footprint - CI/CD pipelines, quick local testing, or systems where you can't install 680MB of dependencies.
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
Shimmy is designed to be "invisible infrastructure" - the simplest possible way to get local inference working with your existing AI tools. llama-server gives you more control, llama-swap gives you multi-model management.

  Key differences:
  - Architecture: llama-swap = proxy + multiple servers, Shimmy = single server
  - Resource usage: llama-swap runs multiple processes, Shimmy = one 50MB process
  - Use case: llama-swap for managing many models, Shimmy for simplicity
MKuykendall··on Show HN: Shimmy – 5MB privacy-first, local alternative to Ollama (680MB)
Hey HN! I built this because I was tired of waiting 10 seconds for Ollama's 680MB binary to start just to run a 4GB model locally.

Quick demo - working VSCode + local AI in 30 seconds: curl -L https://github.com/Michael-A-Kuykendall/shimmy/releases/late... ./shimmy serve # Point VSCode/Cursor to localhost:11435

The technical achievement: Got it down to 5.1MB by stripping everything except pure inference. Written in Rust, uses llama.cpp's engine.

One feature I'm excited about: You can use LoRA adapters directly without converting them. Just point to your .gguf base model and .gguf LoRA - it handles the merge at runtime. Makes iterating on fine-tuned models much faster since there's no conversion step.

Your data never leaves your machine. No telemetry. No accounts. Just a tiny binary that makes GGUF models work with your AI coding tools.

Would love feedback on the auto-discovery feature - it finds your models automatically so you don't need any configuration.

What's your local LLM setup? Are you using LoRA adapters for anything specific?

MKuykendall··on RAG is solving the wrong problem
After watching developers struggle with 200ms+ vector database queries for context retrieval, we realized RAG was fundamentally backwards. Why compute expensive embeddings when you can find context in 0.3ms with SMT-powered reasoning?

ContextLite uses SMT (Satisfiability Modulo Theories) solvers + BM25 heuristics to mathematically prove optimal context matches instead of guessing with similarity scores. We process 2,406 files/second with formal verification - understanding imports, dependencies, and code relationships that vector embeddings completely miss.

No GPU required, no embedding models, no vector databases. Just blazing fast context that actually reasons about your codebase structure using constraint satisfaction and theorem proving.

We're live in production with npm, PyPI, VS Code marketplace, and 8 other package managers. 14-day SMT trial with full formal reasoning, then $99 lifetime license. Enterprise teams get advanced analytics, multi-repo support, and custom deployment options.

The future of AI context isn't more computation – it's mathematical precision. SMT solvers can prove correctness; vector databases can only guess similarity.

Try the math: contextlite.com/downloads