HNHacker News
TopNewBestAskShowJobs

raullen

22 karma · joined September 21, 2011

I am an Co-founder, ex-Uber-er, ex-Googler, Doctor, Cryptographer, Researcher, Runner, Soccer player (MF), Rock-climber and Musician.
submissionscomments
raullen··on Show HN: Rapid-MLX – Run local LLMs on Mac, 2-3x faster than alternatives
Built this to run coding agents locally on Apple Silicon. The main problem I kept hitting: most models fail at structured tool calling, and existing servers are slow on MLX.

Two findings from benchmarking 7 models across 5 agent frameworks:

1. Qwen family gets 100% tool calling across every framework tested. Non-Qwen models (Llama, DeepSeek-R1) vary wildly — 40% to 100% depending on framework.

2. smolagents (HuggingFace) sidesteps structured function calling entirely by using code generation. DeepSeek-R1 goes from 40% with structured FC to 100% with smolagents.

Speed-wise, MLX's unified memory means zero CPU↔GPU copies. On an M3 Ultra: Qwen3.5-9B hits 108 tok/s (vs ~41 on Ollama), Qwen 3.6 35B does 100 tok/s with only 3B active params.

The full benchmark data is in the README. Happy to discuss the MLX performance characteristics or tool calling architecture.

raullen··on vLLM-mlx – 65 tok/s LLM inference on Mac with tool calling and prompt caching
I've been working on a fork of vllm-mlx (OpenAI-compatible LLM server for Apple Silicon) to make it actually usable for coding agents. The upstream project is great but was missing production-grade tool calling, reasoning separation, and multi-turn performance.

  What I added (37 commits):

  - Tool calling that works — streaming + non-streaming, supports MiniMax and Hermes/Qwen3 formats. 4/4 accuracy on structured function calling benchmarks.
  - Reasoning separation — MiniMax-M2.5 mixes reasoning into its output with no tags. Built a heuristic parser that cleanly separates reasoning from content (0% leak rate, was 60%
   with the generic parser).
  - Prompt cache for SimpleEngine — persistent KV cache across requests. On 33K-token coding agent contexts: TTFT goes from 28s to 0.3s on cache hit. This is the single biggest
  improvement for multi-turn use.
  - 1500+ tests — parsers, engine, server, tool calling. The upstream had minimal test coverage.

  Benchmarks (Mac Studio M3 Ultra, 256GB):

  Qwen3-Coder-Next-6bit (80B MoE, 3B active):
  - Decode: 65 tok/s
  - Prefill: 1090-1440 tok/s
  - TTFT (cache hit, 33K context): 0.3s

  MiniMax-M2.5-4bit (229B MoE):
  - Decode: 33-38 tok/s
  - Deep reasoning with tool calling

  I built this to run OpenClaw locally on my Mac instead of paying for cloud APIs. Qwen3-Coder-Next at 65 tok/s with tool calling is genuinely usable — not a toy demo.

  Quick start:

  pip install git+https://github.com/raullenchai/vllm-mlx.git
  python -m vllm_mlx.server \
    --model lmstudio-community/Qwen3-Coder-Next-MLX-6bit \
    --tool-call-parser hermes --port 8000

  GitHub: https://github.com/raullenchai/vllm-mlx
raullen··on Vnsh – An ephemeral, host-blind file sharing tool for AI context
OP here.

I built this because I got tired of Claude choking when I pasted 5,000 lines of server logs, or worrying about leaving sensitive environment variables in my chat history forever.

What is it? vnsh (vanish) is a CLI tool and web app that encrypts data client-side and uploads it to a host-blind storage (Cloudflare R2). It generates a link where the decryption key is in the URL hash fragment.

Architecture:

Encryption: AES-256-CBC via WebCrypto API.

Transport: The key never leaves your device (browser or CLI). The server only sees an encrypted blob.

Storage: Cloudflare Workers + R2.

Integration: It has a native MCP (Model Context Protocol) server. If you use Claude Code or desktop, the agent can "read" these links directly without you pasting the text.

The Stack: Typescript, Hono, Cloudflare Workers, React (Web), Node (CLI).

It's fully open source. I'm looking for feedback on the crypto implementation and the MCP integration flow.

Repo: https://github.com/raullenchai/vnsh Web: https://vnsh.dev

raullen··on Claw (Claude AnyWhere)
Running a long Claude Code session? Need to step away from your desk? Claw lets you monitor and control Claude Code from any device with a browser.

See what Claude is doing in real-time from any screen Send quick responses (yes/no/continue) with one tap Interrupt with Ctrl+C when things go sideways Monitor everything — sessions, windows, panes, git status, system stats

raullen··on Every day at the same time, my internet dies for 1 minute. How do I investigate?
Mine dies for 10min
raullen··on Coinbase Card
Why cointracker starts with Coinbase/Google account, rather than an BTC/ETH addr?
raullen··on I'm Peter Roberts, immigration attorney who does work for YC and startups. AMA
This thread demonstrates how little people care about their privacy...
raullen··on Ask HN: Who is hiring? (January 2019)
IoTeX Network | Palo Alto, California | Full Stack/Frontend Engineers | Full-time/Part-time/Intern | https://iotex.io IoTeX is building the auto-scalable and privacy-centric blockchain infrastructure designed and optimized for the Internet of Things (IoT). Full Stack/Frontend engineers are needed to speed up our product development process. Apply here: https://iotex.io/careers
raullen··on [dead]
IoTeX Network | Palo Alto, California | Full Stack/Frontend Engineers | Full-time/Part-time/Intern | https://iotex.io IoTeX is building the auto-scalable and privacy-centric blockchain infrastructure designed and optimized for the Internet of Things (IoT). Full Stack/Frontend engineers are needed to speed up our product development process.

Apply here: https://iotex.io/careers

raullen··on “hash_salt=” : Salts public on Github
https://github.com/search?q=%22hash_salt%3D%22&ref=searchres... is a better link.
raullen··on HTTP/2 is here. Goodbye SPDY? Not quite yet
Google's HTTP loadbanlancer and CDN have supported H2 for a long while.
raullen··on Show HN: Scan and share apps on your iPhone's home screen
Good idea but need to be polished. Many false match, see http://appetite.io/a/daf97fec

Besides the features mentioned by nmcfarl, it would be great to enable a direct upload from iPhone after taking the snapshot, e.g., send via email.

raullen··on Ask HN: Any volunteering opportunities in YC-backed startups?
Thanks for the heads-up, man. I am checking on Quora now~~~
raullen··on [dead]
Down!
raullen··on [dead]
really depends on what OS you prefer... For MacOS, Sublime Text is the best one per my understanding.