HNHacker News
TopNewBestAskShowJobs

tatef

99 karma · joined July 12, 2020

Software engineer & tech founder
submissionscomments
tatef··on [dead]
Hey HN, along with many of you I’ve been burning through my Claude Code usage and have been wondering how I can save on costs and still make the most of it.

I’m also a big local model enthusiast, but the truth is I don’t believe people should have to binary switch between a local setup and OpenAI/Anthropic.

So instead of making it about hosted vs local, I decided to build a tool which allows Claude Code or Codex to delegate work to subagents running locally on your machine (or network) OR via other provider (OpenRouter).

This is different from the typical subagents you can spin up which will contribute to your usage quota/token costs. It also means you can put your GPU to use for simpler tasks instead of letting Claude burn usage on less important work.

Check it out here: https://github.com/labscommunity/yeschef

Open to feedback :)

tatef··on Sharding a 70B model across 39 Intel laptops
Hey everyone, we’ve been working on a project called https://cascadia.to which allows you to shard LLMs across Intel-based machines and perform inference across CPUs/GPUs/NPUs.

To get it functioning, we split models into shards and pre-compiled them to OpenVINO IR graphs.

Since then, we’ve been working on a lot of optimizations to make sharding and running models on Intel much more performant. Most of our initial improvements are thanks to the use of speculative decoding (with some interesting workarounds there) and also micro-batching and continuous batching to improve the total throughput.

We’re still in the early days, but from our work so far:

- We ran an 8B parameter model on two Intel PCs, serving two users concurrently via iGPUs at ~43 tok/s aggregate - After adding an additional PC and user, it reached 64.67 tok/s - We’ve successfully sharded larger models (e.g. with 70B parameters), and with Cascadia inference was 3.1x faster than basic sharding - And for fun, we got 39 Intel AI PCs, hooked them up via ethernet, and sharded a 70B model across them. Benchmarks aren’t super impressive yet (~1 tok/s on CPU), but we’ve been making a ton of progress there.

If you’re wondering why we’re doing this, there’s a few reasons. There is a lot of Intel hardware out there, and dedicated AI hardware right now is expensive.

Organizations that own Intel computers will be able to pool their computing power to run models locally. For hobbyists with Intel GPU rigs or Intel Xeon Server PCs, you likely want inference to be optimized for your hardware. Cascadia is aiming to be a runtime for anyone wanting to run AI on Intel hardware and squeeze the most juice out of their machines.

Cascadia is open source (Apache 2.0) and available now.

This is alpha, so we’re open to any feedback on the architecture.

Give it a go and let us know what you think :)

tatef··on Using LLMs to make novel research discoveries
Hi HN!

Coming to you today with Autolab, a framework I built that combines Ralph loops and Karpathy's autoresearch to design and run tests and make novel research discoveries. I hope one day a similar framework can be used with hardware, but for now, I've been using this tool in my workflows and find it incredibly useful. Wanted to share.

tatef··on Add 500M tokens of context space to any LLM with <300ms latency
Hi HN,

Wanted to share a project I've been working on aimed at helping solve the context problem called Memoryport. It works across LLM providers/apps, stores conversations locally, is fully OSS, and enables anyone to keep track of their memories over time. I tested mine with 500M tokens of context space (note: not the same as a context window) and only added ~300ms of latency to the inference session. I also built out an open spec called AMP that standardizes the communication protocol for how memory systems should interact with LLMs (see repo).

More than happy to answer any questions and hope you all find this as useful as I have.

tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Thanks for sharing this! If you'd be interested in running the benchmark yourself with Hypura I'd happily merge into our stats. Otherwise will add to my todo list :)
tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Yes definitely. I use a M1 Max with 32gb of RAM daily and it's about on par from a performance standpoint with the new base M5 Pro 24gb. You can check the benchmarks in the repo if you're interested in seeing specific performance metrics, but investing in Apple hardware with as much memory as possible will generally get you furthest in this game.
tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Noted, thanks. I had LLM help positioning this message but I did the initial draft along with edits. Will keep in mind for the future.
tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Hypura reads tensor weights from the GGUF file on NVMe into RAM/GPU memory pools, then compute happens entirely in RAM/GPU.

There is no writing to SSDs on inference with this architecture.

tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Yes, exactly this.
tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
I'm referencing it as being possible, however I didn't share benchmarks because candidly the performance would be so slow it would only be useful for very specific tasks over long time horizons. The more practical use cases are less flashy but capable of achieving multiple tokens/sec (ie smaller MoE models where not all experts need to be loaded in memory simultaneously)
tatef··on Hypura – A storage-tier-aware LLM inference scheduler for Apple Silicon
Yes, definitely agree. It's more of a POC than a functional use case. However, for many smaller MoE models this method can actually be useful and capable of achieving multiple tokens/sec.
tatef··on Nest.land – An immutable, blockchain powered module registry for Deno
We aren't a package manager. We're a registry and CDN (of sorts). Blockchain is actually a huge solution to this problem for three very notable reasons. The first is that Deno module imports are url based, and we don't want code going off the internet, as this would break the code dependent on it. Blockchain solves this because transactions (module code) are unable to be modified or deleted. This means that import links will never break, thanks to blockchain! In addition, it's unbelievably cheap to permanently store data. For reference, we've stored 17,297 files on the blockchain. For proof, you can see our wallet address and transaction history here: https://viewblock.io/arweave/address/tySYSW93nDky1sbCO56PmyE... This permanent and decentralized data storage has cost us right around 5 cents USD. Thirdly, thanks to the blockchain, the module data is completely decentralized across over 340 nodes and counting around the world. You can see the exact statistic here: https://viewblock.io/arweave Again, thanks for bringing these things up. These are great points for us to address publicly.
tatef··on Nest.land – An immutable, blockchain powered module registry for Deno
Actually, this raises a very good point. I'm Tate, a co-founder. Our publishing system works in a way that users will be able to publish malicious modules, yes, but our registry is not decentralized up to a certain point; let me elaborate on this. If a user finds that a module is malicious and wants to report it, we can remove it from the registry completely because the registry is centralized. Though this data will still be accessible from the blockchain and the import url will be functional, we're building a system to warn the user whenever the url is imported from a Deno-specific response header. Now, after a certain amount of time has passed and a module isn't reported as malicious, we're building a system to automatically publish the entire registry to the blockchain as well, so that the registry AND the module are immutable. This is called Fossil, our "archiver." You can see its code here: https://github.com/nestdotland/fossil Again, thanks for bringing this up. I hope this explanation helped. Our goal certainly is not to promote or enable malicious code!
tatef··on Trex: A package manager for Deno
Indeed, this is a massive issue with any url based imports. Because Trex supports nest.land, this is not an issue. nest.land is actually a first-of-its-kind blockchain module registry and CDN. Because we use the blockchain for storing modules, they can never be deleted or altered in any way. This also means that they are permanently and indefinitely resolvable from the web. Because of this, module vendoring is no longer an issue!
tatef··on Trex: A package manager for Deno
Trex is not a company; Trex is a product under the crewdevio organization on GitHub.
tatef··on Trex: A package manager for Deno
In my opinion, Trex is actually working against what npm introduced to Node. Though I don't exactly know what "main problem" you're referring to, I can say that:

1) Trex is supporting multiple module registries, not just one. (Thus not necessarily centralized) 2) Trex is not associated with one entity in particular. This means that they have no company dedicated to hosting modules in house (another reason that they aren't centralized) 3) Node's package.json included many other things than just imports. Using Trex is completely optional, and if you do use it, an import map does not hinder one's development workflow. In my opinion, it makes dependencies easier to manage.