HNHacker News
TopNewBestAskShowJobs

mezark

126 karma · joined February 27, 2023

submissionscomments
mezark··on A walk through of the DeltaNet family of linear attention variants
lol - co-founder of Doubleword here. honoured you think we have sophisticated enough marketing to 'newsjack'. What actually happened is my cofounder wrote it over the weekend because he's a mega-nerd and put it live yesterday. I hadn't even read it until I saw this hacker news thread lol

We're just a group of guys and gals who like inference!

mezark··on You Could Have Come Up with Kimi Delta Attention
And a PhD in Quantum Computing! I'm a physicist so a fan of bra-ket tbh
mezark··on NVLink, NVSwitch, and All That
Another banger
mezark··on Width vs. Depth: Speculating on the Margin
really like the framing of this post :)
mezark··on Anatomy of a high-performance EP kernel
I love this blog
mezark··on Artificial intelligence is not conscious – Ted Chiang
(As someone who cares a lot about philosophy of consciousness / & cogsci)

The whole point of consciousness being a 'hard problem' is that we just cannot make claims like 'X is not conscious'

mezark··on Bringing Up DeepSeek-V4-Flash on AMD MI300X
we think so - but haven't tested it ourselves
mezark··on Bringing Up DeepSeek-V4-Flash on AMD MI300X
Hi! Co-founder of Doubleword here - we've hugely increased the number of models that we offer (partly thanks to work that we've done on hotswapping https://blog.doubleword.ai/fast-sglang-starts.

We're kind of known for our low prices - our prices (our main usage is for our high throughput API - the async tier) is significantly below average openrouter prices - but cached prices is coming soon which will lower them even more :)

mezark··on Bringing Up DeepSeek-V4-Flash on AMD MI300X
We at doubleword are bullish for AMD for low-interactivity inference - it does just take a bigger lift on the software side...
mezark··on UK sovereign LLM inference
If you're talking about UK sovereign LLM inference you need to mention Doubleword... very serious inference optimization lab in london with public endpoints for OS models
mezark··on Should GPUs Make Free Trade Agreements?
We look at how comparative advantage from economics applies to LLM inference - some GPUs are relatively better at FLOPs, others at memory bandwidth. What happens if you let each do what it’s best at?
mezark··on Our Small ML Team Beat OpenAI and Anthropic in a Specialized Domain [pdf]
Huge congrats - and when you look at the latency graphs as well it really shows the value of these specialised systems!
mezark··on Controlled generation of OS LLMs – without impacting latency
TitanML Takeoff Inference Server demonstrating controlled generation
mezark··on Takeoff Inference Server Is Now Open Source
Drop in replacement for HF's TGI server. The fastest and easiest way to inference LLMs locally

Github: https://github.com/titanml/takeoff Docs: https://docs.titanml.co/docs/titan-takeoff/getting-started Discord: https://discord.gg/83RmHTjZgf

mezark··on Falcon 7B running real time on CPU
Hey there - TitanML is these guys: https://www.titanml.co/ . I think the impressive thing isn't actually whether the model is good (although it is a good model especially when fine-tuned) - but how fast this model runs on CPU with the TitanML server compared with before.
mezark··on Falcon 7B running real time on CPU
Falcon 7B running real time on CPU
mezark··on Amazon Titan
Annoying because they stole my company's name (TitanML - https://www.titanml.co/) Fortunately they haven't trademarked it, but still unideal.