Mercury Coder: frontier diffusion LLM generating 1000+ tok/sec on commodity GPUs
inceptionlabs.ai
inceptionlabs.ai
This means that in the worst case the number of iterations is the same as a classic autoregressive transformer.
So they are mostly taking advantage of the fact that the average response is in reality not fully sequential, so the model is discovering the exploitable parallelism on its own.
This is not too dissimilar to a branch and bound algorithm that has a worse theoretical runtime than a simple brute force search, but in practice is solving the integer linear programming problem in almost polynomial time, because not everyone is encoding the hardest instances of problems in NP as integer linear programs.
Without a full paper, it's a bit hard to understand the full details. Does this essentially replace nucleus sampling with diffusion, or does it change the "core" transformer architecture in a major way?
I'm curious about something that has analogues in image diffusion models -- you can see diffusion models, depending on how they are working through their latent space, sometimes try out and then move on from a feature in an image as it fits less with what's around it.
Are there analogues for Mercury? Does it try with a token or set of tokens, and as parts of the response fill in move on from them? Similarly, this architecture seems like it would have real problems inserting a needed token in the middle of a bunch of relatively high confidence generated tokens.
Can you give some insight / thoughts from the frontlines on these?
The moment I heard the synopsis of the technique, I thought of one thing: Style transfer.
This model style should be really nice for translation and style transfer tasks. Takes an existing section text, noises it, and reverses it with guidance like an image diffusion model; A "movement" in latent with a controllable amount of modifications.
The diffusion process enables a wide range of "control" approaches not possible with current transformer models. Perhaps summarizing text can be done differently as well, taking an input and diffusing it into a shorter and shorter section.
I've not been this hyped about a new method since GPT3 itself.
Do these run on actual commodity GPUs like RTX 3090s and what kind of tokens/sec is expected on those?
Also, there's no paper, no open weights, no code. Just an API?
Companies like Groq and Cerebras already hit these kind of numbers over a year ago so I'm not seeing what's HN worthy here.
We're going to be releasing a tech report soon, stay tuned!
As far as I understand (based on reading for less than 5 minutes) it is still a transformer model, but they simply start predicting tokens in random positions, with the possibility of updating existing tokens, rather than producing tokens from left to right.
That isn't too different from say good old stable diffusion.
When o3 solved the ARC[1] challenge, Sam Altman said it was recognizing the image pixel by pixel linearly, and i found it pretty hilarious that they start by using a chip which computes in two dimensions, they use it to compute a line, and that line is a two dimensional image of the ARC challenge.