We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
819 karma · joined March 28, 2014
Multimodal models
Language models for art
We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
Step 1. Model can’t do something challenging Step 2. You try a bunch and fail Step 3. Anthropic trains on your usage data. Your current code base and current problem are now in domain Step 4. Model comes out and you’re shocked when it can tackle the thing you were stuck on Step 4. Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set. Feels of it being worse on similar problems
LLMs and KV caches have amazing performance characteristics with concurrent throughput. It scales very non linearly. So the token throughput within a batch scales WAY faster than the tokens per second of each user.
This is the reason the LLM providers have such crazy margins on their costs.
Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.
There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
Where: E = Efficiency, and efficiency gains come from quality of data, quality of algorithms. C = Compute (Size of model, flops of train run)
So a better company can train a bigger and better model with less required compute which let's anthropic get there first. If another company does the same thing with a worse: model architecture, kernel, optimizer, etc... They will get there as well if they just run there train run with more flops for longer
Mythos was actually ready about 6 months ago. So if you have 6 months later or hardware setup and time to train you can get a lot done.
Small llms are still way more efficiently server on big GPUs.
Sharing server capacity takes advantage of the massive parallel throughput and sharing of memory bandwidth.
You are sharing the GPUs with thousands of concurrent users.
I loved writing GPU shaders or optimizing visualization performance, but most of the time it was wiring up netcode to UI elements that exist.
Ironically as I've moved into focusing on more GPU and kernel programming AI is now lapping me there anyway, however the impact of knowing what sort of algorithsm are state of the art in papers, what is causing memory bandwidth issues etc... does a lot to drive the machine.
And reeks of the same sort of reallocation fallacy that makes people think the rich making too much makes them poor.
If we really just say 4xd everyone’s salary in America. Prices are gonna rapidly rise in everything.
Things don’t get materially better unless we build the material things we need. We need to come up with a way to build more houses. Lower the cost of healthcare. Not just increase everyone’s money supply.
They would be impressed with our technology even if it has downsides. Wisdom is knowing humans and technology and imperfect tools.
Composer 2.5 was a huge leap with minimal compute from xAI.
They can compete with OpenAI and anthropic with xAI scale compute. They have a top notch model team and incredible training data and huge enterprise costumer contracts.
They’re using more compute, a bigger model and tons of training quality improvements to get more out of an equivalent model.
Much easier to create a vm testing swarm of 100 disitributions with llms
You need some amount of parallel compute and some amount of global comparison.
And the rest is basically a ways to parameters and scale.
(This is in theory, in practice you can get a lot of small % stability and efficiency improvements that really compound in algorithmic details of model architecture)
We won't reuse open source libraries as libraries we import, but as design inspiration for the bespoke tools we make.
It's too cheap to make your own stuff and too expensive to be stuck with someone else primitives.
But grounding AI Coding in existing tools is incredibly powerful.
Pre-training scaling laws all support larger models being more cost effeceint to train then smaller models. And distillation is comparably cheap. So you can get the most juice by training the biggest model you can and distilling it.
Most software engineers will just need cheap tokens.
But things like physics and drug discovery have no foreseeable upper bound.
Most software engineers will just need cheap tokens.
But things like physics and drug discovery have no forseeable upper bound.
Cursor has 1B in enterprise revenue. It doesn't matter if people can clone their product, those deals don't move slowly
Can you not see the significance of that?