Together AI raises a $102.5M Series A
together.ai
together.ai
The training revenue stream can sustain the company over the next couple of years as it develops other product lines. They get 2 big bonuses as part of this too: 1. training/finetuning llms for a bunch of companies exposes them to the big problems in this area that customers are willing to spend $ on 2. they're building a big distribution network with AI buyers
This is defensible too. They have a horde of big customers/spenders that are likely incapable of undertaking these efforts on their own and have very high switching costs. Something that you can not say about the big llm providers. Together is sort of flying under the radar serving a great market segment/need while there is a lot of expensive battles being fought everywhere else in the AI space.
This is definitely a company to watch. I wouldn't be surprised to see together.ai becoming a top 3 player in this space.
I'd like some more detail on how they differ from the other 50 or so inference providers currently out there.
I always figured inference was something GCP/AWS would eventually get into their platforms
Inference hasn't really picked up revenue-wise (across the space) comparing to training and it's not a great market to be in. As you mentioned, it's crowded and the barriers to entry are minimal. Anyone with experience in spinning up containers and scaling them can offer this servicer. Paradoxically, it's also the market where the big cloud providers are very well positioned to dominate. Spiky and unpredictable workloads is where their bread and butter is. Their whole economic and infrastructural model is pretty much tailored to this traffic pattern.
Training is a totally different ball game. It is a model that is disruptive to big cloud providers given that it follows very different traffic patterns. Training LLMs involves spinning up 100-1000s of machines for a relatively short period of time and with interconnect that doesn't typically exist in data centers. That is a very unique workload. Additionally you need significantly more specialized ML knowledge in tensor parallelism, optimizations, CUDA etc. That is not as common as scaling a container based workload..
Fun fact: Oracle is surprisingly well positioned in terms of their interconnect fabric. Even Microsoft is partnering with CoreWeave for GPU clusters because they dont have as much capacity interconnected in the right way.
I agree that supply is an issue, but paradoxically the fact that these GPU cloud providers (CoreWeave et al) are partnering with the big cloud players says that the big cloud players are where people would prefer to buy. Once supply constraints are solved, these providers would need some novel offering beyond “we have hardware”, e.g. some specialized distributed training framework. But MS/Google/AWS are also building their frameworks so…
And then the elephant in the room is: compute spend so imbalanced on training vs inference. Why? Is it that there arent enough real use cases? Is it that improvements are so frequent it makes sense to toss out older versions? Is it that privately trained models are a requirement for the highest spenders? My impression is that a lot of corporate spend at scaleups is purely speculative r&d to evaluate capabilities but thats a small sample from friends
My experience with ML projects is that while there is churn in the modeling, most of the effort for a long lived system still goes into the data, but engineers really want to work on the modeling/infra pieces much more than data quality.
Which is to say, I have a lot of skepticism that this is a long term business.
Could they book the money TogetherAI spends on GPUs as revenue now, and then when TogetherAI flops and returns five cents on the investment dollar write it off in a different line of business?
I find this impossible to believe, given they are far from a monopoly here. Where are you getting the figure?
None of the cloud providers are monopolies, but there are few enough they can collude through tit for tat, though at times I think some were run for growth at a loss.
I believe all have multi thousand percent margins on egress bandwidth.
Many examples, but this was the first relevant link on DDG:
https://www2.deloitte.com/us/en/pages/audit/articles/a-roadm...
4B MODEL, PRICE 1K TOKENS: $0.0001
register with an email, test account has 25$ credit, python API as well, the model is smaller, but good to have some fun with the API integrated with other systems, haha...
Per flop the 4B model here is slightly more expensive than GPT4.
If prices scale linearly, we reach 0.03 / 1K tokens (lower end of GPT-4's price range) at about 0.0001 * 300, or about 1.2T parameters (and that's dense - no MoE here).