There’s also what is being called mid-training where the model is trained on high(er) quality traces and acts as a bridge between pre and post training
15 karma · joined August 23, 2025
There’s also what is being called mid-training where the model is trained on high(er) quality traces and acts as a bridge between pre and post training
Quality of response/model performance may change though
There’s also nous research’s Hermes’ series of models, but those are trained on llama3.3 architecture and considered outdated now
It’s a trade off I suppose - you can very well host your own streaming solution, and for the same price you can get a great single node, but if you want good TTFB and nodes with close proximity to many regions you may as well pay for a managed solution as the price for multiple VPS/VM stacks up quickly when you have a low budget
Edit: I think I missed your point about bandwidth pricing lol, but the second still stands
Same experience with reloading on the physical machines. It’s quite amusing
https://huggingface.co/datasets/allenai/WildChat
and