We have heavy concurrent usage, and we haven't had a single issue yet.
180 karma · joined March 31, 2014
We have heavy concurrent usage, and we haven't had a single issue yet.
QA analysis of voice transcriptions. Napkin math: we operate at 2-5% of the cost of running on Equiv Frontier, though this changes near-weekly because pricing is so volatile.
It took us about a month to get the inference configured to achieve these numbers. But if you can get your hands on a pair of B300 GPUs and the context works, it's untouchable for price/performance.
(B200 would work, but you don't have the B300's memory, which lets you run it on 2xGPU instead of 4xGPU... with Dspark, it's like magic)
On a side note, for tasks that don't require the intelligence of DS v4 flash, we're using Nemotron-3-super with incredible success. I'm shocked we're not seeing more adoption of this model, given how easy it is to fine-tune and how blisteringly fast the nvfp4 version is. (A single B200 GPU can produce an insane amount of throughput with Nemotron 3 Super.)
I love it for a few things, but it's gotten really hard to spend any extended amount of time with it because of the lack of mental model I seem to be able to hold while working with complicated problems.
I'm guessing it's just not enough time doing RL on human feedback.
Check out the anouncement of Inkling (https://thinkingmachines.ai/news/introducing-inkling/)... the section in the middle
"Early in RL verbose, grammatical" (if you search) :
We need to understand the operator. The 5D line element is ds² = e^{2A(x)} (ds²_4d + dx²), where A(x) = sin(x) + 4 cos(x), x in [0, 2π]. The internal coordinate is periodic. The background is a warped product: metric g_{MN} where M,N = 0..4. The internal direction has metric e^{2A(x)} dx²? Wait, the ds² is e^{2A} (ds²_4d + dx²). So the internal metric is e^{2A(x)} dx². Actually if the total metric is ds² = e^{2A(x)} (ds²_4d + dx²), then yes, internal metric is e^{2A} dx².
vs. Post RL
We need determine eigenvalue problem for spin-2 fluctuations h_{μν}(x,y) with TT in 4d and depend on x. For metric of form ds² = e^{2A(x)} (g_{μν}(y) + h_{μν}(y,x)) dy^μ dy^ν + e^{2A(x)}? Wait internal metric is e^{2A} dx²? Actually ds² = e^{2A} [ds_4² + dx²]. So internal metric is e^{2A} dx²; warp factor same for 4d and internal? Yes. We need equation for h_{μν}(y,x) = h_{μν}(y) ψ(x) maybe with normalization. …
I can understand it with less cognitive load in the post-RL version versus early in RL. This resonated with my experience using Fable, especially digging hard problems; it feels like I'm reading the "early in RL" version of that model explanation.
There were a number of use cases where we needed to use Gemini (audio modality), and Ultra has been a VERY cost-effective alternative once we got through the nuances.
Nemotron3-super is, without question, my favorite model now for my agentic use cases. The closest model I would compare it to, in vibe and feel, is the Qwen family but this thing has an ability to hold attention through complicated (often noisy) agentic environments and I'm sometimes finding myself checking that i'm not on a frontier model.
I now just rent a Dual B6000 on a full-time basis for myself for all my stuff; this is the backbone of my "base" agentic workload, and I only step up to stronger models in rare situations in my pipelines.
The biggest thing with this model, I've found, is just making sure my environment is set up correctly; the temps and templates need to be exactly right. I've had hit-or-miss with OpenRouter. But running this model on a B6000 from Vast with a native NVFP4 model weight from Nvidia, it's really good. (2500 peak tokens/sec on that setup) batching. about 100/s 1-request, 250k context. :)
I can run on a single B6000 up to about 120k context reliably but really this thing SCREAMS on a dual-b6000. (I'm close to just ordering a couple for myself it's working so well).
Good luck .. (Sometimes I feel like I'm the crazy guy in the woods loving this model so much, I'm not sure why more people aren't jumping on it..)
The only thing I'd grab dspy for at this point is to automate the edges of the agentic pipeline that could be improved with RL patterns. But if that is true, you're really shorting yourself by giving your domain DSPY. You should be building your own RL learning loops.
My experience: If you find yourself reaching for a tool like Dspy, you might be sitting on a scenario where reinforcement learning approaches would help even further up the stack than your prompts, and you're probably missing where the real optimization win is. (Think bigger)
Overall, it's allowed me to maintain more consistent workflows as I'm less dependent on Opus. Now that Mastra has introduced the concept of Workspaces, which allow for more agentic development, this approach has become even more powerful.
I personally have moved to a pattern where i use mastra-agents in my project to achieve this. I've slowly shifted the bulk of the code research and web research to my internal tools (built with small typescript agents).. I can now really easily bounce between different tools such as claude, codex, opencode and my coding tools are spending more time orchestrating work than doing the work themselves.
MCP standardizes how LLM clients connect to external tools—defining wire formats, authentication flows, and metadata schemas. This means apps you build aren't inherently ChatGPT-specific; they're MCP servers that could work with any MCP-compatible client. The protocol is transport-agnostic and self-describing, with official Python and TypeScript SDKs already available.
That said, the "build our platform" criticism isn't entirely off base. While the protocol is open, practical adoption still depends heavily on ChatGPT's distribution and whether other LLM providers actually implement MCP clients. The real test will be whether this becomes a genuine cross-platform standard or just another way to contribute to OpenAI's ecosystem.
The technical primitives (tool discovery, structured content return, embedded UI resources) are solid and address real integration problems. Whether it succeeds likely depends more on ecosystem dynamics than technical merit.
Well played sir! Nice shot man! :D
I'm so tired of arguing with ChatGPT (or what was Bard) to even get simple things done. SOLAR-10B or Mistral works just fine for my use cases, and I've wired up a direct connection to Fireworks/OpenRouter/Together for the occasion I need anything more than what will run on my local hardware. (mixtral MOE, 70B code/chat models)
I also hang out on a few Discord servers: - Nous Research - TogetherAI / Fireworks / Openrouter - LangChain - TheBloke AI - Mistral AI
These, along with a couple of newsletters, basically keep a pulse on things.
It wasn't that I didn't know the stuff, I do, but more helpful with quickly organizing and presenting information in a clean and well-written way. I did have to go through and re-write parts of it specific to our domain.. but it saved me many hours of work doing tedious organization of data.
I also tested it with helping create some SOP's for a new position in our very small company, even breaking down the expected tasks into daily schedules.
It's not that it's perfect, but it generates a bit of a boiler-plate starting point for me which then I can work with from there.
I own a Tesla Model 3, my wife drives a BMW i3, my daughter has a leaf.
The ONLY car we can effectively travel outside of the greater Tampa area without major headache is the Tesla.
The ONLY car that I would try to drive to New York from Florida in is the Tesla. (Yes, we've done it.. but would only try it in the Tesla)
Having a small buffer, even a few minutes of battery power to rely on while trying to get back to the field to land the "impossible turn" [1] would make me feel a lot better and could be the difference between life and death.
There are various groups (Pipstrel[2], Diamond[3]) that I'm aware of that are working on electric GA aircraft. For young pilots that are looking to train, the cost of jumping in a Cessna 172, the gold standard in GA trainers will cost at best $120-200/hr. Electric costs should be 1/5 (or better) of that in reality due to the absolute bargain of replacing a TBO electric engine, scheduled maintenance and relatively low level of complexity.
[1] https://www.aopa.org/training-and-safety/air-safety-institut...
[2] https://www.pipistrel-usa.com/electric-propulsion/
[3] https://www.flyingmag.com/diamond-da40-hybrid-electric-proto...
I own a P3D, facing the same issues of long road trips 1-2x a year I concluded that I'll just spend the few hundred dollars and rent a car for the deep edge cases of my driving and the other 99% of my time I'll enjoy driving my Tesla.
99% of my normal day to day driving I just charge at home at night, I've used the supercharger network on a long road trip and it was surprisingly little-hassle.
It's cute as hell.