Big if. More likely, it seems, is they started with an open LLM model and fine-tuned and repurposed it via their "RLCD" process.
Big if. More likely, it seems, is they started with an open LLM model and fine-tuned and repurposed it via their "RLCD" process.
Where Does Our Training Data Come From?
TypeSafe is primarily a data research lab, which is how the biggest results in AI get made. We make all the data ourselves. We wouldn’t train on your data even if you asked us to (no offense). We do some pretty sophisticated stuff, but if you want to find out more, we’d have to hire you.
I mean, this claim is simply preposterous, and is discountable as ridiculous nonsense on its face.
How do you create "100% synthetic data" that is filled with countless facts, coding patterns, medicine, law, philosophy, etc? The notion is farcical.
This is a ridiculous conversation, but their claim is such laughable bullshit that it's amazing that anyone actually buys it.
... so not made in house.