1,118 karma · joined March 19, 2020
If the conversation was in the training set, there's a high likelihood that the small set of conversations related to solving Navier-Stokes was used by the model. I get Astra to still quote some of my friends' books or blogposts nearly verbatim on certain niche issues.
Much more importantly, we _can_ determine whether a conversation was used in the training data. And if it was, it gives us a great idea whether that logic was captured in reasoning for a novel problem never yet solved.
Given that you don't see any of this as below the belt according to your other comments, maybe your contribution here is more for yourself than a fair conversation about attribution.
I'm not coming from a place of distrust here. This should just be definitively answerable given the weight of the claims here. Surely between you, your lawyers, and other members of your team you can just clear this part up.
Stains an important moment in the history of AI progress for me. The future seems bleak with people like this at the reins.
Amazing results for open weight and that size, but a really long way off, and I'm extremely skeptical of benchmarks that show these smaller models as being anywhere close to Opus 4.8 (or even earlier Opus's).
That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.
I think you might be interested in reading in training data generation, training, and post training papers/articles. I think you might be surprised at how much intervention there is on some of these levels.
Every startup I know that hired some level of jr's and encouraged them to use AI found themselves in a code review bottleneck for basic code quality and architecture decisions. Many startups I know are largely forgoing jr devs.
Code quality regarding evolvability, reliability, maintainability is still part of the development process. All of the people who claimed src code is just assembly on Twitter some months ago have started chiming in that they were wrong.
I do wonder what this does to the talent pipeline like mentioned in the post.
Particularly for reusable liquid engines :)
I use both ChatGPT and Claude for engineering work on a daily basis, touching performance critical code to application backends to frontend work, and I've found that DeepSWE scores don't reflect my reality when I assess high quality output from the models/harnesses.
Not that Opus always beats GPT 5.5., but that 5.5 is ahead of Opus on a general benchmark smells off to me.
FSL (vs a copyleft license or just plain old OSS) implies they want to turn this into a revenue source for themselves ultimately, unfortunately.
>Maybe I should put one of those buy me a coffee links on the repo
Absolutely :) Cool project.
Not sure why it's such an issue to discuss the political views of the beneficiaries of services we use. I understand it's mostly uninteresting as far as comment sections go, but it's always bizarre to see a defense of political association when often the impetus for sharing this type of information is for people/consumers to exercise their right to associate with business based on their political outlook.
I have no idea why you're making a comparison to a TV show; nothing that was described was anything akin to that. I just made examples out of insufferable and clueless forum comments, that very clearly detract from discussion more than they contribute to it.
I don't think you should assume that describing meaningless and unrelated anecdotes as "uninteresting" is equivalent to users calling for a forum ban, which is seemingly what you're doing when you point to forum rules when encountering a critique.