832 karma · joined June 1, 2023
> That's why I created isitsaas.io: your platform for determining whether your llm wrapper has product potential, or whether people would rather just use claude directly instead. Simply pop in your business idea, and our award winning, proprietary technology will ask claude if you can make money off of it.
> Trial memberships start at $10/month
I clicked through the leaderboard. It seemingly isn't sorted by best or worst performing, but rather by absolute change in performance - improved and worsened models are interleaved. Regardless, I selected the model that worsened the most, and investigated what they claimed their fine tuning accomplished; in fact they claimed nothing, their model sheet was the default trl model sheet without even the training set listed. The second worst performing was the same, but they at least listed their training set. It was a set of millions of word math problems. So we would expect the model to get better at those problems and worse at everything else.
On the other hand, the most improved model on the table had an actual model sheet, where they claim the model has been trained to be "more helpful". They claim that by ablating model refusal the model ends up doing better on benchmarks as well. That claim seems to have been borne out.
Obviously if you train a model for some purpose it will not necessarily do well at that purpose. But I would expect popularily used models to be ones that actually do what they say they do. I don't think people need to be told that sometimes models are badly trained. And the claim that this project actually has data to support, that finetuning on one task does not necessarily improve performance on another task, is so obvious that it doesn't need to be stated.
There's nothing plausible that can do that, and he knew that, and was basically just saying "we're going to put ads on it" in a way that challenges you to explain what a better option would be.
I appeciate short letters like this that get straight to the point...
2. As you can see in the quote, the appeals are complete
Such charges are outright nonsense, for one thing. And I think that's relevant since it's the only possible jutification for the utter forfeiture of rights.
Holy crow what a mess. No wonder I'd never heard of this.
There should be a proposal for browsers to expose a property on the keydown/up/press event containing a code for the key combination. Something like "CTRL+S", "CTRL+ALT+S", etc. The programmer could then switch over this property rather than having to check key codes and modifiers manually.
I would also propose to any web developers that they build this property themselves in their own code and check against it instead of checking modifiers directly. Not only would it protect against bugs like in the OP, it would also be a lot more convenient to use.