1,589 karma · joined April 1, 2010
In addition to all the country codes TLDs that do 3rd-level registration, the PSL does also include stuff like github.io. (Maintenance of the list involves manual volunteer labor, so scaling is a real problem...)
(And of course the PSL wouldn't work well for the .name situation, where it's sometimes 2 and sometimes 3, and it can change over time. But that's no excuse for this clusterfuck of just suddenly dropping a bunch of domains that are paid up years in advance.)
No, they use the same tokenization as everyone else. There was one major change from early to modern LLM tokenization, made (as far as I can tell) for efficient tokenization of code: early tokenizers always made a space its own token (unless attached to an adjacent word.) Modern tokenizers can group many spaces together.
I think the pseudocode in the paper is very hard to beat.
This is fairly easy to see, if you consider a stream with some N distinct elements, with the same elements in both the first and second halves of the stream. Then, supposing that p is 0.5, the first instance will result in a set with about N/2 of the elements, and the second instance will also. But they won't be the same set; on average their overlap will be about N/4. So when you combine them, you will have about 3N/4 elements in the resulting set, but with p still 0.5, so you will estimate 3N/2 instead of N for the final answer.
I have a thought about how to fix this, but the error bounds end up very large, so I don't know that it's viable.
- Random bad luck.
- As you say, failing to control for something -- although, if you then treat the lowest datapoint as being effectively the default risk, this would suggest support for radiation hormesis (that people who got a bit more than background radiation actually did better.)
- Some kind of data collection artifact. Perhaps the people with the absolute lowest dose, in a radiation-worker dataset, are selected for being ones who are not getting an accurate measurement (i.e. sloppy about wearing dose badges or something), and those people genuinely do have worse outcomes.
If you're fitting a function which grows asymptotically (i.e. is monotonically increasing at least past a certain point), the best (polynomial) fit absolutely cannot have a negative quadratic as the leading term. If your model gives one, it is 100% guaranteed to be an artifact. Treating it as "suggesting some downward curvature" is a pretty bad misunderstanding.
If you have doubts about this, consider what would happen if we added datapoints at higher doses. Every single datapoint we add to the right side of the graph will make the fit of a negative quadratic significantly worse. Ultimately, if you continue the graph indefinitely to the right, the fit of a negative quadratic is guaranteed to be infinitely bad. Any hint to the contrary is inherently an artifact of the limited dataset.
(It may well be the case that, under certain conditions with a range-restricted dataset like this, such a finding might indeed be more likely if the true function has some downward curvature. But that's not statistics, it's voodoo. All the associated statistical parameters, p-value, likelihood ratio, etc., are absolutely meaningless nonsense.)
As I recall, it's fairly random and fairly rare, but apparently it's been seen a bit more during COVID.
(Context: I maintain an app for an event that runs about one week a year. The ability to use free dynos the other 51 weeks means we retain the ability to do one-off analytics queries, minor development work, etc. during the off-season, without having to delete and recreate the app or something every year. Eco isn't quite as good as free, but it means we can still have separate staging and production instances during the off-season, without paying extra for staging to be idle, and without having to destroy and recreate staging every year.)