305 karma · joined December 6, 2012
"The 2 May 2023, 6 months later, the regulation started applying and the potential gatekeepers had 2 months to report to the commission to be identified as gatekeepers. This process would take up to 45 days and after being identified as gatekeepers, they would have 6 months to come into compliance, at the latest the 6 March 2024.[8][32] From 7 March 2024, gatekeepers must comply with the DMA. [33]"
Legitimate scrapers, maybe. Everyone else does it to circumvent the API limitations, by posing as real traffic. APIs imply API keys which can be traced and banned.
The paper mentions accuracy i.e. (true positives + true negatives) / total examples. And it's actually 100% accurate i.e. there are no false positives _or_ false negatives.
But the big caveats are:
1. this was tested only on 180 examples, which is a very very small dataset to draw conclusions on, and
2. this is obviously an adversarial space so any classifier will be obsolete with the next training run
I'm bearish on any attempt to distinguish real content vs. AI generated content (on any medium, text, image or anything else). This is an adversarial game and the AIs can incorporate your fancy algorithm to fool you better. In the end these projects only end up improving the AI models in terms of realism.
- calorie counting or diets: just write down what you ate, and it should be good enough to compute calories, macros etc.
- gym logging: write down when you went to the gym and what you did, and it should give you tips, help you maintain your routine and other helpful things
We did no specific training for these exams. A minority of the problems in the exams were seen by the model during training, but we believe the results to be representative—see our technical report for details.
This is absolutely spot on, with the caveat that you do need to disaggregate from accounts to people, which is the hard problem. Having people call a phone number is definitely not going to work as a way of achieving this disaggregation. I'm pretty sure I could create a system to bring that call center to a halt with fairly minimal cost in less than a week of coding.
As an attacker, you can also hire people in call centers to make phone calls at scale for you.
> Again, it's not a "request" [..] suspension notifications can also be automated.
Can you clarify what you mean by "protecting" them? I'm not sure suspension notifications qualify as meaningful protection
Pick two.
Different companies do different trade-offs. The optimal solution depends on how the internet community weighs each individual axis
Maybe. A few problems here:
1. payments come with privacy concerns, unless maybe you're talking about zero-knowledge-based blockchains, but we're a LONG way from such functionality being widespread
2. $0.001/email is actually very reasonable for an attacker; they'd probably gladly pay even up to $1 or more, depending on their exact needs, especially if that comes with an elevated privileges account
3. all of this is easily defeated by fanouts. E.g. if they sign up with bob@gmail.com and then are able to use bob+1@gmail.com, bob+2@gmail.com etc. to sign up for a different service, this defeats the purpose
E.g. if a spammer can pretend they're 10 million different people, and each of those "people" requests an explanation, the whole system grinds to a halt.
This is the reason behind a push for more KYC-like verification on these platforms (e.g. asking for IDs). But this comes at a huge privacy cost for legitimate users. So one way or another people who are real, legitimate and with good intentions somehow pay the cost of the harm that is being done on the internet. This is a hard problem.
Source: am thinking/working on this sort of stuff; not representing my employer, my opinions are my own etc. etc.
I did exactly that with a hobby project (https://www.acceleratul.ro/). The funny thing is that I had to add a short (~50ms) artificial delay with a spinner before showing search results, because people completely missed the fact that the page had refreshed
As of April 20 2020, there were 5 in clinical evaluation and 71 in preclinical evaluation. I'm a layperson, but very interesting to see the different techniques and testing methodologies used.
Plus:
- The type system. It can make your life a huge pain, but in 99% of the cases, if the code compiles, it works. I find writing tests in Haskell somewhat pointless - the only places where it still has value is in gnarly business logic. But the vast majority of production code is just gluing stuff together
- Building DSLs is extremely quick and efficient. This makes it easy to define the business problem as a language and work with that. If you get it right, the code will be WAY more readable than most other languages, and safer as well
- It's pretty efficient
Minus
- The tooling is extremely bad. Compile times are horrendous. Don't even get me started on Stack/Cabal or whatever the new hotness might be
- Sometimes people get overly excited about avoiding do notation and the code looks very messy as a result
- There are so many ways of doing something that a lot of the time it becomes unclear how the code should look. But this true in a lot of languages