10,797 karma · joined August 25, 2014
They don't need the resellers help for that! I've already been burned twice thinking "I'll use a raspberry pi for this small project" and having it die within a month because reliably reading/writing to an SD card is hard or something.
Between the notorious unreliability, crazy price hikes, trademark or whatever disputes, and this anti consumer ewaste generating policy, I can't think of any reason I'd ever patronize them again.. my only regret is it took me two purchases to learn this lesson!
> Find unused code in our web app. Exclude generated files and test fixtures, and check for indirect references before recommending a removal.
If I "/goal remove unused code" in Claude today, I would not even think to specify "check for indirect references" and "don't consider generated code dead code," those sorts of intuitions have been "built in" to the frontier models for a while now.
So I'm pretty confused by this, what does it do that I can't with an agent swarm?
It depends what you mean by "reliably." If you mean, "we should be comfortable relying on this kind of tool at scale to identify and punish students, professionals, and writers who may have used AI," absolutely not.
If you take "reliably" to mean "1 in 200 false positive rate" as they disclose on their front page, absolutely that is possible (they are doing it today!). If you think there are more than 200 assignments turned in over a given year at university, you probably do not consider a tool like this fit for purpose. It's an open question whether those procuring said tool are aware of this
Unfortunately their marketing is really insisting on the former, and trying to push it into the zeitgeist that detection of AI-generated or edited text is reliable-type-1 now and long-term. They fail to make it clear that this is merely a tool that strongly suggests text follows patterns known to us at the present time of known LLMs. However, that fingerprint will drift over time, as LLMs get better, human writing style evolves, and the line between human and "smart autocorrect" becomes even blurrier (does speech-to-text push the model into "AI assisted" mode, because it tidied up your punctuation, for example?)
"What color are your bits" is good reading today as it was 20 years ago: https://ansuz.sooke.bc.ca/entry/23
While that is not quite my bar of confidence when implementing wide-reaching technologies that have numerous unexplored knock-on effects, I guess the calculus must have been different on Infinite Loop recently.
I don't think we should have this, for that reason alone (but many others too).
See the headline of this article for Exhibit (I've actually run out of letters in the alphabet) why adblocking is a moral imperative.
> Crack down on unauthorized distillation / prevent weight theft
Actually hilarious to put that in writing, given the genesis of this entire business model.
Laws like this provide cover for abusers and deceivers, by preemptively spoiling objective evidence and making any accusations depend on hearsay instead.
Skimming the PDFs it seems much more dramatic than that? It sounds like at least one of them is concerned OpenAI "solved" the problem by having their internal model use the chats of the independent researchers and want to claim the credit instead? I don't know. The tone is pretty accusational though:
> the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard “forced,” it was a bright red flag.
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien “very little human input” had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. [0]
I was saying if the goal is "text editor," the web platform is the wrong foundation entirely
I don't have some nefarious desire to scare people away from the TLD of their choosing. Really I'm bringing it up to be like "why would you even, like, want some 3rd rate domain instead of getting a .com" so I don't think there's anything to correct
> We have no plans to modify the .name entries at this point in time. We are aware of the implications of adding a wildcard, therefore we won't.
Joe Smith and John Smith can independently register joe.smith.name and john.smith.name, do browsers have a wildcard suffix list for the 2nd level of `.name` specifically, or can Joe set a cookie on all of .smith.name?
What's simple because it's "hard" is replacing parts 2 & 3 with a network appliance like TrueNAS running a zfs pool that syncs to backblaze every night. Yeah you have to learn a bit but it won't fall in weird ways like the hard drive part here will just fail to mount one night and not back things up for 3 months until you notice. My 2¢
I'm kind of sad the author stopped shedding unneeded complexity there though... we're not really building a text editor yet, we're building a website with a fancy input field. If we want to build a proper text editor we must eschew the bloat that is the web browser too.
My charitable read is a legal cartel that allows the small club to switch to Marathon instead of Sprint mode, and drip feed us frontier models at inflated prices, while preventing open source and overseas labs from releasing models because they're unsafe (for the Blessèd Fews' profits).
My uncharitable read is somehow even less constructive..
I think we fundamentally disagree on what "working for me" means, but I remain steadfast in saying we should not accept tools that have ulterior motives beyond producing the output desired of them by me, the user.
> Watermarking the outputs themselves is very different and much more effective compared to how tools like Pangram work.
At the end of the day the only artifact is text that you can do statistics on. It's the same problem as today, with the probability shifted slightly more in one direction. This does not assuage my concerns at all.
> they are incredibly unlikely with SynthID
I kept my commentary focused on text watermarking specifically because I agree, a synth ID image watermark false positive is highly improbable. There's plenty of noise to robustly hide whatever you like in an image. Text is simply too capital I Information-sparse and fragile.
> good faith watermarking attempts bad.
I would sooner call it "ignorant faith" (if they don't know what they are emboldening) or worse "don't care" faith (there will be false positives and they accept this to further some illustrious and arbitrary goal of Text Purity). Whether that be to prevent model collapse or help you not waste time arguing with bots online, to me the principled stance of "tools work for the user" wins..
If it worked perfectly, maybe you could make this argument in a vacuum.
It does not work perfectly. (It cannot. It is by definition a heuristic). That means there will be false positives. There is a chance those false positives ruin someone's career. See [0] for just how easy it is to push SotA "AI text detectors" in one direction or another.
Now, with watermarks, instead of everyone to some extent understanding that AI text detectors are wishy washy woo, they are now Anthropic certified to detect an official AI watermark.
With that kind of false confidence in hand, the people who trust the "computer says you plagiarized" machine are never going to believe you when you say "it can make mistakes," they're just going to fire you/take away your scholarship/cancel your grant/...
This is all beside the fact that we should demand our tools work for us and not for some shadowy master. "Universally good," absolutely not.
[0]: https://freddiedeboer.substack.com/p/i-wouldnt-say-pangram-i...
Why is that so clear? I can think of hundreds of real daily problems I'd rather my legislators be focused on than deciding how old my kids have to be before they're allowed on some website.
I would like affordable groceries and public transport, not whatever this is.
If they really want to "do something" then require an 18+ ID at a cell phone points of sale. That'll be something that won't affect me and actually will temper a ton of underage Internet use. Then they can study the actual effect of that and whether the harms were exaggerated after all.
And, reminder, so no one here loses the plot: this is a backdoor feel good measure lobbied for by big tech to push us into a world with mandatory device attestation so their ad impressions are worth more. They could care less about your kids.
You have the power, and should exercise it, to rate limit bad actors
If that's what the business really believed, they'd put up a paywall instead of literally giving away their content. Some businesses do believe that. Most of them realize it's more profitable to sell data about you instead. So screw 'em and screw this victim-blaming "well we wouldn't need all this invasive tech if it weren't for those pesky adblockers, you'll need a TPM to browse the internet because you didn't play along with my bad business model" nonsense. The internet existed before big adtech, and it will exist after. Long live free computing.
> Like if you don't want to be kicked around, you need to be the one kicking.
This is a punk stance but somehow the argument is "be punk and whip out the credit card?" My argument is "be punk and adblock."