There's no way a CSR has any power over this.
12,226 karma · joined December 14, 2013
There's no way a CSR has any power over this.
It's worse than that -- they don't even know the result! They never tried to run it!
* LLMs don't use Markov chains, * LLMs don't predict words.
The other notable thing about the situation is that three companies ended up simultaneously responsible for a large part of the PC platform, originally -- IBM, Microsoft and Intel. They all worked in various ways to encourage competition to each other -- the reason we see OS competition on the PC platform is that IBM and Intel both found it in their interests to allow other OSes on the platform to reduce Microsoft's leverage over them. IBM in fact created one of the competing PC OSes out the gate, OS/2, which was originally an IBM/Microsoft joint project until they started feuding. Now, OS/2 is dead, but IBM's interest in being able to support their own OS instead of Microsoft's is a big reason the PC platform was built in an OS agnostic way. People criticize UEFI for locking down the PC platform more than the previous BIOS implementations, but UEFI is still _way_ more open than basically any other platform, most of which don't have a standard for bootloaders at all. It's really the absense of a standard for bootloaders that keeps most Android phones locked down. Two Android phones from the same OEM might have different bootloaders, much less two phones from different manufacturers. We've yet to see an alternate OS with the resources to support implementing their own bootloaders for a majority of Android phones.
ARC-AGI-2 keeps a private set of questions to prevent LLM contamination, but they have a public set of training and eval questions so that people can both evaluate their modesl before submitting to ARC-AGI and so that people can evalute what the benchmark is measuring:
https://github.com/arcprize/ARC-AGI-2
Cursor is not alone in the field in having to deal with issues of benchmark contamination. Cursor is an outlier in sharing so little when proposing a new benchmark while also not showing performance in the industry standard benchmarks. Without a bigger effort to show what the benchmark is and how other models perform, I think the utility of this benchmark is limited at best.
- "Here is a spec for an API endpoint. Implement this spec."
- "Using these tools, refactor the codebase. Make sure that you are passing all tests from (dead code checker, cyclomatic complexity checker, etc.)"
The clankers are very good at iteratively moving towards a defined objective (it's how they were post-trained), so you can get them to do basically anything you can define an objective for, as long as you can chunk it up in a way that it fits in their usable context window.
And this is exactly why I think the question of "what is the correct way to regulate car ride services" shouldn't hinge on incumbency bias towards taxis, but actually ask the question of what is best for participants in the market (which doesn't just include taxis and Ubers but also includes public transportation and its users, for instance). But that doesn't fit neatly into Doctorow's enshitification narrative.
``` To navigate all of these technical minefields, you need the help of a third party. In a modern society, that third party is an expert regulator who investigates or anticipates problems in their area of expertise and then makes rules designed to solve these problems.
To make these rules, the regulator convenes a truth-seeking exercise, in which all affected parties submit evidence about what the best rule should be and then get a chance to read what everyone else wrote and rebut their claims. Sometimes, there are in-person hearings, or successive rounds of comment and counter-comment, but that’s the basic shape of things.
Once all the evidence is in, the regulator—who is a neutral expert, required to recuse themselves if they have conflicts—makes a rule, citing the evidence on which the rule is based. This whole system is backstopped by courts, which can order the process to begin anew if the new rule isn’t supported by the evidence created while the regulator was developing the record.
This kind of adversarial process—something between a court case and scientific peer review—has a good track record of producing high-quality regulations. You can thank a process like this for the fact that you weren’t killed today by critters in your tap water or a high-voltage shock from one of your home’s electrical outlets. ```
And this is central to Doctorow's point, right? The narrow question of the legality of Uber's current service offerings is actually pretty well litigated, and if Uber was as flagrantly illegal as he claims, "we're an app" wouldn't have kept them in business. Doctorow argues that this is happening through regulatory capture -- the case isn't primarily that Uber is violating the currently existing set of laws, regulations, court precedents, etc. It's that Uber is violating what the regulations _would be_ in a world where they had less market power with which to influence regulations.
And so it's not enough to argue about how the apps get around _current_ laws. By Doctorow's own arguments, we're debating the merits of a counterfactual set of different regulations that we would have if you changed current conditions. And at that point, it is absolutely fair game to ask if this counterfactual set of different regulations is actually better for market participants.