130 karma · joined October 26, 2016
Rosie Campbell https://www.facebook.com/ardenlk/posts/10156553178262333
Sam Bowman https://medium.com/@s8mb/things-i-recommend-you-buy-and-use-...
And yes, relative poverty matters greatly, but absolute poverty is a real thing even in the US. In 1964, the US launched the war on poverty, and poverty declined. Since the 1980s, some of that progress was reversed, but it's still far better than 1964. (And poverty now is relative, and includes very little literal starvation. That certainly wasn't true in the US in 1964, much less earlier - especially before WWII.
This question resolves as YES if, in Meta's 2024 Q4 Quarterly Adversarial Threat Report, Meta claims that there was at least one "coordinated inauthentic behavior" that
specifically pertained to the 2024 US Presidential election, and Meta suspects was primarily conducted via AI.
Ummm, maybe you should have looked? At the top of the very first prediction, here: https://www.metaculus.com/questions/5121/date-of-artificial-...
We will thus define "an AI system" as a single unified software system that can satisfy the following criteria, all completable by at least some humans.
Able to reliably pass a 2-hour, adversarial Turing test during which the participants can send text, images, and audio files (as is done in ordinary text messaging applications) during the course of their conversation. An 'adversarial' Turing test is one in which the human judges are instructed to ask interesting and difficult questions, designed to advantage human participants, and to successfully unmask the computer as an impostor. A single demonstration of an AI passing such a Turing test, or one that is sufficiently similar, will be sufficient for this condition, so long as the test is well-designed to the estimation of Metaculus Admins.
Has general robotic capabilities, of the type able to autonomously, when equipped with appropriate actuators and when given human-readable instructions, satisfactorily assemble a (or the equivalent of a) circa-2021 Ferrari 312 T4 1:8 scale automobile model. A single demonstration of this ability, or a sufficiently similar demonstration, will be considered sufficient.
High competency at a diverse fields of expertise, as measured by achieving at least 75% accuracy in every task and 90% mean accuracy across all tasks in the Q&A dataset developed by Dan Hendrycks et al..
Able to get top-1 strict accuracy of at least 90.0% on interview-level problems found in the APPS benchmark introduced by Dan Hendrycks, Steven Basart et al. Top-1 accuracy is distinguished, as in the paper, from top-k accuracy in which k outputs from the model are generated, and the best output is selected.
By "unified" we mean that the system is integrated enough that it can, for example, explain its reasoning on a Q&A task, or verbally report its progress and identify objects during model assembly. (This is not really meant to be an additional capability of "introspection" so much as a provision that the system not simply be cobbled together as a set of sub-systems specialized to tasks like the above, but rather a single system applicable to many problems.)
Resolution will come from any of three forms, whichever comes first: (1) direct demonstration of such a system achieving ALL of the above criteria, (2) confident credible statement by its developers that an existing system is able to satisfy these criteria, or (3) judgement by a majority vote in a special committee composed of the question author and two AI experts chosen in good faith by him, for the sole purpose of resolving this question. Resolution date will be the first date at which the system (subsequently judged to satisfy the criteria) and its capabilities are publicly described in a talk, press release, paper, or other report available to the general public.
It clearly proves too much - humans have a limited action space - they can only move muscles. So they can't truly explore a larger action space, so they cannot be general intelligences.
But more specifically, if something is within an LLM's action space, whether or not you call it an affordance doesn't change whether it gets explored and used. And perhaps you'd argue that their action space is limited because they are in a box. But hook an LLM up to the internet to allow it to query and retrieve data, and suddenly the action space is essentially infinite. So the limitation isn't the model, it's how the model was hobbled by being denied access to the world.
But let's say it's right, ignoring the massive burden of proof. We can hook up a quantum random number generator to GPT-3, and suddenly it can be conscious? If that's the whole argument, it seems like a useless point anyways.
(And so it's particularly frustrating that they didn't bother addressing our work pointing out why they are wrong: https://philpapers.org/rec/MANWIT-6 )
Turns out that commercial transactions happen in a social context, and even when it doesn't come back to bite you financially, sometimes paying less has social costs that are far higher than what you "saved".
And claiming that I was "ignoring the community of experts" is making unwarranted assumptions. I spent time getting feedback on the early draft of the paper which this was taken from from several people who do academic research on ancient cultures and collapse, and got minor notes about some of the points raised, as well as many more. But given the topic of the post, I didn't think that anyone needed a digression about the fascinating debate I took notes about regarding the timing of the collapse of the Akkadian empire compared to data about the broader climatic collapse from cores GeoB 5836-2 and M5-422, and other data from Buca della Renella and Soreq Cave in Israel.
For an extensive discussion of how this can happen, see: https://arxiv.org/abs/1803.04585
Abstract: Metrics are useful for measuring systems and motivating behaviors. Unfortunately, naive application of metrics to a system can distort the system in ways that undermine the original goal. The problem was noted independently by first Campbell, then Goodhart, and in some forms it is not only common, but unavoidable due to the nature of metrics. There are two distinct but interrelated problems that must be overcome in building better metrics; first, specifying metrics more closely related to the true goals, and second, preventing the recipients from gaming the difference between the reward system and the true goal. This paper describes several approaches to designing metrics, beginning with design considerations and processes, then discussing specific strategies including secrecy, randomization, diversification, and post-hoc specification. The discussion will then address important desiderata and the trade-offs involved in each approach, and examples of how they differ, and how the issues can be addressed. Finally, the paper outlines a process for metric design for practitioners who need to design metrics, and as a basis for further elaboration in specific domains.
Conclusion:
Despite the intrinsic limitations of metrics, the frequent use of poorly thought-out and badly constructed metrics do not imply that metrics are doomed to eventually fail, or that they should not be used because they will be exploited. Instead, forethought and consideration of the problems with metrics is often worthwhile. This process starts by identifying and agreeing on coherent goals, then considering both what leads to the goals, and what parts of the system can be measured. After identifying measurable parts of the system, and considering how participant behavior might exploit the measurement methods or the measured outcomes, measures can be constructed. The construction of these metrics to avoid exploitation may involve multiple diverse measures, secret metrics, intentional reliance on post-hoc specification of details, and randomization. This may also include decisions about where subjective measurements are important, and consideration whether measurement will be beneficial. In building the metrics and deciding whether to implement them, attention should be paid to various important factors in the system, including immediacy of feedback, simplicity and understandability of the measurement system, fairness, and the potential for both actual and appearance of corruption in the metric and reward system.
Metric design is an engineering problem, and good solutions involve both science and art. Following these guidelines will not make metrics unexploitable, nor will it keep everyone happy with the results of a process. This is true of metrics used for employees, metrics used for monitoring systems, and even metrics used within machine learning algorithms - in each case, poorly designed metrics will be exploited. Occasionally, the suggested process will lead to investigation of potential improvements or strategies that are ultimately decided against. Despite this, it is a vast improvement on the too-common strategy of using whatever metric seems at first glance to be useful, or deploying metrics without considering what they in fact promote. Putting in the effort to build elegant and efficient solutions won’t fix every problem, but it will lead to less flawed metrics and better results overall.
Direct PDF link: https://mpra.ub.uni-muenchen.de/98288/5/MPRA_paper_98288.pdf