I wonder if this waters down the “distillation attack” claims by Anthropic. They have their own RL environments! I guess the caveat is that the RL datasets are still opaque, nothing is really proved.
263 karma · joined September 26, 2024
I wonder if this waters down the “distillation attack” claims by Anthropic. They have their own RL environments! I guess the caveat is that the RL datasets are still opaque, nothing is really proved.
There’s a more nuanced discussion to be had about AI when it comes to potential and existing benefits, open source, the geopolitical dynamics, self-hosting, the financial dynamics, the sanctity of human thought, a future of even more personalised advertising and so on. We do need to elevate the discourse around AI and shift the public narratives around what US frontier labs are pushing, understanding their biases but also acknowledging substantive claims, potential benefits and risks.
I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!
I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.
Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.
An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.
This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.
But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.
But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.
If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible.
My understanding is that the HuggingFace incident would not have occurred with a model that was not an unfiltered internal preview instructed to roleplay an attacker, with access to abundant compute, resources and a slack sandbox to reach its goal.
Embedded human auditors will improve safety standards, but the structural solution is mandating accountability for actual agent operators.
> AI models are dangerous. They can help bad people do dangerous things. They may be capable of autonomously executing dangerous things. They may cause unwanted effects on the labour market.
> More capable AI models are more dangerous, but require more money to train.
> Money requires investors with expectations that the model will generate a profit over its operational lifetime.
> The operational lifetime value of a model is decreased if every model is public and can be hosted on any infrastructure.
> If investors see less operational lifetime value from model companies, model companies receive less capital and therefore train more capable models at a slower rate.
Follow-ons: There are immediate risks in releasing capable cyber models that can be ablated and then launch cyberattacks. Concentration of power moves immediately into the infrastructure layer for inference.
Models will be kept private for longer, if not indefinitely, given there is less incentive to release them publicly.
Enforcement globally will occur because accessing the lucrative US market means using an open source model.
The entire argument hinges on strict enforcement of this policy, when AI model routing can already be opaque.
It also hinges on investors being rational and expecting free cash flow from AI companies, rather than reaching a criticality threshold of model capability for recursive self-improvement internally and parlaying that into a global mega-corporation.
New labs and companies without infrastructure connections will no longer be able to raise money, given investor expectations, and therefore not be able to increment AI progress.
So a win for the infrastructure layer, models being kept private for longer, new labs being unable to compete, immediate risks in the rollout (I suppose could be mitigated by a staged rollout i.e. policy active in 2030, pricing effects now), enforcement being tricky, and in the event that foreign competitors lead the AI frontier and then start closing models, just conceding the market to them if US consumers and companies find workarounds to pay for foreign AI.
But the truth is, much of the utility of the model is allowing it to grep across the codebase and explore for context.
I have great interest in zero knowledge inference, but as far as we know, it is difficult to build any sort of efficient representation.
Operator error exists with or without AI. Installing any dev tool that could exfiltrate your information means you are responsible for securing your environment.
Use a devcontainer. You can find starting examples on the official claude code and codex github repos. Configuring a firewall script to block egress. That being said, you will likely allowlist openAI domains so I'm not certain if the sites feature will be blocked. Worth testing.
I do not think we can trust private companies, no matter the virtue signaling they put forth into the world, to effectively regulate themselves and inform the public and scientific communities about risks. Their ongoing conflict of interest poses serious credibility risks.
For anything health related all AI models show high levels of anchoring bias. I would not use it as a confidant, and be skeptical of claims. Even so, human doctors are also fallible and prone to cognitive bias.
I think the obfuscation is because human intelligence has been projected onto AI model capability. AI models only have a limited dimension of human intelligence, and in some axes orthogonal, and when I say distillation I refer to this.
Most of the discussion around AGI is highly speculative. I am not saying AGI could not exist, and it is a term that has historically been loosely defined. Decades of coming science and research will tell.
I agree with your last statement.
On a funny note, I think their prompt was:
"Hey Fable. Please attribute every piece of scientific and economic progress to AI until 2040. And predict every major geopolitical event. Make no mistakes."
I am not sure if alternative reality fiction is the best way to approach real and serious AI risks.
I am also not sure, with the amount of emdashes and the style of prose, that the entire article was not AI generated.
AI is going to be a mature scientific field. There are going to be efficiency improvements in training and inference. New paradigms are going to emerge with better multimodality, real time streaming and real time interfaces. Models are going to converge on the limits of our data available for pre and post training, improvements will be incremental and spiky in domains.
I am not sure who the AI 2040 article is for. I suspect it is intended to be a digestible piece of media for the financial class.
AI is going to be a useful technology and its impacts across the economy and global will be broadly distributed. Because AI represents the distillation of the very best human knowledge and expertise. AI is compression of human capabilities, the very best ones. Maybe the argument is that in verifiable domains, such as model training, AI models can supercede humans. I don't think so. A human's high level thinking, our incredibly more efficient semantic/neural compression, our ability to switch tasks and achieve the creative insight is not replicated through the current paradigm.
My comment wasn't supposed to be a jab at Proton. No service can have 100% uptime.
But I signed up for Proton and paid them just for a secure and stable contact email for my domain registrar. And when I couldn't log in for 15 minutes, with opaque errors and requests just being timed out, it was definitely a surprise.
I'll keep using their service. The refund remark was probably a bit polemic.
I bought the yearly paid plan because they won't close it for inactivity. I'll look into mailbox.org or another solution, maybe even self-host if I need to migrate my work email!
I assumed an email service was supposed to be stable first, given how important it is. I was going to use proton mail as my contact email with my domain registrar.
This outage may change my mind.
REFUND?
Today I decided to sign up for a proton mail account to move my domain registrar contact email to a provider other than Gmail. I changed my registrar contact email around 30 minutes ago.
Now it seems like Proton mail is down, or that I am unable to authenticate. The issue is at their end, my outgoing auth requests get hung and timed out. Tried all the usual fixes, multiple networks, cache, incognito firebox, google chrome, etc.
Now the status page shows they are investigating an incident.
I am now second guessing my decision to use proton mail at all.
Obviously, my domain registrar allows me to change my contact email at any time so this is not a real issue at the moment but I have neverexpeirenced a gmail outage.
The timing is rather serendipitous, or whatever the polar opposite of serendipity would be here. FUNNY!