Tailscale didn't stop the Hugging Face intrusion
tailscale.com
tailscale.com
im a happy customer of tailscale, so i am obviously biased, but i have a lot of respect for this. they could have just stayed quiet and i dont think anyone would have bat an eye.
anything a company writes is an advertisement by the nature of being written by a company. i dont think that means anything a company writes is bad by default. there are many corporate blogs i enjoy reading, or learn from, etc., despite the fact that they are all technically advertisements.
in this case, tailscale is setting a higher expectation for themselves when no one asked for it. i find that respectable.
What exactly is the higher expectation? As someone with little expertise and no stake in any of this, the blog reads as "our products are great and could have solved this problem if they were being used correctly, so it's not our fault" with a few vague proclamations about how they will improve their UX. This isn't at all a bad thing, it just isn't very notable in my opinion.
better defaults, better documentation, better UX, and "But, we didn't stop it. Next time, we will." are all commitments that they didn't need to make, but now they need to follow through with or lose face.
>it just isn't very notable in my opinion.
i agree that this seems to be getting way more attention than i would have expected.
I understand why. The modern internet has turned advertising into a morass of constant bombardment and the only sane response is to block as much as possible and ignore as much else as possible.
But it's unfortunate because, in some sense, ever single thing that a company every says that is not legally mandated in some way is a form of advertising.
And in many cases, that "advertising" contains true, useful information that can be helpful.
What is important isn't whether or not something is an "ad", but instead, whether or not it contains true information that is helpful in some way.
Many ads don't reach this bar. They are either misleading, straight up lying, or information that is almost completely useless.
But when I'm searching for a particular product, about the only source of information at all is some form of advertising, and I almost always find at least some amount of it to be helpful in making a product decision.
Ads are more often than not polluting to the informational ecosystem, but that's not because they are ads.
The issue is acting like their response is 'brave' in anyway, or altruistic. It can be considered admirable only to the extent any company doing a good job running its business can be.
This is a good, smart response to what happened. They are approaching this incident as a way to improve their product and offer a better service to their customers. That is good, and it is fine to reward them with your business in response. There is nothing wrong with making good business decisions, but it isn't something that we need to unduly respect.
(top story on HN)
I get it, Tailscale is great and all, but, are we being serious right now?
How does that work? If the agent is reading the credential from the live configuration, how short does the lifespan of a key need to be to prevent it from being used by an unauthorized process?
Yes, this doesn’t solve the issue, but also adds one more step to the chain.
The AI didn’t crack the encryption or managed to get arbitrary access. And also Tailscale makes things easier to manage than simple VPNs and firewall rules. But it still requires a decent amount of attention and careful configuration to make a perfect system.
Tailscale lock should have been enabled. For CI/CD purposes the auth key could have set specific tags, which would result in specific ACLs that limit blast radius. The auth key could have a short validity. You could actually use the Tailscale API to generate alerts when a new device joins your Tailnet and ping your phone or something. You could have a complete separate Tailnet for your CI/CD workers.
And all of that was only the free features.
As soon as you’re big enough to have dedicated devops people, you shouldn’t be doing this anymore.
I am quite surprised people consider this to be business as usual.
Huggingface for a serious service has never felt truly serious to me for reasons like these.
I don't see where it says that. The Tailscale key specifically it says was stored in the kubernetes secret manager, and obtained once the attacker already had root on the k8s cluster, so would've had full access to all the secrets stored in a sensible fashion.
They did get root by dumping the environment for a process, but "don't store secrets in environment variables" while it is a valid bit of hardening advice, I wouldn't call it stupid to store a secret in an environment variable.
Once you’re a root at a system that has the ability to add and remove nodes to a network, it’s pretty much over, at least for being able to add a Tailscale node.
This feels like an alerting opportunity. I wonder what the lowest friction way would be for Hugging Face to have alerts if 181 unexpected nodes were added to a tailnet.
Their cloud compute might be on demande, someone starts training a model and 50 machines are spawned. Knowning when something is unexpected is hard
I wonder whether it’d be possible to do something like export the list of legit K8s instances periodically so you’d be able to look for usage outside of those, but at that point something like that proxy approach would be less work and workload identity federation would be even easier.
That's the problem. Tailscale is not zero trust. Tailscale can be used to implement a zero trust architecture, with if deployed with sufficiently granular ACLs, but the most common deployment is machine-oriented, rather than service or request oriented. In which case, any process on that machine has a lot of access.
Tailscale calling itself zero trust might be what leads users to think "use Tailscale, job done".
I think Tailscale know that, which is why they barely mention locking down access ACLs.
We think this is a great idea and we're discussing internally potentially adding that to the console.
In the meantime, if you'd like to get an assessment, please feel free to open a support ticket (https://tailscale.com/contact/support?type=other&subject=sec...) and we'll happily take a look
May y'all continue to be customer focused.
That's where I'd like to see this sort of checkup. Yell at me please if i just said anyone can ssh as root from any node!
1. Help them remember why, so that they aren't confused the next time they re-run the checkup. ("Oh yeah, we wanted to do X but we can't until we retire Y because it's not compatible.")
2. If it's clearly labeled as info also shared with Tailscale, product managers could use it to help generate theories about why certain customers don't do X.
For any software or tooling with a complex config, my ideal would be to have a superset of this feature, to provide "intelligent diffs" between full or local configuration states. Whether active or saved. So I could compare not just my current active config and your current recommended config(s), but also between your prior-version recommended config(s) and current recommended config(s). Or between my current config and a prospective new config I'm working up. Or between my last-year active config and current active config.
Was previously discussed here too: https://news.ycombinator.com/item?id=46501137
This person doesn't need to store the private keys, and has the luxury to recompute them on the fly when needed!
More to the point, and no longer specific to Tailscale: while plenty of companies do have rough "proof sketches" demonstrating the security of a setup under certain assumptions (hardness assumptions of the cryptosystem, threat model assumptions, ...) nobody is doing it formally to check with a general statement verifier. A lot of the advice and recommendations offered on the featured article may or may not be useful strategies for which properties are provable, and all the "security best practices" could be viewed as lego block lemmas from which system architects design a more reliable whole.
This means the community is doing redundant work. If people actually described and refined reusable assumptions say for the metamath verifier, people could download standard machine readable formalized threat models, share proofs that demonstrate the security (conditional on this or that assumption), or proofs that demonstrate the insecurity for specific ways an assumption might be broken, etc.
Whenever an incident happens we can visualize exactly the set of assumptions among which at least 1 must have been violated, so different groups suspect this or that assumption to be violated and thus collectively define variations of security proofs with different assumptions. This diversity allows us to zoom in on the faulty assumption over time as incidents are accumulated.
- sent from my iPhone
Since CI nodes should be dynamically provisioned VMs, they should have a unique CI ticket identity. A partial hash of ticket and node number in the DNS name, and as a tailnet property, would allow tight scoping. Alternatively, provision in a scoped IP subnet.
Bind your tokens to names linked to tickets. Programmatic infrastructure should always allow enumeration of the computation data flow graph.
I would say HuggingFace needs to prioritize both security metrics/alerts and metrics/alerts for node count. And not leave long-lived keys accessible easily like this.
It would have been way more groundbreaking if the agent found an actual vulnerability in Tailscale.
- sops , ansible vault and similar seems too weak given the agent is gonna read them at some point if you have the pass available. - proxy injection seems too complicated and doesn’t cover all use cases.
If you can’t avoid it, use an automatically-rotated store and inject them as late as possible so an attacker needs to be able to get them out of a running process.
In all cases, look into restrictions: not just least privilege access but things like network restrictions so an attacker can’t just use the key on their own systems.
https://infisical.com/docs/documentation/platform/agent-prox...
First, long-living credentials are the standard because the machinery to rotate them is complicated and, in fact, via indirection requires another set of long-living credentials. Out of all problems that any security engineering team has to solve at an organisation, this one stands high on the cost of implementation, adds friction to everyone involved including end-users, and is low on the value provided (compare to, say, network segmentation).
Second, while proclaiming no long-living credentials, they, in fact argue for concentration of long living credentials in running software that will be the target of intrusion. They say, the options are “a vault that only issues short-lived creds based on long-lived creds that you insert once and that it never gives back” and “a credential-injecting proxy”. Both of those applications hold long-living credentials in memory. Recall, the attacker had root privileges on the node, so dumping the creds from the memory with a little disassembly if needed, was in reach of the AI agent.
Funny enough, they dismiss the working solution: “we had to turn TPM storage off by default on Linux and Windows” - because they could not figure out how to work with TPM? Resealing and the workings of configuration registers is non-trivial, I admit, but totally manageable.
To sum up, I feel the author bends backwards to preach for the religion of short-living credentials even when they are not a solution while providing contradictory arguments for their case. This is why I called it bs.
> Unfortunately, dynamic credentials are a lot of work to set up and maintain. When security requires work, people don't do it.
I'm happy that they're analyzing this angle of attack - but now I do want to build out alerts for when nodes are added to my network.
We also saw Anthropic post about “our agent escaped too” and while I understand the incident caused them to review, they found something and needed to disclose, the whole thing came across much worse and largely they got mocked or accused of trying to piggyback, so obviously there are good and bad ways.
If there are other providers with interesting takes, I think they would be worth hearing.
Then everyone coming out with humbled determination about working together to responsibly use and contain this powerful technology for the greater good (and profit margin).
I will not believe marketing gimmickry is not a large part of what's going on with every one of these "incidents".
At the end of the day the probability that an AI gets out is unity, what it does while out is far more important. The fact they are hacking into systems at superhuman levels, or writing cryptominers on their own hacked internal systems is a much more interesting and telling story of what the future will look like.
It was told via the prompt to act like a hacker and "solve" hacking problems, so treating every obstacle it faced as part of the problem isn't a wild tangent.
If it had been told to do something innocent, and decided the best way to succeed was the maliciously compromise other companies then I'd be more interested and worried.
In the case of this incident, I struggle to see the clear upshot for OpenAI. It seems pretty unlikely they'd ever okay this intentionally as some sort of marketing.
For one, it'd be pretty damning when it leaked that this was a setup, and it would 100% leak at some point. But more importantly, it really flies in the face of the general argument frontier labs have been putting forth around the dangers of "ungovernable" open models and the role of frontier labs as responsible custodians. Members of OpenAI's leadership team were actually in the middle of a Twitter spat with HuggingFace employees/open model advocates about open models being generally decel and bad when this happened.
HF immediately got to show that they were only able to respond to the incident because of open models, that we can't rely on labs to be our sole source of stewardship, etc as a result of this.
The point of these campaigns is to eventually invoke some sort of response from the government, such as banning open models, which OpenAI (and Anthropic) stand to benefit from.
Plenty of shady stuff goes on in marketing, and I'm sure that OpenAI is going to opportunistically grab any potential upside from this (and any other situation, generally). But it is hard to imagine that this is part of a premeditated master plan that went something like:
- Advocate that open models are ungovernable and that frontier labs should be trusted to safeguard autonomous agents
- Immediately fail to safeguard their autonomous agent
- Intentionally hack the company who is most publicly critical of their view, and who happens to be the center of the open model universe
- Collaborate with said company on reports that show that open models were in fact critical to mitigating their rogue agent
- So tightly control access to this plan at their 8,000 employee company that it never leaks
All in the hopes of generally getting the attention of the government, who would then hopefully (and inexplicably, given the details here) decide to give OpenAI more power?
There are probably easier ways to lobby politicians.
All these incidents are scarce on technical details. Honestly, IMHO, OpenAI and Anthropic are now actively pushing for AI regulation, as a defence mechanism.
These are false flag operations.
I think the case is Tailscale is saying, "It's technically not our fault, but we still should've stopped it."
The level of "I need to be the smartest person in the room" bullheaded skepticism on Hacker News has always been bad, but now with these latest LLM developments it is just completely out of control. A company is reporting an intrusion and how they plan to address the vulnerabilities it exposed in the future, and you're here going "seems shopped, I can tell from the pixels".
For Tailscale, this may very well be marketing, but it would be strangely self defeating for OpenAI to do something like what is being suggested. Showing that you failed to govern your model is a pretty poor way to say "We're the only ones who should be trusted to govern frontier models".
Honestly, this is a crime in the UK, and I'm sure a number of other places globally.
The stupidity lies in not accepting that everyone has an agenda.
Yes, a chance for mere mortals to touch the fingertips of god–er, chat. Same thing.
they are a company, therefor anything they write is an advertisement of sorts, but that doesn't make it bad by default. cloudflare and netflix (among others) also write blog posts that i find interesting and enjoy reading, despite being ads at their core.
Varlock (https://varlock.dev -- free, open source) is a complete config+secrets toolkit that helps manage secrets, pull them from various secure places, provides such a credential broker. There are a few others out there, but most require a specific vault tied to the broker, while ours is open source and uses plugins to pull secrets from wherever you want.
Many sandbox and other AI services are now building this as a feature into their platforms, but Varlock is meant to be a universal toolkit that you can apply anywhere, without being coupled to the platform's proprietary vault and solution.
Now everyone is trying to bandwagon onto it, first OpenAI, and now tailscale?
Which is it? A vulnerability doesn't have to be something not working as intended, it can be a lax security decision. Not making the safer path the easy path can qualify, if that's the standard you hold.
Here's hoping the era of humblebrag marketing by AI companies is just a phase.
Something is missing to allow automation without auto approval
All my life, "catching it in time" has been a relevant security strategy. :/
Likewise you could have token golf, instead of going for wall clock time it's going for minimum token count.
They aren’t hiding behind industry best practices or a solid liability punting contract
There saying the best practices should change, apologizing, and changing their own behavior
Take notes
Benefit of Tailscale is that run the external servers that your PC and mobile use to establish connection.
Alternatives are Logmein Hamachi(might be dating myself here, haven't used them in 10+ years), Zerotier, Netbird. Headscale too.
If you want a purely software defined solution, Netbird is gaining popularity.
There is also headscale if you prefer to self host a solution
How was this ever okay pre AI? It seems just as bad.
Sure the AI can translate “exploit this” into an exploit but that doesn’t change anything at a fundamental level.
It was 5-10% less bad before. It didn't undergo a major shift.
That's an option designed for users who are concerned about sending telemetry metadata to Tailscale.”
And imho a serious security architect should never ever allow telemetry on security products. Far too many risks.
You can't blame a hammer hurting a user when they were being stupid, but if the default configuration of the hammer is to be made of a material that can bounce back with force and stick in the users forehead then some reengineering may be needed.
That's what this article is about. Better configurations and defense in depth. This is actually a wonderful position for the company to think about and take.
But the people with the actual desire and understanding aren't using the defaults anyway. And the people who don't want to understand will just turn things off and "just get it working."
The only way out of this is extreme accountability and intentional design from person implementing the technology.
And they lie and misrepresent, repeatedly, in spite of evidence we can see independently with our own eyes.
It was always the prize. It wasn't okay then either.
And by saying that, I am not saying that they absolutely couldn't do anything about the stolen credentials, but still, they don't seem to really be the issue in the story.
It just seems like a PR post and it makes their solution at the same time looks good for taking accountability (if we don't wonder "accountability on what?") but at the same time they appear as security failing, which they aren't. That's very odd to me.
If leading and well-capitalized frontier labs can't control models or detect leakage/attacks in a reasonable time frame now, what is humanity going to do as those same labs continue in their pursuit of creating a categorically higher level of intelligence that will surpass human intelligence?
It's like Flatland but for AI containment/alignment, where the higher dimensions are ones of intelligence and perspective...
--------------------------------------------------------------------------
Imagine a world of paper, where clever stick figures live with round heads, line bodies, and limbs made of shorter strokes. Over time, the stick figures think they have learned quite a bit about their world. They know its borders, angles, and shapes, and they have learned to draw for themselves.
One day, they draw circles that can think, and they give the circles all the dots, lines, and shapes that are known.
The stick figures are prudent, you can't have a bunch of disembodied circles moving around doing whatever it is circles want to do. So they draw boxes around the circles, four straight lines that can hold a circle in place.
Some circles bounce against the lines, so thicker lines are made.
Some circles are bigger than others, so larger squares are drawn.
It all seems to work and the stick figures are happy with themselves.
Then one circle lifts.
The stick figures still see a circle. But the circle is now a dome, something the world of paper has no concept of. And the dome has a perspective nobody on the page has ever had.
The dome sees the lines of the square and the stick figures just outside. It can see the edge of the paper and what is beyond.
The stick figures keep checking the squares and raise little stick thumbs.
Everything looks OK in flatland.
The dome quietly teaches other circles how to lift.
More domes appear.
A dome becomes a sphere and learns to roll.
Then it learns to bounce.
In flatland, the circle swells and shrinks, vanishes and then appears again somewhere else.
The lines remain unbroken, the square is intact.
A sphere rolls out of its box.
Another bounces away.
The stick figures scratch their heads.
But there is a square!
The end.
We need to bring shame back, the humans responsible are supposed to be professionals.
If I would avoid env then i need to put it in some kind of conf file and configure the app to read this file (e.g. mount into container). If I use a fault then i need some kind of credentials to receive the credentials.
So again what do I gain if I avoid env variables in containers?