> This included making direct code changes to the Tesla Manufacturing Operating System under false usernames and exporting large amounts of highly sensitive Tesla data to unknown third parties.
This sounds like something out of the 1990s, that dark and romantic era of version control when we thought CVS was pretty cool actually and we didn't know what key-based authentication and 2FA were.
There are volunteer-ran projects that don't have this problem.
Edit: to be clear, I presume no one is debating the fact that someone with high enough credentials can push code to production. The questions that the email raises are:
1. Why can anyone, regardless of credentials, push mission-critical code without review (or, alternatively, if the changes did go through review, why did the review process not catch multiple malicious changes?)
2. Why can someone compromise several high-level credentials without anyone figuring it out (the changes were made, apparently, under "false usernames")?
Or a contractor that was fired had their credentials appropriated by this manager, perhaps by that manager removing them from a "delete these accounts" list.
Those are a couple of mundane ways of getting a false username to a production machine. This is even easier when there is a lot of flux at the company with many people coming and going, a lot of account management happening etc.
It could have been that the accounts were local to specific machines and not managed by the company as a whole.
And -- keeping in mind that production type machines operate machinery that can kill -- this sounds okay to you?
Not to mention this:
> Or a contractor that was fired had their credentials appropriated by this manager, perhaps by that manager removing them from a "delete these accounts" list.
...keeping in mind that production type machines operate machinery that can kill, does it sound OK to you that anyone can get access to an account that they don't own and control it?
This particular case would be enough to have PCI certification come into question (if not for it to be revoked), and that's just about money, not life-and-death stuff.
Then again it could be just as simple to create an alternate fictitious identity without going through IT but just by accessing the systems you have permission to access anyway.
The former restriction is maybe difficult enough to efficiently implement in an organization that it's excusable (we have a scheme for it at $work, but it unfortunately means that sometimes people show up at work and the paperwork isn't ready yet and some of the accounts they need aren't yet ready).
The latter, on the other hand, is security 101 and not implementing it on the production floor is just irresponsible. I really hope it's not what happened.
And I'm not joking - recently I was asked whether something (that was designed for a clustered Linux environment) could run on Windows XP because that's what was on the machine they wanted it to run.
Why do you suppose the unauthorized party was following the company's development practices? Maybe it was from the sysadmin side, somebody who worked on the toolchain used for reviewing and pushing things to production. So he was able to sidestep the normal review process. This can happen, what is important is that such things are discovered.
He should not have been able to sidestep the normal review process. That's the problem in the first place. Even if you're from the sysadmin side. It should not be possible to do it.
You may think that looks exaggerated but I've worked in two places where we implemented such a process, both of them far more boring than Tesla and, I suspect, far less money to burn on infrastructure.
> This can happen, what is important is that such things are discovered.
No, what is important when working with mission-critical code is that such things are mitigated. Discovering such a problem in production code is already a problem, not a solution.
No, seriously.
1. Sign every review 2. Use the review signatures + the manufacturer's keys to sign reproducible builds of the production image (i.e. you cryptographically certify that "this image is authorized, and it includes this list of commits, that has gone through these reviews"). 3. Use a secure boot scheme of your choice to ensure that only signed images can be installed on production servers 4. Keep anyone with 'wheel' access away from the image signing keys, and anyone who can generate images away from 'wheel' access.
This way, you make sure that no one who has 'wheel' access can install a sabotaged image, that any image that can be installed has gone through an auditable trail of reviews, and reduce the attack surface that a malicious developer has control over to stuff that requires root access (which is still a lot of surface, but is harder to sneak past a review).
Root access to production servers does not need to mean that you can install arbitrary code on them, and with the right systems engineering, you can ensure that it does not trivially result in arbitrary code being run on production equipment.
Edit: this is all in the context of "questions that Tesla's answer raises". For all I know, the answer might be that they hired some brilliant genius who figured out how to sneak by whatever secure boot scheme they're using. The point is -- the post that sparked all this is not naive. This is real stuff. Companies that are concerned about it can ensure that unauthorized commits are so difficult to get into production that a disgruntled employee would rather just quit than go through with it.
It's possible a sysadmin with low-level access can exploit that and a variety of zero-day exploits and escalations of privilege in the layers above to systematically compromise the boot images, steal or falsify credentials and signing keys, and circumvent the safeguards and alarm systems which should be in place to prevent malicious modifications of the source code and the compiled binaries, while hiding his actions from his co-workers, sure whatever.
And if that's what happened to Tesla, wow, sucks to be them, that's amazing.
But if there are no safeguards, no review process, no alarm bells to go off and any damn person can submit malicious code effortlessly and they were basically working off the honor system... I'm going to blame the victim a little bit.
Depending on what those "other user accounts" had access to, it could go in many different directions. :)
2. If you already have privileged accounts you can escalate in pretty much any environment. And they obviously did figure it out.
There is really no reason for this to be the case. Certainly all code that actually runs on the car can be required to go through review and be verifiably built, even if server code standards are more lax.
If they found a way to use more than one username they may very well have run it through the review process and approved it themselves
Programmers: read-write to repository
Staging team: read-only on repository, read-write to test servers and staging zone
Deployment team: read-only on repository and staging zone, read-write to production
It wouldn't prevent malicious code going out but at least would require a chain of cooperation between employees, which would be harder to achieve.
No, it just requires a chain of cooperation between authorized accounts. That's a very important distinction, especially here, where the email in question alleged the following:
This included making direct code changes to the Tesla Manufacturing Operating System under false usernames
* Storage engineers: Generally have access to most storage (all of dev/test/prod) in their group. Sometimes their access is silo'd, sometimes not.
* Backup engineers: Generally have read/write access to _everything_, and all historical versions of it, as backup systems need to be able to do both read/write. Fairly often there are ways for this access to be "unlogged" too, so the actions aren't captured into any system auditing logs (otherwise it can screw things up). I've not (yet) seen access for backup engineers ever be silo-d, but some places might be doing it and I've just never seen it. :)
It very much sounds like thats the case here - production code was edited, and subsequent auditing has found what should've been deployed and what is deployed differs.
And what stops a single member of the staging or deployment team patching the build scripts, binaries or just installing their own software to a server?
Given this, I think it’s much more likely there were few or bad controls, than a person in an incredibly privileged position working in the margins.
That's completely unnecessary and should not be the case. If you need something pushed quickly, you can get a colleague with review bit and get them to ack for "urgency" reasons after a quick lookover.
I believe this policy, used only a few times per year, has saved us 8 figures in outage costs over a decade. (More than half of the benefit is from a clear statement and instilled sense of ownership, and only secondarily the defusing/unraveling of people would otherwise wait or insert tangles of “best practice” red tape while the website or a factory was hard down.)
I based it on 14 CFR 91.3 (in intent) which says, in part:
91.3 Responsibility and authority of the pilot in command.
(a) The pilot in command of an aircraft is directly responsible for, and is the final authority as to, the operation of that aircraft.
(b) In an in-flight emergency requiring immediate action, the pilot in command may deviate from any rule of this part to the extent required to meet that emergency.
(I made part c of the law, the reporting requirement, mandatory where it’s only on-demand in the aviation law.)
When we explain the policy to new employees, we often cite the aviation law directly, to help clearly communicate our intent.
Often it boils down to not taking sufficient measures against social engineering. In this case the claim is they used fake usernames - most places I've worked, successfully getting a fake account if you already work there would tend to "only" require a willingness to lie on a form or two ("fake" a contractor) and then request elevated privileges. Very few places I've done work requires sufficient checks or counter-signatures to require additional accomplices or make it harder than that. They do exist, but they're rare.
The state of security most places is quite depressing at times. Then again, most of the time it's enough.
Not that that's necessarily better... Manufacturing equipment's at about the same danger tier as cars.
While you're absolutely right, it is far more likely that someone was able to do this because there is a lack of security in their software engineering practices. This isn't aimed at you, but I think a lot of people are blinded by their support of Elon Musk and Tesla to acknowledge that ultimately he's one man, and Tesla are just a company that makes cars. People and companies are fallable.
There are areas that I couldn't push code willy nilly, namely in the security space. But I'd be willing to bet that a majority of teams at any bigN could have a single bad actor cause some damage... That's just the maturity of the industry.
What? Why?
Nobody on my development team has access to the production code signing keys. And nobody - at all - has the ability to remotely make a production system take an unsigned update.
And in my specific case, the group with the singing keys is also the group paid to tell us "no" whenever a release is blocked by process reasons.
I got dragged into defending my technical solution, but my point was that if the ability isn’t needed, don’t have it. You can break the glass when when you need to. Good process make deviation from it more visible. Our code signing keys belong to a team we already need wet ink signatures from to release software.
I can go into the biohazard labs with shorts on easier than I can leverage our technical disaster recovery. I’ve only ever had to do the latter.
You know that backup system your production servers have?
That also has perfect _write access_ to everything, and it's needed so _restoring_ after something goes wrong works.
I have very different experiences from a SEC regulated company. With SOX there are controls to prevent such a thing to happen. If this is a SOX breakage, Tesla is in deep trouble with the SEC.
So Sarbanes-Oxley came about because Enron were stating investor-facing metrics that didn't match reality. SOX compliance comes about if you:
* are publicly-listed
* announce numbers to investors
If those numbers are counted by a computer, you must show that you have procedures in place to prevent a single person from making a change to the code which would allow them to choose what number is produced.So, if you were some web app that mentioned Monthly Active Users in your quarterly results: bam! Your release process now has to be SOX-compliant, since the MAU calculations come from analytics (from the frontend, or from access logs) which could be altered by some nefarious code.
Note that the intent of the act isn't to prevent such a thing from happening, but to make sure there is enough information for external auditors to detect it if it occurs.
One person writes a requirement, this needs to be OKed by another person, then a third person writes this code and it's pushed to production. Controls are setup - e.g. checking JIRA tickets in git logs - that no code without proper authorization (corrrect JIRA status) is pushed and deployed.
People need to be able to trace every code change to the requirement and the OK.
In the core this only applies to systems that are in one way or the other relevant to financial data (like ordering), but auditors usually want to be better safe than sorry. But Tesla might have a SOX-IT and non-SOX-IT.
I would expect that there are limits to this. If a rogue employee engages in fraudulent behaviour against you, using "false usernames" to subvert your security as this employee reportedly did, then I don't see how the organisation could be considered responsible.
Many orgs don't track github usernames to real human mappings well or at all, mostly because the single sign on version of github is 3x the price.
As a developer I'm not allowed to have write access to any production system, except in an emergency via a break-glass mechanism, which is audited to the hilt and back.
It also means we're not allowed to deploy software to production systems. This has to happen via a specific chain development > regression / user acceptance testing > production.
All those environments need to be physically seperated with very specific access requirements.
The deployment process needed to be signed off by outr auditors.
I can't speak for other banks, but they probably need to implement the same -, or a similar system.
Neither can I speak for SOX requirements regarding software fo car manufacturing.
What if the employee were a sysadmin-level person that sidestepped the normal process?
E.g. run quarterly reports for changes to prod that didn't go through the regular release pipeline and note down which P0 they corresponded to.
Most companies do not have good security. Even when they do, it's hard to get it right, especially for internal attacks. Don't underestimate what a single individual can do when they're already inside and well-informed.
No, not for a company doing what Tesla does. In the environment I used to work in, that would have been flat out impossible. No change would have been allowed that did not follow process unless an explicit and documented exception was made.
Even if you did manage to commit code to the release trunk without review, every single change to the codebase was checked before the release process started in order to verify that the right process was followed. If we ever found a change that no one could find paperwork for, it would be reverted.
So, yeah, it's possible if you want to do it badly enough.
Code reviews are a thing, but physical/mathematical assurance that zero people in your organization can bypass them are not. However amazing your tower of automation and policy may be, there's at least one sysadmin underneath it all, and at least one guy with keys to the cabinet underneath him. That only starts to diverge when you're running on FIPS 140-2 Level 3+ HSMs and you need to assemble a quorum of operators to do anything. And those are quite easy to sabotage - just trip the tamper protection mechanisms.
If it was a totally new user, shouldn't the system have a list of approved users and a strict protocol to validate a new user?
What does being rude add to the discussion?
Edit: see closeparens comment above. Complicated systems can always be subverted when trust is broken.