Gitlab servers are being exploited in DDoS attacks
therecord.media
therecord.media
Handles user input data??
Uses ‘eval’???
I'd accept this for maybe pre-2000, but people should really know better.
this doesn't work in perl 5.6.2! grrrr
So this code is at least 20 years old, and probably pre-2000. Sept. 26, 2008 - Version 7.44
- Added read support for DjVu images
There were probably enough systems running perl 5.6.2 around 2008 to cause bug reports, or the code was migrated from an older piece of code and added to ExifTool.It was not uncommon to manually ./configure, make, make install tarballs locally in those days, especially not on systems like Slackware so I can see it being possible to have old packages installed that were not automagically updated.
I looked at the front-page of that library. It says it cleans metadata from a huge number of file format. Frankly it looks more like something you would use on your own, known safe, files before sharing them online.
I'm not sure the tool is presented as a sanitizer for untrusted input. At least, it does not claim to be.
Why does Gitlab need to clean metadata from DjVu files? Wtf are DjVu files?!
It doesn't. It needs to clean metadata from JPEG and TIFF files. They didn't properly check if the files were actually of those types, and Exiftool performed its own content type detection to end up in its DjVu code.
> Wtf are DjVu files?!
DjVu is basically an alternative to PDF.[1]
Today with modern JIT compilers its probably not much faster...
x["a"]["b"]["c"]
And the developer decided that this was best evaluated by eval. During the code review phase I talked to them and asked why they were using eval, and they didn't know it could be evaluated directly as they were a little unclear if javascript supported that syntax. var array = ["foo", "bar"]
I was expecting xml or json (like the rest of the endpoints), but I realized that they just served this as text, and then eval'd it on the frontend...We are already avoiding / trying to avoid way too much interesting and useful things. All for the sake of security and only to encounter new ways to be attacked. Instead of "avoid" how about actually organizing worldwide intolerance and hunt for those attackers.
"eval is evil", if you will.
Lambda allows arbitraty network access and may allow access to your AWS resources.
If you have to do this, the best approach is to containerise it, use capabilities to enforce restrictions and run in a virtual machine as isolated as possible.
It's still not great though. Some languages (eg Java) have additional features that help with this though.
I didn’t mention containers since they don’t provide strong isolation and some people misuse them as though they do. There’s no harm in using them as another layer of defense, but hardware virtualization provides much better security.
In my case, code supplied by the end user is compiled into a different language such that I think I can prevent intentionally-malicious activity.
Nonetheless, spinning up a VM to create an environment in which potentially untrustworthy code is executed before then destroying the VM seems the safest option.
Have you tried writing your own interpreter?
As for mostly-full-featured `eval`: iirc Perl itself has a facility to create restricted sub-interpreters and run scripts that can't do certain things. (Though I might be confusing Perl with PHP here.)
Besides that approach, simply rolling your own parser using a parser combinator library is super simple. The word combinator makes it seem complicated, but it's actually the opposite, using parser combinators is a lot simpler than writing a parser the traditional way you might have learned in formal education.
Implementing a simple DSL like for example an event-filtering language should cost a competent but fully inexperienced programmer maybe 1 or 2 weeks for a proof of concept, and then 3-6 more weeks to get it production ready depending on the feature set of course.
Of course, that's more time than simply running the V8 interpreter over your input string, and maybe running the V8 interpreter over your input string is an awesome way to empower your (trusted) customers.
- the problem appeared in GitLab 11.9.0
- the problem seems to have been fixed in GitLab 13.8.8
- the vulnerability uses ExifTool, so to exploit it, a user needs to be able to upload images
- if an update is not (yet) possible, DjVu format file uploads can be blocked to avert this vulnerability
- this vulnerability isn't relevant for GitLab instances that just have 1 user, or are not publicly accessible on the Internet
Now, i'm not saying that the above is entirely true, but after reading something like the above, one should be able to figure out how to best act: - if you have a public GitLab instance with open registrations, consider updating it immediately (with backups in place, of course)
- if you have a private GitLab instance with many users in your own corporate network (that somehow isn't updated yet) - this is a good reason to put updating it into your agenda today, even if your users aren't necessarily hostile
- if you have a private GitLab instance or one with registrations closed (e.g. you're the only user or people that you trust use it), mark this down and update whenever possible, however it's probably not necessary right this moment
Of course, i can't say the above with 100% confidence, because the article itself lacks this actionable information to aid in decisionmaking and so i'm left to piece things together on my own, because of which i could be wrong.On an unrelated note, DjVu is a pretty interesting file format, though sadly i've only seen it be used very sparsely, on some Russian forums for tractor manuals or something: https://en.wikipedia.org/wiki/DjVu
A lot of private GitLabs contain mirrors of public repositories or vendored copies of public libraries. Our GitLab is private but practically speaking there's probably several hundred people, most of whom we couldn't identify, that could "upload an image" to it.
Can you provide more info? I skimmed the upstream ticket but didn't see how. Getting access to anything other than the login page on an accessible-but-private instance seems like a security bug regardless of this CVE.
The following describes the entire unauthenticated attack:
https://attackerkb.com/topics/D41jRUXCiJ/cve-2021-22205/rapi...
And, if you like that sort of thing, there is a metasploit module you can use to reproduce the unauthenticated attack:
https://github.com/rapid7/metasploit-framework/commit/6f4aa5...
I'm no security guy, but this seems... incredibly dumb? Like even for perfectly secure code, the asymmetry in resource usage alone to submit an image vs. get them to dump a file, shell out to a scanner, and rewrite that file would probably be enough to seriously hurt smaller GitLab VMs.
https://www.amazon.com/Problems-Mathematics-education-instit...
Note that the issue was enabled by GitLab not verify the file format, ie that a .jpg is a JPEG and not a DjVu file for example, before handing it over to ExifTool.
So a simple extension/mime check won't cut it.
Please see this post on the GitLab forum for details how you can determine if your instance has been compromised through the exploitation of CVE-2021-22205: https://forum.gitlab.com/t/cve-2021-22205-how-to-determine-i...
Ah, the good old "File upload vulnerability". File uploads remain one of the hardest problems to solve when it comes to security.
But also, file uploads should be handled in a jail or box of some type - and never let their analysis make network calls.
A hackjob usually has less deploys than my own stuff (which, outside of Windows 2000 components, is less than a few millions)
Just looks at how PHP got so popular. Clearly it was useful and thus became popular. I think it's hard to argue that it wasn't a hackjob when it first started.
ExifTool is essentially for (image) metadata what ffmpeg is for video.
It isn't a hack job, but it started out as one, like so very many other things. Version 1.00 was released end of 2003, while the problematic code was added in 2008. Mind you, the problematic code does not just eval whatever it sees, it tries to make sure the input isn't dangerous first. That check failed spectacularly, defeated by a newline combined with how '$' in perl regex works without any special flags[0]. Using eval was a bad choice to begin with (but I am told something you would commonly see in perl software of the time, yes, even in 2008 still), it was the lazy choice of re-using perl to unescape C-strings instead of rolling your own unescaping code.
So what do you suggest? Use libexif[1]? Exiv2[2]? Where would I run it? Can you suggest any operating system that never had a "stupid" RCE?
>It appears to be a complete piece of shit.
Let's see your code then. All the code you ever wrote that is possibly still in use somewhere. So if you never ever fucked up or got lazy, feel free to cast the first stone, otherwise I would suggest you dial down your rhetoric when it comes to taking massive steaming piles on other people's work.
Yes, Phil Harvey had a "WTF?!"-class security bug here[4], shit happens, "goto fail", let me deRail your yaml and the Debian random number of the day is: 6.
He patched it promptly compared to other vendors and projects (public release on April 13th, while April 7th was the initial bug report, to Gitlab not ExifTool, which Gitlab then passed along[3]).
You can blame him for the bug, you can blame Gitlab for not running exiftool in some sandbox. But that half the gitlab instances remain unpatched some 7 months after patches became available, that you'll have to put on the people running these instances.
[0] https://github.com/exiftool/exiftool/commit/cf0f4e7dcd024ca9...
[1] http://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=libexif
[2] http://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=exiv2
[3] https://hackerone.com/reports/1154542
[4] I cannot be sure if he wrote it, or if somebody else contributed it, but at the very least he didn't catch it during review. Looks like he wrote it, tho.
Why? It seems like they should have read/write but no execute. What goes wrong?
Usually what goes wrong is parsing or processing the files. It's hard to get programmers to safely validate a 20 byte email address string. It takes a lot more care to safely parse a 4,000,000 byte image file in a complex format.
Off the top of my head, a lot of old console exploits on the Wii and GameCube revolved feeding malicious save files to games. Same idea, missing bounds checks or whatever when deserializing some field lets you shell the process. Parsing random files from users is just dangerous.
> When uploading image files, GitLab Workhorse passes any files with the extensions jpg|jpeg|tiff through to ExifTool to remove any non-whitelisted tags.
> An issue with this is that ExifTool will ignore the file extension and try to determine what the file is based on the content, allowing for any of the supported parsers to be hit instead of just JPEG and TIFF by just renaming the uploaded file.
> One of the supported formats is DjVu. When parsing the DjVu annotation, the tokens are evaled to "convert C escape sequences".
1. No filesystem access
2. No network access
3. Input passed on stdin (or a pre-opened fd)
4. Output passed to stdout (or a pre-opened fd)
5. A hard timeout specified before the process is killed
Suddenly bam, dramatically safer.
If you're looking for a tool that can do all of this for you, check out firejail:
https://firejail.wordpress.com/
It has a ton of options, but you can do all of what I suggested and more, really easily.
A user input facing software should always use fuzzing to uncover such bugs.
A usecase for WASM's nanoprocesses (capability-based security) perhaps? Of course, until such a time someone exploits the WASM runtime itself.
Used in Firefox to sandbox some libraries, including image handling IIRC.
The blockers, to me, right now are:
1) I mostly write Go, and the Go runtimes didn't seem to be particularly maintained when I last looked. So it just hasn't been worth it to me to do plugins. (I have done "provide your own code to a Go application" before -- "gojq" and "expr" got the job done. Less features than a full WASM runtime, but still pretty powerful.)
2) It's unclear to me which programming language APIs should target. You add a plugin system and you want developers to use it -- what are the popular languages that target WASM? Go and Tinygo look great here, but I have a feeling that the average programmer wants something a little more dynamic for their small plugins. AssemblyScript obviously wants to be the standard, but it's probably too different from Typescript to make it a no-brainer for Javascript developers. Some sort of Perl/Python/Ruby that compiles to WASM would be great, but I haven't seen much progress on that front.
As for running untrusted code in general, I don't think WASM needs to block you. gVisor simulates the linux kernel for containers, providing stronger isolation between them, and is designed to protect you from things like this. (I think the original usecase was running ffmpeg to transcode user-provided video files?) And you can always go full VM on these things. Or take the nuclear option -- carefully audit the untrusted code and build up that trust ;)
> Thanks in no small part to your recent findings, [GitLab] are rolling back our policy about paying half-bounties for third party findings. These have great impact on GitLab and we want to continue to incentivize research for high+ severity issues in that area.
At least they're taking these things a bit more seriously now.
Seems GitLab agreed on that one, they registered two CVEs:
Yes.
I've had some experience with bug bounty programs. The situation with vulnerabilities in dependencies is complicated.
On one hand you want to incentivize researchers to find vulnerabilities in your dependencies, which likely don't have their own bug bounty programs.
On the other hand, you run into situations where researchers discover a bug in a 3rd-party library and, instead of reporting it ASAP, they keep it secret as long as possible while they build up an arsenal of bug bounty exploits to report to multiple companies that use the library. This is a weird disincentive to fix the underlying bug quickly because the researchers know they only have so much time to exploit it in bug bounty programs before one of the companies fixes it upstream.
We also had a problem where amateurs would spam our bug bounty program any time a CVE came out for one of our dependencies, even if they couldn't exploit it. It was relatively easy to close these bug reports because they couldn't provide a proof of concept exploit, but it still wasted a lot of time arguing with them, especially when they'd try to blackmail us on social media for not paying them out (for a bug that didn't exist in our product).
I do not miss my days of dealing with bug bounty programs. Met a few great researchers, but most of the (attempted) participants were trying harder to exploit the bug bounty program than find exploits in our code.
Sure, my fault for not keeping it up to date. But there is much noise to filter through in the many tools we juggle these days, especially if an organization prefers to self-host.
No judgment. I’m paid to make things, not apply patches. This is however why I don’t use self-hosted, pros and cons, etc.
And I'm not sure about gitlab, but their are often mailing lists for security updates for major software packages.
https://www.linode.com/docs/guides/how-to-configure-automate...
I've never had any problems with it, although I just run a couple of servers :).
If everyone enabled this one thing I'm sure the Internet would be significantly safer.
Obs: You can select security updates only which I believe are unlikely to break anything!
The way to keep up-to-date on critical security updates for GitLab is to sign up for our Security Alerts mailing list: https://about.gitlab.com/company/preference-center/
I once had a CVE RSS feed, but it was mostly noise even after I filtered it to only tools/libraries we used.
If an organization is too overloaded to patch for six months, maybe they should re-evaluate if self-hosting is the best course of action. Seems like this is a foot-gun of your own creation.
The fact that you do need to upgrade it yourself regularly is indeed a drawback. On the other hand, an Omnibus upgrade has only failed me twice in the last five years or so, so there's little reason to not do automatic upgrades at night and fire off an alert in case something doesn't work as expected afterwards. Their releases are typically solid, so kudos to the team.
[1] https://status.gitlab.com/pages/history/5b36dc6502d06804c083...
If you are a 50,000 people company where your own employees could be the adversaries maybe not. But in this case you also have the budget to have a proper security team. Apart from this case, with proper virtualization and no direct internet exposure, you'll find out that 99.99% of CVEs are not a risk to you.
If you are under-budgeted, it's fine to neglect such internal services and check vulnerabilities once a year or less. However always stay on top of the CVEs of internet facing services. My point is your message is basically the propaganda of cloud services "doing it yourself is HARD", "email is HARD", "this and that is HARD" lol.
Self-hosting in jails not being directly exposed to the internet is a productivity booster as it allows you to not fix what works FOR YOU. Just keep using this 4 years old version if it works well for you. But make informed choices as much as possible, try to stay on top of CVEs even if you decide not to patch these internal services 99% of the time. But even if you can't stay on top of your non-internet facing jails, it won't be a real risk 99% of the time.
Suffice to say, we fall under "insanely paranoid".
Also let's not forget that self-hosted still offers features SaaS does not, such as server hooks, which are absolutely not uncommon in grown environments based on gitolite etc looking to migrate.
It's not a silver bullet but in many cases it will be the difference between getting hacked and not getting hacked.
It is totally possible to design applications that can be safely exposed to the public internet but it requires some real effort.
I prefer requiring TLS mutual authentication with a corporate PKI and issuing employees client certificates.
Doing both wouldn't be a bad idea either.
Executing everything on an isolated container with no permissions? Audit trial etc/good logging? If someone comes up with an RCE you're basically done for, you can only mitigate it but not completely stop it.
These types of programs are relatively simple, and this is a case where a formal proof is much better than reliability.
Is anyone aware of research on this?
This is obviously "expensive" though. Doesn't scale very well.
Unlike this issue then, going by the 1Tbps attack it's reportedly causing...
Linux itself has several features that can be used to isolate processes, and there are use friendly tools like bwrap [0] that make configuration easy.
It should be entirely possible to sandbox something like ExifTool itself such that it has no network access and is limited to reading and writing files in a particular directory.
- It's a separate interface with a different attack surface than your system, so compared to a locked-down version of the normal syscall API, it provides better defense-in-depth.
- It's designed to be a fully self-contained sandbox, by default. If you're locking down everything but reading and writing previously opened file descriptors, you can build a secure sandbox atop syscalls fairly easily. If you need more nuance than that, WebAssembly seems more likely to remain secure, while syscall sandboxes seem more likely to fail-insecure if you get a detail wrong.
- It seems easier to sandbox otherwise-unmodified code that way. If you have code that needs some access to system resources, I think WebAssembly makes it easier to give it just what it needs and nothing else.
(Also, note that I'm not talking about running in a browser; I'm talking about standalone WebAssembly runtimes like wasmtime.)
https://gitlab.com/gitlab-org/gitlab-workhorse/-/commit/8656...
It's hard to find a linked detailed requirement for this. I would certainly prefer if GitLab didn't silently mangle uploaded images (not least if I'm working on an EXIF library..).
Bonus points for a commit that includes the words "perl" and "exec" not also having a detailed security review attached.
What you don't do is pulling an ages old perl codebase to run over complex formats.
What are alternatives to automatic immediate updates and kill switches besides not exposing the service to the internet?
The solution here is to offer a security notification service, which GitLab does. It's up to the admins to maintain these systems. It's GitLab's job to give them the information they need to do so, which they have.
I do occasionally get an email to say my account had been locked for security, but had assumed that was just noise from random login attempts, and the 2FA didn't make me look any further.
Was Gitlab was extracting the metadata and using it for some purpose. If not, what is the reason to accept images with metadata. Perhaps they assume their customers prefer less "security" and more "convenience", instead of vice versa (less "convenience", more "security").
That better security goes out the window once you start using eval(), of course.
This approach doesn't work in general. An attacker could craft a polyglot file - and in that case it's a matter of which format is tried first. Valid tiff's could potentially be processed as something entirely different.
https://github.com/angea/pocorgtfo/blob/master/contents/arti...
For more context see https://about.gitlab.com/blog/2021/11/04/action-needed-in-re... If you are using GitLab.com you are not affected.
The attacks on VOIP vendors mostly used UDP amplification, which relies on having a server that can fake its source IP due to an incompetent (or complicit!) network provider, while this is a botnet (that is only about a week old).
It it much easier to pinpoint any problem if you are aware the update happes and choose the time to do so.