Username ending with MIME type format is not allowed
gitlab.com
gitlab.com
Examples of MIME types: "text/plain", "text/html", "image/png" "application/pdf", "video/quicktime", ...
If I was prevented from using the username "wcoenentext/html", then I wouldn't really be bothered by that. (Although I might question the design decisions that would necessitate such a restriction.)
I’m the author of the issue on Gitlab (small world, isn’t it ?)
Yes the message is confusing and I agree that .mov isn’t a MIME type but I was merely reporting the error message shown ( plus, they added .mov in their list of file types and had aliased it to .mp4 format, please see: https://gitlab.com/gitlab-org/gitlab/-/blob/master/config/in... )
That’s weird. Why’d they do that. They should make a separate entry for mov and associate it with video/quicktime
Guess it might be something related to https://stackoverflow.com/a/44785870 but like they point out, mov is a container format that can contain one of many different codecs used. And isn’t mp4 just a container too? Referring to mov files as video/mp4 seems straight up incorrect to me
If you look at a .mov file and a .mp4 file in a ISO bmff viewer, you'll generally see the only difference is the ftyp box is different ("qt " for .mov, "isom" for .mp4). Indeed, if you ask ffmpeg to make a .mov file and .mp4 file of the same content, literally the only difference is the contents of the "ftyp" box, every other byte is identical.
Not quite. There are some boxes acceptable in one but not the other. Strings inside MOOV are length-prefixed in MOV but null-terminated in ISOBMFF. There are a variety of differences like that.
The set of codecs allowed, also differs between the two.
Thanks for the suggestion :)
"""the name would have violated California regulations as it contained characters that are not in the modern English alphabet,[321][322] and was then changed to "X Æ A-Xii". This drew more confusion, as Æ is not a letter in the modern English alphabet.[323] The child was eventually named "X AE A-XII", with "X" as a first name and "AE A-XII" as a middle name.[324]"""
It sounds as though they allowed the middle name to contain a space, which doesn't match my mental model but perhaps that's how they do it in California. It invites the question: do they allow a first name to contain a space, so that ("X", "AE A-XII", "Last Name") and ("X AE", "A-XII", "Last Name") are different names?
EDIT: The funny thing isn't allowing spaces in a middle name; it's have a separate field for middle name(s) at all. Modern passports have just two name fields: "surname" and "given names". Both may contain spaces. But according to images on the web, Californian birth certificates really do have three name fields.
Source: I have a space in my name and some of my different identity documents have the name as "<A> <B C>" or "<A B> C", which causes all sorts of administrative problems.
Nowadays the meanings of our last names have largely disappeared, so you have countless Coopers who have never touched a barrel in their lives, whose children will be called Cooper also, despite that. I think it's a little sad that so much of what people call us is semantically equivalent to a random UUID with tons of namespace collision. With that in mind, I'd say "da Vinci" is more a last name than most of us have.
I can understand that in some cultures this might seem weird or antiquated but here in Germany these names are reality. Sometimes people with such names are descendents of royalty and sometimes someones last name "from family-name" is thier last name and happens to historically correspond to one of germany's state names or city names or just a little town.
One a side note: In Germany in 1919-1920 royalty was no longer a legal aspect that changed how laws applied to you[1]. When that happened titles that were reserved for ruling functions (king, grand duke) were removed and all other titles were moved to be part of the persons name (such as prince etc.) and could not be decreed on anyone new. These titles still exist in Germany but are simply naming "conventions" in a formerly royal family.
[1] https://de.wikipedia.org/wiki/Adelsrecht
[edit] In Leonardo's case perhaps not but still i wish to elaborate a litte on the situation here.
Seriously, though, it just feels apropos.
A UK passport has a "surname" field and a "given names" field. But a recent UK birth certificate has a single name field in which the "surname" part is distinguished by being written in capitals, so "Peter James ADAM SMITH" would have a surname consisting of two words. But what if a word consists of a single letter? For example, some Irish surnames look like "O Briain", and I think there is a Vietnamese name that consists of a single vowel, so presumably you can't always tell from the birth certificate which part is the surname.
The legal name however consists of First and Family names only, without the other names. Therefore many people have two names in the First Name field, usually separated by a space. The disadvantage is that in all official forms they have to spell out all the first name(s), no omissions or abbreviations are generally admitted.
Typical Portuguese names have around 5 names, two first names and three surnames from both parents.
In fact your description fits quite well how they are used in Portugal, plus a few other nuances.
So, Heinrich August Schmidt has a last name / family name (Schmidt) and two first names / given names.
*"you are" meaning you'd be a German citizen with a German passport.
When you marry, you get to define what is going to be the family name and then children born from that marriage get to be registered with that family name.
In the case of foreign-born people, they can keep the naming rules from their original country. In the case of the marriage between foreigners from two different countries, you have to choose which rules are you going to follow, but the family name stays fixed.
To us (Brazilian marrying a Greek) it was a very interesting process. I have two last family names, and Greek names are gender-conjugated (i.e, the last name changes whether you are a boy or a girl). It the end the simplest thing to do was to just keep only one my last family names.
This is possible because you can make a “name declaration” where you choose to apply the naming law of another EU country.
I have two middle names; most financial service providers decline to recognise my second middle-name. Same goes for the taxman and my pension provider. It seems that a "full name" is no longer a canonical identifier for a person, just really a kind of nickname; the canonical identifier is now an account-number, employee-ID or whatever.
(Although probably a lot more than I know about.)
Someone tell the Encyclopædia Britannica.
Or that æ is a character formed by combining two separate letters?
Both support the idea that Æ is not a letter in the modern English alphabet.
Unless Musk tries to name future kids in Emoji, European nations can probably already cope with any of this sort of thing.
Good question about the partner; I guess that a narcissist can pair well with a very passive person (or maybe another narcissist).
It's probably a good idea if government computers can keep track of the name, so a minimum requirement would probably be "can be represented by unicode glyphs". So attempts like Prince's should probably be disqualified. And even in the unicode set there are characters that may be problematic - Record Separators, Zero Width Space, Pile of Poo emoji springs to mind as examples, even if the later one might be doable. It's hard to address a letter to a person whose name just consists of a mix of different whitespace characters (especially when their neighbor is a different mix whitespace characters). So, yes, government probably should care about this, at least a little bit.
That said, Æ should probably be allowed. It's a standard character in several living languages.
> That said, Æ should probably be allowed
This may cause all sorts of issues, from typing a letter to issuing a passport.
Else than that, exactly what I was thinking.
And practically, in a world where people cross borders, where people come to this country for refuge and opportunity, how does it make sense to force them all to have only basic English characters in their names?
You’ve never interacted with the government or a large corporation, have you?
But in seriousness: it is their concern because changes that seem minor can require major changes to electronic systems, and allowing special cases that necessitate lookup in ways the electronic systems can't handle (perhaps even requring manual search through hardcopy!) really throw a star-mangled spanner in the works.
Naming is already a special case problem in Japan, where people are accustomed to having no idea how to pronounce most people's names because the parents used non-standard readings of the kanji characters and/or used kanji from a special exempted list of archaic kanji that almost nobody can read. If you're wondering why the government doesn't require people use only the official "common use" kanji, at least just for this one problem, then I should tell it has been tried — enough people raised a fuss about being unable to register their child's "perfect name" that gradually a list of "allowed only in names" kanji was created and expanded.
And if I recall correctly, when registering a name, you can specify a totally unrelated pronunciation using the simpler non-kanji phonetic characters, so even just "common use" kanji are almost "all bets off". A relative few kanji have such common pronunciation in names that they can 'usually' be guessed.
And the problem is the same with many Japanese place-names, having little or no correlation between written form and pronunciation or meaning.
So, why is the spelling of a chosen name any concern to a government? It gets crazy out there, in Name Land. How mäný variatǐons of spelliñg cån ße ællowèd before people give up on pronouncing it?
What's fascinating to me is that none of you consider that government might not need to have a list of everyone it governs, or that such a list might not need to be centralized or computerized. When faced with a facet of humanity that's too complex to be easily reduced to consistent data, private groups either do their best and work with what data they can extract, or else have people handle it, with our flexible, tolerant minds. When government is involved, though, the immediate response is, "We have to force people to be less complex!"
I'm beginning to see technocracy as the biggest threat to a diverse, human, and humane society.
There are quite specific rules:
The new name must be in the format {substring of original name 1}{?-}{substring of original name 2}, or you have to go through the much more arduous and expensive full name change process (though you get a 2 for 1 discount).
That specious, meaning appearing true but actually false. It is a diphthong expressed by a ligature of two letters. The claim was never made that a diphthong is a single letter, and it is easily expressed in modern English, in a computer using ascii and outside on a piece of paper.
In English, I'm pretty sure the ae digraph is always pronounced as /i/, so calling it a dipthong would probably confuse most people.
[1] merriam-webster.com/dictionary/ligature
Edited to add: A "famous" wtf-worthy explainer from Oracle: https://asktom.oracle.com/pls/apex/f?p=100:11:0::::P11_QUEST...
Relevant snippet:
'' when assigned to a char(1) becomes ' ' (char types are blank padded strings).
'' when assigned to a varchar2(1) becomes '' which is a zero length string and a zero length string is NULL in Oracle (it is no long '')https://www.latimes.com/archives/la-xpm-1986-06-23-vw-20054-...
https://www.facebook.com/public/James-Null
as well as Abcde
Neither part of the MIME type format even has to match any existing (commonly used) file name extension anyway. E.g. it can be `text/plain`. Even when it does it is just a coincidence (although very common), it actually references the format name (IIRC `image/jpeg` was used even when almost nobody were using `jpeg` for the extension and the convention was to use `jpg`).
At least part of that is not true. The ‘popular’ MS-DOS 8.3 form derives via CP/M from DEC OSes which used 6 character file names and 3 character file types, due to their use of RAD50¹ to fit 3 characters in a 16- or 18-bit word. The type field always existed, so the word ‘extension’ most likely refers to its presentation on the end of the file name, rather than an addition to a previous format.
Never thought about it but file extension name really is a word. Someone replied saying this is three words, but it is not is it? It's an open compound word or maybe a "set phrase", I wanted to call it an idiomatic expression but that was clearly wrong.
Not helped by the fact that even with explicit cajoling* to give steps to reproduce, the reporter wrote:
> Steps to reproduce
>Try to login/create an user in Gitlab (on-premises/Gitlab.com) where the username ends with a MIME type format
Goddamit, those are not steps to reproduce! Has the current era of "social coding" and its terrible software development practices completely turned people's brains to mush?
(In fact, such cajoling shouldn't even be required. If you don't know you need to provide STR without someone going to the lengths of coddling you by going out of their way to create a template, you don't have any business using a bug tracker.)
The other person is, I think, trying to express that it can potentially be done more than one way, so additional steps are required.
From a QA perspective, I prefer not to guess what the reporter might have intended. It's much better to have tons of detail, but I can sympathise with being the user and _thinking_ you have been totally clear and am know I've sometimes done this myself.
(( Unrelatedly, "ok boomer"? Since that seems such a non sequitur, I'll take in another unrelated direction and raise you an "ok athena" and "waddaya hear, starkuck?". ))
> From a QA perspective, I prefer not to guess what the reporter might have intended
There's not really any other perspective. That thing which you say you "prefer" is much more than a preference. It is the only reason why the bug tracker is configured to even ask for STR. If you have a "no solicitors" sign on your front door and a solicitor comes knocking anyway, or you have your door locked and a burglar climbs through a broken window and and takes all your stuff, you wouldn't respond by saying, "from my perspective, it's really disruptive and time-consuming to deal with solicitors who ignore the sign" or, "try to understand that from my perspective, it's really inconvenient when you take my things, because I have to work and spend money to replace them when I could have used that time and money doing something else." To do so tacitly legitimizes an illegitimate position held by the other.
Please look up "DARVO" to understand why explaining yourself like this is a bad idea.
I initially suspected that a regex was involved and someone forgot to escape the period, but it looks like that wasn't even the case -- the erroneous code was literally checking if the username ended in any recognized extension.
https://gitlab.com/gitlab-org/gitlab/-/merge_requests/65954/...
fixed it!
I'd rather a technically oriented site like github be "unrealistic" and correct and instructive, than foolish and wrong and misleading. Are you really saying it's better to use the wrong term because github users might not understand the more widely known correct term?
Just how does leading the user on a wild goose chase looking up the definition of "MIME Type," causing them to waste their time and misunderstand the error message, when it's really a file name extension (a term which more people understand anyway), help the user achieve their goals?
The bottom line is that github disallowing "MIME Types" or "file name extensions" in user names, just like a bank disallowing "select" and "drop" and "from" and "null" and "delete" and "bobby" and "tables" in passwords, is a symptom of a much larger more terrible problem, and whoever wrote that stupid error message instead of fixing the underlying bug that caused it has much worse problems than poor English language writing skills.
The bug affected Gitlab, not GitHub.
It's a very odd error. Apparently .nro is a file extension used by the Nintendo Switch video game console; .o is obviously the output of compilers, so I'm not sure why my username wasn't rejected. Maybe it would be if I tried to register now.
It seems pretty obvious - based on the failed username in question, and to someone with fairly deep technical knowledge - that they mean anything they consider a file extension. Which is not an excuse for this marvel of awful UI slapped on top of a poorly-thought-out workaround (for some unknown vuln (that's been patched for over 2 months and is still private? quite strange for an "open" company eh?).
Edit: as for what's actually used, looks like [2] it's the ruby mime-types gem [3], which is based on both IANA registry and various other recommendations [4], AIUI.
[1] https://www.iana.org/assignments/media-types/media-types.xht...
[2] https://gitlab.com/gitlab-org/gitlab/-/merge_requests/65954/...
Registrants can include info like file extensions with their registration, but those extensions are not registered (not by IANA, anyway).
And indeed, what IANA registers there is media types, so I've only mentioned associations. I think it still works for a-dub's argument, that it's obvious (though I'd word it differently, such as it being a reasonable guess) that if "MIME type format" is mentioned and "mov" is shown as an example, what's actually meant is filename extensions associated with registered and/or otherwise known media types (which turned out to be the case).
[1] https://www.iana.org/assignments/media-types/application/xml
[2] https://www.rfc-editor.org/rfc/rfc6838.html#section-4.12
I don't know exactly what list GitLab is using (because, as has been mentioned, there's no central registry of file extensions), but there are many extensions more obscure and less obvious than even mov, which would further obfuscate the actual problem.
Yes. I am well aware that certain software, such as Windows, MimeMagic, or Apache, do include their own lists. That is not a "registry" and in fact you will likely find that every such list is either identical to, a fork of, or incompatible with every other such list.
Whether you call such a list a registry or not is fairly trivial, I think. Their web server uses such a mapping, and for some reason usernames with suffixes on that list confused something. Call it a registry, or a mapping, or a list, the meaning of the statement is the same, and I don't think it prevents most people from understanding what's going on.
In this case, a registry would be a mapping, however a mapping is not a registry. A registry is the single, centralized place from which you can look up registrations, to guarantee that there are no conflicting records. Literally none of that is (or can be, at this point) true of file extensions.
Which is definitively not what is meant there. But I think it shows that Gitlab is not a company I'd go to if I wanted linguistic precision.
So, I find strange this meaning that you give for a "draft" = that it is something that has connotation of throwing away when building the final product.
I think you are confusing a "draft" with a "sketch" and they are not the same.
I googled the term "draft" and here is what I found:
> "a version of something (such as a document) that you make before you make the final version" [1]
> "A preliminary version of a piece of writing." [2]
While "sketch" means:
> "a rough drawing representing the chief features of an object or scene and often made as a preliminary study" [3]
> "A rough or unfinished version of any creative work." [3]
> "A rough or unfinished drawing or painting, often made to assist in making a more finished picture." [4]
In the case of sketch I see some keywords like "preliminary study" or "assist" or "unfinished version" that indicates that the sketch will not be the final product.
So while it is true that a Draft could be thrown away if someone has new/better/difference ideas while working on it, it does not seem to imply that a Draft should be thrown when building the final product.
As far as I understand it is more that a draft will evolve into a final product or might be abandoned.
[1] https://www.merriam-webster.com/dictionary/draft
[2] https://www.lexico.com/definition/draft
As a native English speaker, the first time I encountered the word 'draft' was at primary school. We used it to describe a piece of writing where presentation was not the focus, instead, content and accuracy in terms of spelling, grammar, and punctuation would be the focus. Once the drafts were complete, we would 'copy these up' in our neatest handwriting.
However, if I prepare a 'draft' of some document or other for my boss, I expect it to be essentially an unapproved version of a final document, perhaps needing some minor modifications before release, but also perhaps not.
Although personally, I would use 'sketch' to mean something disposable which illustrates a more perfect version, my grandmother is an artist, and she refers to the initial drawings she makes on the canvas as 'sketches', which she then paints over in more detail.
Essentially, I don't think there's a big difference between draft and sketch - both could (in my opinion) represent either a version which will be discarded or which will be developed further.
Overall, I think here, the 'Work In Progress' label is the clearest and least likely to be interpreted differently by different users.
In a highly technical environment this is the kind of work I associate with drafting, but not necessarily with the word draft.
The wikipedia page for "Draft" has a veritable smorgasbord of different things that are considered "Drafts"[2], which kind of illustrates that getting pedantic about the definition of the word is losing proposition.
But a native english speakers as well, I agree with you, that the everyday colloquial definition of the word "draft" is an incomplete piece of work that needs further refinement before it can be considered complete or final.
[1] https://en.wikipedia.org/wiki/Drafter [2] https://en.wikipedia.org/wiki/Draft
That is it for me too and why I think this change is so annoying.
WIP means something is being worked on. [0]
Draft can mean the same, but it also has a bunch of other possible meanings. Note that if you search in both of your sources, you'll find "sketch" as an explaination. It can mean the same, but it can also mean a bunch of other things. It's a strictly worse name.
A draft in the visual arts (drawing, painting), is typically abandoned, like a sketch. Perhaps it uses a different medium to the final version, and in any case can't easily be adapted.
A draft in writing is typically incrementally improved, or at least we think of it that way now that we write on computers. We don't need to start again even for substantial changes like adding a new paragraph.
Draught (sounds the same as draft) is the wind through a crack or a type of beer. A door or window might be draughty but a beer wont.
Drought looks similar but sounds like "drowt" and is what you get when there is a long period of time without rain.
Be careful with draft, draught and drought!
I think it's only software developers who urge others to first write a draft, then throw it away and start working on the real thing. Usually, a draft precedes the "real" version and it's a status attached to something. Eventually, a "draft" becomes "published" or something similar.
Authors, scientists, email writes, report creators, movie/music producers all create drafts that (maybe) eventually become the real thing, I don't think many of them throw away the draft but rather work on the draft until it's not a draft anymore.
Interesting. What industry does it that way?
I guess I'm dense, but I actually thought it's about users such as Mr. Joe R Text/Plain. When I read about "mov" in the actual issue, then it became clear.
They keep on having high-severity security bugs being fixed every month (e.g. auth checks not being done everywhere). Then there's all these odd edge case bugs everywhere.
As an outsider, it just seems to me that GitLab isn't being engineered in a principled way: on sound abstractions and with separation of concerns (e.g. auth should be some universal middleware, not ad-hoc per call). Just really basic stuff.
Due to a security concern in which a profile containing a file extension would not load [3], we do not allow usernames that end with file extensions (ex: .mov). As noted by many folks here, these are associated with a MIME type but are not MIME types themselves. It is not related to preventing an injection or any such attack vector.
The error message for this check incorrectly included MIME type rather than file extension. This has been updated [4].
Additionally, there was an issue with how the actual check as it did not include the leading dot. The leading dot was added to the check in a subsequent MR [5].
Thanks for all the feedback.
1 - https://news.ycombinator.com/item?id=28535739
2 - https://news.ycombinator.com/item?id=28538166
3 - https://gitlab.com/gitlab-org/gitlab/-/issues/26295
4 - https://gitlab.com/gitlab-org/gitlab/-/merge_requests/70374/...
5 - https://gitlab.com/gitlab-org/gitlab/-/merge_requests/65954
Why did you not fix your routing engine to not consider file extensions where the username / group name should go?
Besides blacklisting certain usernames or breaking a bunch links to profiles, how would you fix that?
EDIT: IIRC, github has the same issue, but they have profiles as lower priority. So if your username conflicts with an existing URL, your profile page doesn't work.
1. deny-list only usernames that are actually existing conflicts
2. Change the URL for only usernames that have conflicts, to `https://gitlab.com/u/<username>`.
3. Change the URL for _all_ usernames to `gitlab.com/u/<username>` as this collision points out the flaw in the original URL design in the first place, because of possible collisions. 301 redirects could of course be used for any non-colliding usernames.
I am now wondering how _github_ takes care of it though. Github also has `github.com/<username>` urls. What does it do with collisions? Github pages don't even all end in `.html` or contain a `.` at all, so gitlab's particular solution would not work. For instance, there is a page `https://github.com/topics`. What happens if you try to create a github user called "topics"?
If I try to create one, it says "Username 'topics' is unavailable." Same for say `marketplace` or `trending`. Perhaps they've deny-listed only actually-existing github urls? That does seem tricky, whenever they want to create a new top-level /page on github, they can only do it if there isn't already a github account with that name?
But if as someone else says `/dashboard.html` is just a weird non-canonical alternate for `/dashboard`, which already had to be reserved, maybe gitlab is already doing (1) anyway? Then why do they need to also deny any username with ending in any valid extension? Unclear.
It still makes me wonder if they have a routing precedence problem, which they worked around by just forbidding any username that triggered it, instead of fixing the actual issue.
Hopefully we'll know more once their security ticket becomes public.
My comment, and the issue that was submitted here, is about *.mov
https://myservice.com/my.super.duper.crazy-ass.user.name.pdf.exe
and still have it return a proper HTML document that covers that user's profile page. Hell, the username could be some insane zalgo-tier shit and still function properly.I see some comments defending arbitrary "bandaid" architecture and I think that this is not defensible for something the scale of GitLab. This is basic HTTP stuff.
Examples:
https://gitlab.com/gitlab-org/gitlab/-/merge_requests/70427....
For an old post on a somewhat related topic, see:
https://ryanbigg.com/2009/04/how-rails-works-2-mime-types-re...
I could imagine the mix of rails #respond_to and "file extensions" at the end of urls might make a mess (think /users/profile/smith.html vs /smith vs /smith.json vs /smith.txt - essentially what might have been /smith?format=json etc).
Ed: current documentation: https://apidock.com/rails/v6.1.3.1/ActionController/MimeResp...
Note to self: Never say Ruby when you mean Rails on HN
https://news.ycombinator.com/item?id=28540665 has all URLs and issues available to get started in the code.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Co...
As for actually performing this change myself, it would take me quite some time to grok the codebase.
Maybe I get bored tonight and see how hard this would be to resolve.
Please, for the love of your own mothers, people: stop pretending like you know how to parse urls like they're strings. You don't. And even if you can, you won't do it right every time.
It's been demonstrated over and over again.
Just use the already tested url facilities to give you the host/path/query parameters
Still, what is the attack vector?
Not sure if it's an attack vector per se, or just that the behaviour is incompatible with allowing usernames containing . and then having urls where the username is the last segment of the url
seems like a badly designed url scheme :)
This has always disturbed me, considering that HTTP has had content negotiation for ... oh, basically its entire history [https://www.w3.org/Protocols/HTTP/1.0/spec.html#Accept].
On a similar topic HIBP allowed people to request versioning via a custom HTTP Header, a Accept Content-Type, or a version segment in the URL path and approximately everyone went with option 3.
Take my `Accept: text/html` and give me my HTML-ified JSON, dangit! ;)
A variation on the Scunthorpe Problem[0] then, eh?
It wasn't a regex, they just did a generic "ends with" check.
'The problem was named after an incident in 1996 in which AOL's profanity filter prevented residents of the town of Scunthorpe, Lincolnshire, England, from creating accounts with AOL, because the town's name contains the substring "cunt".'
Right. Regardless of the specific pattern matching function, in both cases, the results were both incorrect and unwanted. Which is why I consider this instance to be a variation on the same issue.
- The merge request which originally added this check is inaccessible (https://gitlab.com/gitlab-org/security/gitlab/-/merge_reques...)
- In the issue comments the Gitlab employee says "Sorry, I cannot go into details right now. I will link the issue here once it goes public, is it ok?"
I'm not sure how one could exploit it though...
[1]: https://docs.microsoft.com/en-us/previous-versions/windows/i...
Update: tested this link in FireFox 92 - it still performs sniffing in 2021: http://www.debugtheweb.com/test/mime/script.asp (based on the content, not extension)
1) web servers, browsers, proxies 2) graphical os shells 3) email
every file a webserver returns has a mime type in the header, and that is how the browser knows how to present it.
But even there is a header for that, there will still be situations that clients don't care and guess it themselves (probably because software is too old or other reason).
The story is that in some cases, Internet Explorer could be tricked to ignore the Content-Type header. Instead of fixing the core bug, Microsoft decided to add a header (presumably, they didn’t want to break existing websites). A decade later, people discovered two ways to bypass X-Content-Type-Options. This time, Microsoft fixed the core issue instead of adding X-Content-Type-Options-For-Reals.
FWIW background https://www.youtube.com/watch?v=8t8JYpt0egE
I think even browser nowadays still do it. A file in <script src="" /> will be treated as script even mime is wrong. A file ended with .mp4 will be play as video even it says it is a text/plain file. Browser guess mime from contents at many places, just you may not notice it.
unless XCTO is set
https://gitlab.com/gitlab-org/gitlab/-/merge_requests/65954/...
Slightly tangentially reminds me of the "More Magic" switch of GLS fame.
Somewhere in the Gitlab code base there is a MIME_TYPES map with common extensions as the map key. No idea what it is used for but that module is very likely the target of a recent security issue.
The first fix to combat the "publicly unknown" vulnerability was to prevent usernames ending with any of the keys in the MIME_TYPES map using a simple "ends_with" strings check. Of course the map keys did not have periods so the ends_with would also match "Asimov" with the "mov" suffix.
The second fix in this PR is to extend the ends_with check to add an extra dot.
The actual vulnerability is still unknown but I suspect it's something like an intermediate component that performs special handling based on interpreting URLs and that could bypass security/ACL checks.
Can you elaborate on what you're referencing?
Tried googling but couldn't find anything.
Of course, being a teenager, the term often came with a raised eyebrow from me since it was commonly used as a slur for lesbians. He was the only person I've encountered in my life who used that term in that manner and over the years I just assumed it was some bit of obscurity related to where he started working in systems. I guess not!
[0] I'm sure the right Gooble query would have gotten me an answer but it teetered on the edge of "I don't care enough to bother" until the answer was presented.
Rails has a special track record for convenient magic implicated in security vulnerabilities.
Another commenter gave a good example theory implicating a convenient-but-questionable out of the box behavior of Rails: https://news.ycombinator.com/item?id=28537562
The good news is Rails has been slowly moving away from the magic over the years - it used to be a lot worse.
This does really look like a "too much magic" situation.
Maybe just playing it safe?
Good times ..... (rocking in corner ....)
- There was a yet-undisclosed security vulnerability in Gitlab usernames
- Staff member made a change to disallow usernames ending with `Mime::EXTENSION_LOOKUP.keys`, which I assume is a set of recognized file extensions (hidden – https://gitlab.com/gitlab-org/security/gitlab/-/merge_reques...)
- This was overly broad since it caught a lot of common names (like "asimov") (https://gitlab.com/gitlab-org/gitlab/-/issues/335278)
- The check was updated to additionally look for a "." before the extension (https://gitlab.com/gitlab-org/gitlab/-/merge_requests/65954)
This is a hell of a thing to commit and pass code review in 2021 on a project like GitLab. I understand that the staff member was fixing a security issue and was probably not thinking deeply about the ramifications, but even so. How many "Falsehoods programmers believe about names" articles do we need?
The bigger question is how this passed code review and testing.
Use unique prefixes for users, groups, projects, assets in your webapp design, kids.
So I'm probably missing something and I'm really curious for the underlying vulnerability.
Something like "the username can be part of the URL, and if the URL contains .mov, some browsers will misinterpret this and assume it's a movie file, leading to bad things™".
Or: "the username is sometimes used as a folder name, and our syncing software contains rules to exclude certain file extensions, so these folders were never synced, which lead to issues on production servers"
I'm guessing it's something along these lines. Something that you control, but not really, leading to these kind of haphazard workarounds.
If only github had an established system and procedure for doing code reviews before releasing security fixes...
That they don't seem to understand what a MIME type is just adds to this perception.
Github has the exact same issue, and they solved it by just restricting usernames to alphanumeric characters and hyphens (but IIRC, there are some existing profiles with unreachable URLs from before they made the change).
Gitlab decided to just restrict known file extensions, which IMO is not the best idea, but it's also not that unreasonable.
> That they don't seem to understand what a MIME type is just adds to this perception.
That's not fair, people misspeak/mistype things all the time.
This change absolutely seems like the wrong place to fix any real security vulnerability and the fact that it affected a bunch of legitimate usernames is the icing on the cake.
This comment https://news.ycombinator.com/item?id=28540665 helps with more content and issue URLs including the problem discussion.
There is no reason for this. They should always hit the same controller and be served as text/html. Why would the username ever influence this?
If that breaks an existing page then you either shouldn't have allowed the username to be created with the same name as an existing page (exactly what GitHub does) or shouldn't have created a page with the same name as a username [depending on which came first!] - or better yet, had them in a separate namespace so they cannot conflict in the first place.
I don't think browsers infer meaning about file types based on the URL. The Content-Type is always what is being used.
If you have backend side code that maps URLs Content-Type header mime types. Don't. Instead simply always return text/html for user profiles. Then the extension shouldn't matter.
— (July 14) https://gitlab.com/gitlab-org/gitlab/-/issues/335278#note_62...
Has anyone traced down an explanation of the (presumably security-vuln-related) thing that was going on, motivating the original restriction?
I would think it would be public now, with the restriction removed, but maybe not?
It does seem like an unfortunate UX to exclude the lastnames of millions of people as usernames.
just a regular .ends_with()
What is the reason for this?
The proper fix is to disable this mechanism at least for the username segment of gitlab path but perhaps GitLab developers are too lazy or unaware or just in rush.
I mean, barring some tricks that could break layouts during rendering if people got really creative with them, why isn't any valid non-blank Unicode string a valid username?
Reminds me of this xkcd:
it's because HTTP is stringly-typed
"Did you really name your son Robert'); DROP TABLE Students;-- ?"
"Oh, yes. Little Bobby Tables, we call him."
Basically something to extract boatloads of money from enterprise customers by annoying THEIR customers so they can't write "<script>" in texts in their application.
Or, less tongue-in-cheek, a way to harden web applications against known attack patterns like sql injections or xss-attacks. As they work on pattern recognition and don't know anything about your application they sometimes get in the way. But they'll probably check some box for some security audit so they're used.
Cloudflare for example offers one at https://www.cloudflare.com/waf/
A WAF is useful for when a zero-day is found for that legacy application you just can't get patches for anymore because the team has "moved on".