In fact they should have added their own honeypot company names to the DB to force companies to parse robustly.
In fact they should have added their own honeypot company names to the DB to force companies to parse robustly.
The contents of this field link here: https://community.letsencrypt.org/t/adding-random-entries-to...
I think Let's Encrypt have the right idea. I honestly don't think that trying to tip-toe around poorly written code is generally the right thing to do; it seems more like the UK Government is prioritising short-term security (trying to block "bad data", whatever that even is) over long-term security (forcing people to write better code).
Only took a day or two of randomly shuffling around column orders on every write for them to see sense!
.gov should offer these detection services, and NSA should be providing an ambient baseline of pentesting.
Absent government action I think it’s a net-positive action though.
I find it harmful assuming that some externally-sourced data will match any arbitrary format (e.g. contain only allowed characters), even if it’s really supposed to be so. (Inverse for outputs - one has to conform as strictly as they can.) Ignoring this leads to mental dismissal of validation and correct handling, and that’s how things start to crack at the seams. I have seen too many examples of “this can never be… oops”.
Add: Best one can safely assume when handling a string is that it’ll be composed of a zero or more octets (because that’s what typically OS/language would guarantee). Languages and frameworks usually provide a lot of tooling to ensure things are what they expected to be. Ignoring the failure modes (even less probable ones, like a different Unicode collation than is conventional on a certain system) makes one sloppy, not practical.
We sanitise input all the time. This is not particularly unique. There isn't a great loss in this restriction of company names.
No we don't.
Companies like the aforementioned were made illegal because nobody sanitizes input.
SQL query injection and other forms of malformed data entry is still one of the most common attack vectors in the year 2024.
If everybody sanitizes their inputs (in undefined ways) then companies like the one mentioned would be randomly blocked from administrative processes.
This is not what we (as a society) want.
If Bobby Tables isn't a valid name the legislation should make it invalid, instead of rubber stamping it at the government registry and let poor Bobby get random errors when making requests to various public bodies. ("Sorry, our school does not admit persons with semicolons in their names.")
Is it astonishing? "Don't sanitize your own strings; always use a library" is common advice for handling SQL and HTML, which implies to me that it is in fact pretty hard to do correctly.
How about things like parsing strings for serializing to binary storage?
Can everything be an injection attack?
> Can everything be an injection attack?
What does this question even mean? I guess we must say "for any system accepting arbitrary input: yes". Not even sure if the "arbitrary" qualifier is necessary.
It never does, because abstractly speaking, there is no such thing as a secure computing system. This goes double for any computer that is switched on.
Practically speaking, it depends on how critical your application might be. If you're storing values for neurosurgery or automated dispersal of life-saving (or potentially life-ending) medication, you'd better be sanitizing on the way in, validating on the way out, and have some additional layers like audits and comparisons to known good values at rest. Look into defense in depth, and never trust the computer to make a decision, because the computer cannot be held accountable.
If you're storing quiz results for someone's favourite colour, or it's not internet connected, you can probably be a bit less paranoid about it.
> Can everything be an injection attack?
But yeah, anything and everything could be an injection attack if the attacker is determined enough. It's just a matter of how difficult you want to make it for them.
const csv = rows.map(cols => cols.join(','))
.join('\n')
because we are too lazy to write the more correct, const esc = cell => `"${String(cell).replace(/"/g, '""')}"`
const csv = rows.map(cols => cols.map(esc).join(','))
.join('\n')
(And perhaps something slightly more efficient but slower that only quotes each cell when it needs to be escaped.)I caught myself doing it the other day, Go has a JSON library and here I was too lazy to define a struct,
w.WriteHeader(500)
fmt.Fprintf(w, `{"error": %q}`, err.Error())
Is %q a JSON-compatible format? I have no idea without reading some source code! Almost certainly it won't \u-encode weird characters. That might be OK, I think the only stuff you really have to escape in JSON strings is newlines, backslashes, and double quotes? And %q probably handles those. Maybe it breaks on ASCII control characters...But yeah, we are meant to always use a library because we have deadlines and we are willing to compromise a whole lot of quality to deliver on them.
Specifically json and unjson I make globally available in all my projects. If I used csv more often than once in a decade, I’d have csvesc(s) too.
Sometimes you read some stdlib reference and wonder what they were thinking with things like System.out.println and without one-line one-arg readtext(), tojson(), fetch() and so on. It’s like a kitchen with all appliances still in boxes and all utensils in a tight vacuum cover. Everything is there, but preparation friction makes it absolutely unusable.
People think hard things should be easy and with less "friction". If I want to output a string why should I have to know what the difference between stdout and stderr is? If I write CSV to a file why do I need to know the difference between CRLF and LF, and UTF-8 and UTF-16 or what a BOM is? At the end of all of this you end up with a company named 'W""oopWoop;' crashing the banking industry.
So no, you should know all of that, and more or get the fuck out of my industry.
I think the high horse here is a bad point cause it simply claims it must be hard for no good reason. It’s not even complexity-wise hard, you just have to (metaphotically) unpack your instruments every time you use them. That’s bs at all experience levels and it must be obvious to anyone who works in a shop. Ime, the problem isn’t knowledge, but inconvenience.
What's astonishing is the popularity of the way of thinking that producing the cheapest code possible that still works along happy path (and simply doesn't fail too badly when it does) is is considered not only a valid practice but even some business virtue that needs to be protected.
The more I think about it, the more I like the idea of an EICAR-like records like this SCRIPT one - in the official database. It must be fully benign, of course (in a sense the script source should point to the same agency, and contain only a warning but no harmful code), and it must be well-known - effectively a test case for production systems. Rather than a pinky-swear "company name will should be okay, don't worry" that allows neglect, it's a "hey, this is a special weird case - specially to make sure you're doing things right" friendly guidance.
Are we still passing SQL statements and data to the SQL back end as single string instead of passing them separately? Why would you even need to escape SQL data in 2024?
[0] https://cheatsheetseries.owasp.org/cheatsheets/SQL_Injection...
Most are to do with ones which could be misleading, eg you can’t have ‘bank’ in the name unless you are, well, an actual bank.
The absolute best case scenario here is that the bureaucrats successfully block all possible actually-malicious injection attacks but the vulnerable consumers still get broken occasionally by a random apostrophe that gets thrown in.
It's WAY less pragmatic to test every company name for potential malicious actions in other peoples code that you don't own.
So you have a transitioning issue. You suddenly allow this company name sending a script to a domain they control then it is too dangerous.
Test data like you mentioned is a great idea to increase resiliance. However I don't think that rises the overall ecosystem of consumers of this data to the right level to release actual exploits into the dataset.
Downvoters are probably thinking purely. They are thinking "everyone in the world should make their systems 100% secure against common exploits and let a company name be an arbitrary string".
The problem is that is not realistic.
It works at a corporate level but not across all actors who interact with this dataset and the global internet. You can "should" at them all you like but no one has control over this.
The government can choose: more exploits in the wild or fewer. Allowing script URLs they dont control in company names is the former.
I think we can forgive the young William Gladstone (who was President of the Board of Trade at the time) for not fully anticipating how difficult robust string handling would turn out to be!
So you're right, this could only ever be approached as a transitioning issue.
Also, there can be a problem with who/how decides what is code. There are myriad of programming languages already, and for trolling or legal attack purposes, one could build interpreter using arbitrary words as keywords (to make problems for arbitrary company)
Blocking names that look like code is part of a defence in depth approach, it's not a standalone silver bullet.
Laws eventually are use not as intended, but as written.
“defense[1]”, “if happy begin something end”, “if”. All of these technically are code (somewhere). Also check out some esoteric language like: https://en.m.wikipedia.org/wiki/Whitespace_(programming_lang...
This is not how the real world runs though. In the real world (outside the bubble of programmers) things are messy and a lot of stuff barely works, many people are incompetent etc.
Said otherwise, it's defense in depth.
"Should" doesn't factor in. You can't make everyone competent at the wave of a magic wand. But you can control what company names are allowed. You can't control how they will be parsed. There is one law about company names, but a myriad systems that may parse them.
This is a huge blindspot of programmers.
This koolaid with protecting real world only helps perception (“I made it work now with this simple rule”), cause moving the bar down relaxes issues a bit and they don’t instantly accumulate at the new level.
It doesn’t matter where the bar is, they will always find enough competence and budget to follow it in a moment. You just have to hard-break what half-works in advance.
You can't make everyone competent at the wave of a magic wand
You can make their incompetence fail by adding random honeypots like someone suggested above. That would be a smart move. Your “out of bubble” move is just an instant gratification button.
My point would be, I'm not sure if this wouldn't be too damaging to the mental health of programmers if everyone was doing shit like that.
Not executing user input strings?
IMO, this is like making human names illegal because people with certain accents or native languages may struggle to pronounce them.
Our government officials are so stupid it's astounding. This doesn't make anybody safer, but there's now another minor charge after somebody has broken the law.
> “A company was registered using characters that could have presented a security risk to a small number of our customers, if published on unprotected external websites.”
Emphasis mine.
Maybe you’re the stupid one?
Be right back, gonna rename my company real quick
Company names are not a game of hack-a-mouse. You think you're being smart, you're just being another annoying Ackshually guy
They are names that should be useable across many systems and use cases.
Let's say the UK registry fixes their systems, but now you need to have your company name across other suppliers/vendors systems. Congrats, you played yourself
We are grown ups, we can disagree without resorting to ad homenim. (Might be time for you to review the HN code of conduct.)
/s because sadly I feel it is needed here.
- somehow ensure all software is bug free (at least when processing company names)
- outlawing things
- just let it happen
The first option isn't that far away from hacking the matrix and making buggy software physically impossible. The second option seems to be better than the third.
That's actually a really good point.
Nex the UK will ban knives. Oh wait...
Do you have bars on your windows? No? The potential cost is a breakin?
You you expect restaurants and stores to pat you down before you are allowed to enter? No? The potential cost is an attack on the staff.
Should we ban cars because they can be used as lethal weapons? No? The potential cost is a terrorist attack.
Deterrence through consequence is a thing and generally less costly for society than to make crime 100% impossible.
Absolutely. In exchange, however, I get better visibility, and lower cost windows.
Those advantages are meaningful enough that my house does not have bars on the windows.
The law does not prevent attacks it lowers cost of prosecution by clearing up the ambiguity about whether this was illegal.
I'm not sure I love that, but that's how it always seems to work. Otherwise it's just another "job killing regulation".