Secure by Design
henrikwarne.com
henrikwarne.com
Size. A payload of one million characters should probably
be rejected without further analysis. As well as checking
the total size, it is good to check the sizes of the parts.
Another thing to check, that is often overlooked, is Quantity.Every loop should have a limit. There should be no unbounded object allocation. Usually, it's not the big things that get you but the sheer number of small things.
For example, a 20 MB email with 4 million empty attachments:
https://snyk.io/blog/how-to-crash-an-email-server-with-a-sin...
Further examples that affected ClamAV and SpamAssassin:
https://blog.clamav.net/2019/11/clamav-01021-and-01015-patch...
http://mail-archives.apache.org/mod_mbox/spamassassin-announ...
NB I don't know what the answer to this is, but pretty much any time a system I have been involved with contains a "reasonable" limit then people want more than that pretty quickly!
In fact, I think the problem is less relevant to a Quantity check, where an order of magnitude headroom above normal usage goes a long way, more so than a Size check.
For example, do you know of people who receive 10,000 attachments per email? This limit would be far and away above any reasonable usage and yet provide decent protection at the same time.
Even defining "normal use" is intractable. For instance, most docker layers are a few MB, but some people are deploying 3P software packaged as a container with 10 GB in a single layer. You can't fix their container. They can't fix their container. Your definition of reasonable changes, and you bump your maximum to 1 TB. Then someone is trying to deploy docker containers that run VMs, which have 1.5 TB images. It's to interface with legacy systems that are infeasibly difficult to improve. But the vhd is a single sile, so now you have a single layer maximum size of 1.5 TB. But since the 10 GB body size is a possible attack vector in and of itself, what's the security benefit of having any maximum size limit at this point?
It's the wrong approach. Instead, your system should gracefully handle objects of arbitrary size. Security should be enforced by cryptographically enforced access controls and quotas.
e.g. When I thought of a "reasonable" maximum number of attachments on an email I'd probably say 64.
We should not be talking of a "reasonable" maximum at all.
Rather, the maximum should be orders of magnitude away from the "reasonable".
At this distance, any arbitrariness quickly fades into meaninglessness.
Detached from the "reasonable" usage, the maximum then becomes informed by processing cost and complexity analysis, which are more concrete.
I think that may not always be true though. Especially for highly variable workloads.
How many requests per seconds does your site handle on average vs how many can it handle? What about spikes?
If your previous peak is at 10 rps and you can probably handle 20 rps, when should you start dropping connections to keep the site up?
Most systems are designed to have small margins, for cost purposes, so I imagine this sort of situation might come up quite a bit.
The correct response to a breach of this limit would be for a reply email to explain the limit and why the email was rejected.
The user then could send multiple emails with their attachments, and the system can be sized to handle e.g. 24 threads processing 64 attachments of a maximum size of e.g. 4096kb
That's how you ensure a system sized to your hardware and ensure maximum throughput for all users.
Sure, I think we are actually in agreement. That's exactly what I meant by "processing cost and complexity analysis".
Some sort of lazy or other on-demand evaluation can help a lot here. You set the limits large, and then only evaluate the part of it that's actively being accessed.
Would an answer be to have limits that are tunable? I.e., something like "Attachment limit exceeded. Enter new temporary attachment limit for re-scanning: _____"?
Is this data available? One simple one: How much long must be a name in a database field?
I’ve recently created some mini-languages (DSL’s) for everyone users to control various software and I’ve found that using a synchronous transformational language[1], which has no unbounded loops (or anything) and is therefore guaranteed to complete, and which runs in a stop-the-world fashion (ie each invocation acts as if atomic, whether it actually is or if implementation tricks like transaction rollback + retries are used) makes it easier both to develop and for end users especially ones who aren’t programmers. I’m a huge fan of this approach for correct and secure software.
[1] https://en.m.wikipedia.org/wiki/Synchronous_programming_lang...
I would hope that was from auto generated code somehow..?!!
Got a feature that lets users add multiple additional options? Someone will use it to add 10,000 options, and now pages which display them in a <select> widget will slow to a crawl.
There should be a law named for this.
What about Limit's Law?
See also total functional programming, which applies the same logic to runtime.
I love when people write articles like this. Thank you Henrik!
I think people do it because it is quite some work to implement proper value types in common languages like JavaScript, Java and C#.
Languages like F# or Kotlin give it to you almost for free.
The rule of thumb is exactly as you observed: if values have different semantics they should have different types.
[1]: https://fsharpforfunandprofit.com/posts/conciseness-type-def...
It's reasonable to put a little more effort into it along API lines. But there's a reason that the compiler doesn't make it easy to define an integer type that can hold values between -7 and 923091.
Good advice!
> Repave servers and applications every few hours. This means redeploying the same software – if an attacker has compromised a server, the deploy will wipe out the attacker’s foothold there
Yes, but keep in mind that the same attacker will be able to run the same attack successfully again, as long as they have an attack vector.
> Repair vulnerable software as soon as possible (within a few hours) after a patch is available.
Very good advice, and for that you need somebody to update vulnerable libraries and OS packages - like what Linux distributions do - unless you want to maintain hundred of packages by yourself.
I'm reminded of 'fileless malware'. A virus can reside exclusively in volatile memory. Worst case here would be to have a set of servers continually reinfecting each other even as you continually repave. I imagine the solution to that would be to repave the whole set and then switch over, like double-buffering. (Of course, not using files isn't exactly a strength here, but it seems apropos.)
> Yes, but keep in mind that the same attacker will be able to run the same attack successfully again
Exactly. Detecting that deployed files were tampered with can be tricky (one has to take updates into account, and some attacking code may be able to detect this analysis and nurture it with the original version of the files).
No auto-correct culprit here, that's plain and simple self-incorrect (!): English isn't my native language.
>> For example, say that you want to represent the number of books ordered. Instead of using an integer for this, define a class called Quantity. It contains an integer, but also ensures that the value is always between 1 and 240
The Ada code to implement this is:
type Quantity is new Integer range 1 .. 240;
>> instead of just using a string, define a class called UserName. It contains a string holding the user name, but also enforces all the domain rules for a valid user name. This can include minimum and maximum lengths, allowed characters etc.
The Ada code to implement this is:
with Ada.Strings.Bounded; package UserName is new Ada.Strings.Bounded.Generic_Bounded_Length (Max => UserName_Max_Length);
Dynamic predicates or even a string subtype could be used to further refine the UserName definition depending on exactly what restrictions are needed.
While it's not perfect, Ada does make it pretty easy to specify constraints on data types and will complain loudly when the constraints are violated.
Common Lisp does too:
(deftype quantity () '(integer 0 240))
(defun foo (x)
(declare (type quantity x))
(1+ x))
(foo 1) → 2
(foo -1) → ERROR
(defun valid-username-p (string)
(and (< 8 (length string) 24)
(every (lambda (char)
(find char "abcdefghijklmnopqrstuvwxyz0123456789-_=./" :test #'char=))
string)))
(typep "foo" 'username) → NIL
(typep "foobarbaz" 'username) → T
(typep "foobar-baz" 'username) → T
(typep "foobar-baz " 'username) → NIL
Common Lisp is pretty awesome. type
Quantity = 1 .. 240;
Colors = (Red, Green, Blue);
Pixels = array [1..1024, 1..768, Colors] of Byte;
Just a small example how its type system already so much better than C, although not as good as Ada.And overall, Pascal's type system is NOT so much better. Not strictly better at all, if at all any better. It's a chore to do the simplest things, starting from the mess that is the various types of strings, to a confusing memory management story, continuing with extremely verbose type declaration syntax (which requires to add many additional names), to the mess that is 0-based vs 1-based indexing, and let me not start with the messy object systems that were put on top in Delphi.
If you ask me it's definitely WORSE over all, although for example Delphi has nice aspects to it, especially in the IDE.
Oh yeah, and if Ada was ever adopted by a significant adoption of programmers, then they probably have committed suicide in the meantime.
By the way, better check your github issues.
I don't know how much effort you put into finding this, but the result feels almost like a certificate of quality to me and confirmed my opinion that this style of coding is pretty f*ing safe in practice. And that while I spent definitely less than 1% time on debugging memory issues (running valgrind twice, according to my git history, compared to working on this project during 7 months, initially 2.5 months full time, leading to an estimate of about 500h of development time).
There's a whole well-reviewed book on exactly this: https://pragprog.com/book/swdddf/domain-modeling-made-functi...
See the same program in flow [1] (nominally typed) and TypeScript [2] (structurally typed).
In the case of flow the type can only be constructed with the class -thus enforcing the guards - whereas in TypeScript I can accidently (or deliberately) bypass all guards by having a class with an equivalent structure.
[1] https://flow.org/try/#0MYGwhgzhAEAKD2ECWAXJA3ApgSQHYswHNMAna...
[2] https://typescript-play.js.org/#code/MYGwhgzhAEAKD2ECWAXJA3A...
type Brand<K, T> = K & { __brand: T }
type USD = Brand<number, "USD">
type EUR = Brand<number, "EUR">
const usd = 10 as USD;
const eur = 10 as EUR;
function gross(net: USD, tax: USD): USD {
return (net + tax) as USD;
}
gross(usd, usd); // ok
gross(eur, usd); // Type '"EUR"' is not assignable to type '"USD"'.
[1] https://michalzalecki.com/nominal-typing-in-typescript/#appr...* It uses an arguably invalid construct "K & { __brand: T }", where K is not an object, is an empty intersection. The fact that typescript allows casting a number to this type is concerning.
* Typescript currently allows "{} as USD" for non-object "K"'s but this will throw up serious issues down the line (obj is not number etc.); this is a likely error after validating JSON for example.
* Similarly typescript will allow you to bypass the guards for primitive types by using the structurally invalid value "{__brand: 'USD'}", or bypass them for object types using a structurally valid form with a "__brand: 'USD'" member. Which is more concerning I don't know.
* The type system now believes you have a member "__brand" that you don't actually have.
* In summary, you cannot enforce the guards through the type system.
That said, this is an interesting hack that, assuming your developers aren't trying to hurt you and you don't use it for primitives, could help get some extra safety in there. However the absurdity of the intersection looms heavily over it, I wouldn't bet on this working in a few years...
I just recently was working on a backend application and was modeling the database entities. The project uses UUIDs as primary keys. Should each entity have its own primary key type? `Location` gets a `LocationId`, User gets a 'UserId', etc, etc, where they're really all just wrappers around UUID?
Honestly, I thought about doing that a bunch of times during the start of the project, but I was pretty sure I'd get some harsh, sideways, glances from the rest of the team.
It makes sure that you never pass a `LocationId` where a `UserId` is expected; the type system literally will not allow it.
The problem is a non-trivial one even for 'simple' things like a person's name. Having a rule that takes in languages, special characters, spaces etc is hard.
As someone who loves the idea of dependent types, it seems to me that this is the best theoretical solution, but maybe not the best practical solution. If solving a simple-seeming domain problem involves modelling, not just first-order types, but the whole theory of dependent types, then I think people are going to start looking for escape hatches rather than upgrading their mental models.
E.g. C has a very weak type system, which is static. There's a lot of implicit conversion going on. Also the expressiveness of the type system is very limited (in C++ also).
OCaml, F#, Haskell and other functional candidates are strongly and statically typed, with very expressive type systems.
Idris with it's dependent types would be ideal and goes even further than the above.
In embedded most likely ADA and Rust offer strong enough static type systems.
It is true that the subset that C++ shares to C has too many implicit conversions, but you can do much better.
For example in C++ you can use enum class to define strongly typed integrals that do not implicitly convert to the basic types.
I just wish I could use rust or ADA at work.
Rust does a great job in that regard and is not too slow.
How is this helpful? To automate it means there is another system that could be attacked and it‘s a valuable one as it manages all the secrets, or not? What‘s the story?
If you automate this and run it on an automated schedule < 30 days then it is pretty likely that it won't be causing downtime unexpectedly, that you'll have monitoring in place to make sure it actually gets done, that, even if you forget to trigger it for a specific reason (e.g., aforesaid person leaves the organization) it will happen within a reasonable period of time.
In terms of securing such a system... you need to make sure that you separate the system into appropriate pieces with limited access. So, for example, you want a job that is run with an account that only has access to rotate the credentials. It can't use them for anything, just rotate them. Services that consume those credentials should not be able to update them, just use them. You can then ensure that the process that rotates credentials executes in a highly locked-down part of your infrastructure.
Indeed, automating this process also encourages you to create processes with limited access, rather than relying on administrators who have so many responsibilities, you probably just throw them in the equivalent of wide-open sudoers file and call it a day.
It sounds complicated, but if you have decent abstractions, this kind of stuff is actually pretty easy to accomplish.
I'd be interested in seeing any end-to-end examples of how people are doing this in practice.
For example, suppose you're maintaining a SaaS application and you have a private key to access some third party API that certain parts of your back end code need. How do you automate this process, so you change your private key on a regular schedule and update all affected hosts so your application code picks up the new one?
Ideally this needs to avoid introducing risks like a single point of failure, a new attack surface, or the possibility of losing access to the API altogether if something goes wrong. Assuming the old key is immediately invalidated when you request a new one via some API, you also need a real time way of looking up the current active key from any of your application hosts when they need it, again without creating single points of failure, etc.
No doubt this could be done with enough work, but it doesn't feel like a trivial problem.
That doesn't solve your problem, but it means that you should take this complaint to whatever API service offering you are using.
1. The system that gives a service to the public is publicly accessible by default. The system that rotates its credentials it doesn't need to be; it can sit behind your firewall listening only on one port for ssh, with firewall rules allowing access only from a bastion host.
2. The credential rotator also connects to your server and drops credentials into it; you don't call from your server to the credentials rotator, because see Point 1 above. This limits the attack surface.
You're right that everything is potentially vulnerable, and so would the credential rotator system. However, by rotating secrets this way, you shift the locus of security to a smaller place that you can defend much better.
Same goes to 2 apps connecting with a shared secret, making both of them change it wont expose any more components but will add additional layer of security (like google authenticator)
During a breach, if each service gets their own secrets, it becomes easier to trace the entrypoints and which secret go compromised. Once the system is closed again the attacker automatically looses access to everything after 4h.
Credential rotation processes are yet another layer of “defense in depth” — the more layers you have, the more secure you can be.
As you rightly point out, it adds complexity. Having a credential generator+rotator means you are more resilient to the fallibility of humans choosing bad credentials or too busy/lazy to do the task.
* http://cr.yp.to/qmail/guarantee.html
* http://cr.yp.to/qmail/qmailsec-20071101.pdf
* https://hillside.net/plop/2004/papers/mhafiz1/PLoP2004_mhafi...
* https://blog.acolyer.org/2018/01/17/some-thoughts-on-securit...
I was expecting something more of security related content and mitigation tactics.
I will agree that a well-trained and well-coordinated set of individual software builders can occasionally pull it off economically, but it requires too many things to go right in terms of organization and staff: a lucky accident. Most IT shops are semi-dysfunctional because the usual frailness of human nature wins out over rationality the majority of the time. Dilbert™ is life, life is Dilbert™.
“Any integer between -2 billion and 2 billion is seldom a good representation of anything.”
I think that's something Ada got really right.