Software Security Is a Programming Languages Issue (2018)
pl-enthusiast.net
pl-enthusiast.net
Regardless of language, all software has bugs; that's such a banal argument it barely even needs to be made. Offensive software security is the practice of refining and chaining bugs to accomplish attacker goals. You can make it easier or harder to do that, but in the long run, more bugs classes are going to be primitives for attackers than not.
Rust is a fine language to teach people. But what Rust accomplishes for security is table stakes. People building software in Java in 2002 had largely the same promises from their language; Rust just makes those promises viable in a larger class of systems applications. But Java programs are riddled with security vulnerabilities, too! So are Golang programs, and, when Rust gets more mainstream, so too will Rust.
Aside: I'm really frustrated by the advice to validate all input. That's not useful advice and people should stop teaching it. It begs the question: validate for what? If you know every malicious interaction that can occur, you know where all the bugs are. And you don't know where all the bugs are.
This works in the real world. For example, see airliners.
A large part of my job as a language designer is to look at the causes of certain kinds of bugs, and find ways to design them out of the language. There's no reason today that array overflows should still be happening.
Then a follow-up question would be what indications we have that other classes of bugs can or can't also be eliminated with help from languages. (Of course there's also potentially a lot of space between "it's inherently impossible to write programs that are incorrect or unsafe in a specific way", which is something that things like memory safety and static typing can help with, and "it's possible for a programmer to do extra work to prove that a specific program is correct or safe in a specific way", which has been a common workflow in formal methods.)
I would be interested in your view of this; after memory safety and perhaps type safety, what kinds of bugs do you think could be eliminated by future language improvements, and what kinds of bugs do you think can't be eliminated that way?
I remember arguing on Reddit about the value of serious type safety languages, and the notion that they could eradicate things like SQLI bugs (ironically: the place where we still find SQLI bugs? Modern type safe languages, where people are writing raw SQL queries because the libraries haven't caught up yet). I'm both not sold by this and also stuck on the fact that there's already a mostly-decisive fix for the problem that doesn't involve the language but rather just library design.
Setting aside browser-style clientside vulnerabilities, which really are largely a consequence of memory-unsafe implementation languages, what are the other major bug classes you'd really want to eliminate? Scope down to web application vulnerabilities and it's SSRF, deserialization, and filesystem writes. With the possible exception of deserialization, which is an own-goal if there ever was one, these vulnerabilities are a consequence of library and systems design, not of languages. You fix SSRF by running queries through a proxy, for instance.
Also, the Latin plural of "programma" should probably be "programmata" (which also works for agreeing with the neuter plurals "omnia" and "purgamenta", whereas "programmae" would be a feminine plural which wouldn't agree with the neuter adjectives). Most often Latin uses the original Greek plurals for Greek loanwords.
https://en.wiktionary.org/wiki/%CF%80%CF%81%CF%8C%CE%B3%CF%8...
> where people are writing raw SQL queries because the libraries haven't caught up yet
SQL is a pretty good way of representing queries, I don't see how a library could be constructed that could do better. But then as an SQL guy I have a hammer so Omens Clavusum Est.
Edit: plural, so clavusuma? clavusa?
No one wants the overhead on current hardware but C can be made memory safe with bounds checked pointers (although not particularly type safe without introducing a lot of code incompatibilities).
It's important to distinguish actual language advantages from indulging the common fetish of wouldn't-it-all-be-better-if-we-rewrote-it.
There has been fairly little rigorous study of programming factors leading to more reliable software (esp in the kinds of environment most software we use is created and used in, as opposed to say aerospace).
Wouldn't it be sad if enormous amounts of effort were spent rewriting everything in a new language only to later discover that other properties of that language, such as higher levels of abstraction or 'package ecosystems' that cause importing impossible to audit millions of lines of code to get trivial features, lead to lower reliability than something silly like using C on hardware with fast bounds/type checked pointers?
Well it depends. That overhead may be quite acceptable in many places. Certainly not 'no one'. However you've not quantified what that might be, IIRC some guy turned on int overflow/underflow checking and it only cost 3% extra time. Guess that's down to speculation & branch prediction.
I don't like any overhead in time or hardware either, but measure first.
"Secure software" as an absolute, objective construct doesn't exist, so that's not really an interesting argument.
What's interesting is if it produces software which is more secure, and I think it's a reasonable argument that eliminating memory safety issues does just that.
One can't defend against unknown unknown, but mitigating known issues (input validation) and delegating to tools that know better than you (rustc) are both valid strategies to produce software that is more secure upfront (and I know you know it, eh)...
What security issues have been recently solved by programming languages?
Things like unsafe memory usage are an anomaly. Most other security issues can't even be represented at the PL level.
The cost in those cases is performance. The cost is offset massively by productivity.
It's less about direct costs and more about constraints and priorities.
Imagine you tasked two rooms of 100 jr devs to query a database, one with EF, one with ADO.NET. Do you honestly think the same percentage of jr devs would have written code open to SQL injection?
string queryString = "SELECT ProductID, UnitPrice, ProductName from dbo.products " + $"WHERE ProductName = {userSuppliedName} " + "ORDER BY UnitPrice DESC;"; SqlCommand command = new SqlCommand(queryString, connection); SqlDataReader reader = command.ExecuteReader();
I've met a ton of devs who would write this code without thinking(some even with quite a few years of experience.)
'Taint mode' ala Perl[1] turns non-validated input into runtime errors, rather than security bugs. Crashing is nicer than being pwned.
"You may not use data derived from outside your program to affect something else outside your program--at least, not by accident."
Bugs still exist (as well as bugs in languages themselves), but languages can help mitigate the surface area.
Do you realise this reads as "don't bother validating your inputs"? Surely that's not what you meant?
> It begs the question: validate for what?
All inputs necessarily follow some protocol. Command line options and configuration files have syntax. Network packets have sizes & fields. Binary files have some format. So you just make sure your input correctly follows the protocol, and reject anything that doesn't.
Now the protocol itself should also be designed such that implementations are easily correct and secure. In practice, this generally means your protocol must be as simple as possible.
And it's probably a good idea to make sure your parser is cleanly separated from the rest of your program, either by the OS (qmail splits itself in several processes), or by the programming language itself (with type/memory safety).
Unless I missed something, you couldn't have meant that input validation is useless. You most probably didn't mean that input validation is superseded by some other technique, and therefore best ignored. You couldn't possibly think that everyone already does input validation, and therefore don't need the advice.
But then I wonder. Why saying "validate all your inputs" is not useful? What would you say instead?
Besides, it's not just about us: other people, (including @carty76ers apparently) would like to know what you would advise instead of "please validate all inputs".
The reason I kept guessing, was because you kept not telling. Until then: https://news.ycombinator.com/item?id=21719065
Finally something I can argue with.
This suggests teaching what validation means, not abandoning the idea of validating. In other words, be especially aware that you're crossing a trust boundary when accepting input, and use proper validation algorithms, like proper parsing instead of ad-hoc pseudo-parsing using regexes.
But getting developers to use parser generators has been a struggle because I'm finding it hard to give dev team tools they can use and integrate into their build systems. It also doesn't help that most developers have forgotten anything they learned about formal language theory, so it's hard to communicate the value of using parser generators.
But we are getting there. Slowly, but there is progress.
At my company we often do bug bounty and hacking events where we have people attempt to break into our systems. Some of our engineering team members always make a big deal about making sure only certain things are in scope and that we only allow hackers to break things that won't cause major issues.
I always tell them that this is impossible. Sure, we can tell the hackers to only target certain end points, but we have no way of knowing what downstream things will be effected, and if we can't be sure there is no way the hackers could know.
If we knew for certain what exactly could be broken by the hackers, we would know exactly where are vulnerabilities were. If we know that, why are we paying hackers in the first place?
Sure enough, at one of our events a system waaay out of scope ended up breaking when a hacker triggered a bug on an in scope system. A bunch of our engineers were upset, but the hacker had no way of knowing that his api request would trigger a series of further actions in systems leading eventually to the breakage. Any system connected to an in scope system is automatically going to be in scope.
But, while I know nothing is a panacea, memory safety is especially not a panacea. We have over a decade worth of experience with software built principally in memory-safe languages; it's hard to find a modern web application --- increasingly, it's even hard to find a modern mobile application --- built in a language as bad as C. And, as you know, all these memory-safe applications are riddled with vulnerabilities. They don't have memory corruption vulnerabilities, but I'm still game-over'ing app pentests, and very unhappy with myself when I can't.
But they do think (like Daniel Bernstein and Edsger Dijkstra), that the usual "pentest & patch" does not work. http://www.pl-enthusiast.net/2015/09/30/penetrate-and-patch-...
I've talked to more than one university professor in the last 2 years about what a first and second course in software security should look like, and if there's a theme to my feedback, it's "not like this; this is what people thought it might look like in 1998".
I was saying this is most probably not the premise of the course. That just like you, they don't believe they're teaching a panacea, just something that will help.
That assessment could be wrong, though. Unlike you, I never discussed security with university professors, and I know nothing of their misconceptions. Could you offer a gist of what they teach, and most importantly what they should be teaching instead?
For example, plan interference always arises in high-level planning, so no language is going to eliminate that bug class without somehow completely reinventing the philosophy of planning and logistics.
We can certainly ask that languages do better around features like FFI, string-building, and error-handling. We keep making the same mistakes. Memory safety is probably the biggest conceptual leap forward in our lifetimes and we are still not yet in agreement as a community that it is clearly good.
It's not that we're not in agreement. It's just become standard for most developers so they don't care about it.
Memory safety is only still relevant in the context where one is using C and C++ and maybe Obj-C. And in that context many do not agree that Rust is the solution for them, even if they appreciate memory safety otherwise.
Our compilers should long have included static analysis on par with optimization efforts, because - as you say - people tend to produce errors in code. Strong type systems and memory safety are a very nice first step though.
I am certain this will change over the years, but then the trend (that I am also prone to) is to go dynamic for the sake of productivity anyways.
Java, C# (not all .NET languages) are both quite strongly typed (it would be nice if there was no implicit toString conversion in C#, I think Java has this, too).
F# is a bit better than C# in terms of type system (especially exhaustive pattern matching on discriminated unions leads to a lot more potential type safety).
C or C++? They are happy to implicitly convert a lot of things, they're statically typed, but very weakly so. And in terms of memory safety C strings and arrays are the reason why static analysis is absolutely necessary here.
>Our compilers should long have included static analysis
The Rust compiler is doing a lot of what would be considered static analysis. And if you're dealing with normal strings and arrays you're not prone to C-like buffer overflow behavior. If you run in debug mode then you even get assertions on integer wrapping. This IMHO is also the biggest issue with C/C++. The fact that the compiler still allows so much to go through and only the static analyzer gives you a warning. Big companies rarely have code bases where most of the static analysis warnings are even solved. And it's hard to hire only "disciplined" programmers.
Where is this definition coming from? Seems arbitrary.
Static tools capture the same old range of security issues which is only mostly relevant for C and to some extent to C++ and also Objective-C and unsafe Rust or unsafe Swift.
Yet the javax crypto libraries are routinely misused. ECB defaults. People fucking up IVs. The whole shebang.
The question is: when we find a common class of vulnerability, what's the best way to deal with it?
Functions must only accept input which they can properly act on.
A function serving web pages needs a different notion of String than a database function. There are Strings which are safe for a SQL use but unsafe to serve on a web page, and vice-versa.
Validation is restricting your input to only those things your function can act on.
For students: learn a lot about the most important classes of security vulnerabilities, of which memory corruption is one important example but just one, and then take the time to learn how to exploit at least simple variants of all of them in a realistic setting, to cultivate the mindset needed to think critically about software security.
Don't write anything in C. Sure. But really almost nobody does that anymore anyways.
Nope. See (for example) CompCert, SeL4, ...
If your software asks users to input their age, to validate that, you would discard any input not in the set 0123456789, and any valid input (a number) that is less than 1 or greater than 130.
It's not so much about input validation as it is about sanity checking the input. Is it sane? If so, then you should accept it. It may still be incorrect (user input error.. entered 24 rather than 25) but it should be safe to treat this input as an unsigned 8 bit number and manipulate it as such.
The difficulty is that some inputs have a large, varied set, but even those can be bounded (and are) in the real-world. So if someone enters a first name that is 500 characters long, that should fail sanity checks.
The problem with CS people is that they obsess over edge cases (100% correct and verifiable solutions) sometimes when they should not. I don't blame them for this, as that's a large part of what they were taught to focus on in school.
User input is not an algorithms problem that needs a 100% correct and verifiable solution, it's real-world, can be reasonably bounded and good enough solutions are sufficient. Edge cases can be handled manually and added to the existing solution, too, if they are more common than what they initially seemed.
Raise your hand if you are executing code in your project that was written OUTSIDE your organization, and has never been reviewed by people within your organization.
This is the ham sandwich problem.
If a stranger walked up to you on the street and handed you a ham sandwich would you eat it? I would venture to guess that most of you would NOT. However many of us are all too happy to grab some random chunk of code off the internet and shove it into production without a second thought.
Personal and social information of 1.2B people discovered in data leak https://news.ycombinator.com/item?id=21606415
Is this another ham sandwich from a stranger that has been eaten? How often is the "bug" implicit trust and poor design or default behavior?
What about Specter? A hardware level exploit was bound to come up again, we have had hardware bugs before (1994 pentium math bug) but exploitable ones are "fairly rare".
------------
The old mantra "security through obscurity" is true, but it has some serious validity issues when there aren't enough eyeballs on the software we are already running, and were generating more at a rate where people can NOT keep up! (This doesn't address the fact that we are now building opaque boxes with ML that NO ONE understands accurately).
"we want to hire some people with no background checks, no interviews and no idea where they live to write code for our critical line of business application. We're then going to put that code into production without reviewing it, and regularly update it without reviewing the updates. Oh also we won't have any form of enforcable contract as come-back if something goes wrong."
You'd, at best, be laughed out of the room.
Yet that's exactly what pretty much every company does with 3rd party libs.
CEOs outsource critical things without meaningful oversight all the time.
For hiring, no chance you'd get that policy past corporate HR in any large company. Try hiring a developer sight unseen to work remotely with no contract in any large organization and see how well that goes.
Yet companies effectively do just that with 3rd party library use. the reason the CEO doesn't do anything isn't likely to be because they've made an informed risk decision on the topic, it's because no-one is telling them the risks :)
If you're trying to talk about hiring for full time employment as opposed to contract work... what happens there is as long as people are able to get through whatever idiosyncratic hazing process was involved in hiring, they're going to be at the company for at least a year. It's perceived as hard / risky to fire people, even if they can't program their way out of a wet paper bag.
This stuff happens all the time, it's not that different from evaluation and use of third party code. "This project has 500 stars on github, it must be good." "This guy used to work for Google, he must be good." Now you're stuck.
Same with hiring, the person may or may not be able to code, but active malice would likely result in firing, and the code they write should be subject to review before being put into production.
Of course you can argue "hey where I work hiring is trash, we write bad contracts and have no internal standards, so this 3rd party stuff isn't much worse" but I'd suggest that's not an argument most companies would make publicly about their processes.
Point taken about active malice... but I've also seen companies with mostly in-house code cover up instances of rootkits on production servers and malfeasance related to credit cards.
I'd rather companies use third party open source crypto, for instance, even if it sometimes gets compromised, because it's a lot more likely to come to light.
In the sandwich case, the user is the customer. You wouldn't eat the sandwich, because it might make you sick.
In the software library case, the developer sees the benefits of using a library, but rarely sees any penalty if it goes bad. They'll earn the same paycheck, and probably even work the same hours. They might not even be the same team that has to deal with security issues in production. They get the benefit of saving time, and accept none of the responsibility.
While code level issues should absolutely be a focus in the SDLC, it's common to find security issues crop up from:
* Hardware, kernel, OS, package, and library vulnerabilities
* Component integration / API contract misunderstandings
* Transitive trust between services and third parties
* Accumulation of access over time
* Demos, hotfixes, and workarounds that are somehow now mission critical
Just because you've written secure components in a safe language doesn't mean you don't have security issues when you run them together.
Perl also has strict mode, to tighten up your programming. Without it you don't have to declare variables, but on it requires variables to be declared before use.
More languages should help programmers like this - sort of a ladder to go beyond just low hanging fruit.
I personally think it's a great idea. Loose types can get a POC or even early production models up and running quickly while you're changing your opinions on the data every hour, then once it grows to a size that's hard to reason about and static analysis can really pay off, start tightening it up.
I wrote about a similar topic from a web application security point of view: Why Framework Choice Matters in Web Application Security* ( https://www.netsparker.com/blog/web-security/why-framework-c... )
Also today there is enough data in the industry to prove this argument beyond any doubt for web applications.
* original article is written about 11 years ago or something this is a republished version
There are a lot of basics in this post.
However, this is really only a small fraction of the issue in application security.
Just thinking back to some recent major breaches recently in the headlines, we have
* failure to update (Equifax)
* unencrypted backup files (Adobe breach of long ago)
* A long-ago root-level compromise of 90 servers of a giant bank.
* The Target breach: vendor access to network
* Numerous breaches related to improper setup of AWS
Also, two interesting SSL/TLS vulnerabilities had nothing to do with anything a language design can address. The GOTOFAIL and HeartBleed. In fact, someone illustrated how to make the same error in Rust (and promptly got downvoted).A good view of front-line security problems is addressed in https://www.youtube.com/watch?v=_4vSurKPl6I (Attack Oriented Defense.)
I have audited applications in many languages, from C, Clojure, .NET, Java, Perl, and Ruby. Vulnerabilities found did not relate to langsec at all.
The article starts out talking about a programming languages course, then tries to justify including security therein. I think they're really just hammering home the PL-specific case of your point that "security is part of the daily programming job".
One thing I do to practice "continuous security", is to accompany each code review with an additional code review, solely focused on security. Else it's too many balls to juggle when doing a general-purpose code review.
Do you know of similar simple yet effective techniques?
Because most software developers have no idea what security is. By most I mean almost all. This point is easily proven. Ask any software developer what security is and compare their answer against the standard answer. All security courses and certifications I have seen define security in exactly the same way.
If people wanted to take security seriously in software they would train their developers on security or require that they be security certified.
> I believe that if we are to solve our security problems, then we must build software with security in mind right from the start.
Yes, but clearly most organizations don't take security seriously. Instead they bolt it on at the end just enough to appease the corporate attorneys the same way they do for accessibility or any other necessary requirement whose absence results in class action lawsuits.
It should be noted that type safety itself does not solve this problem. For external input, you need explicit schema validation or the type needs to be enforced at the protocol level using something like Protocol Buffers with implicit schema validation.
I think that security has nothing to do with the programming language and everything to do with the developers who are writing the code.
Secure Design: A Better Bug Repellent Christoph Kern, IEEE SecDev '17
https://s3.amazonaws.com/cybersec-prod/secdev/wp-content/upl...
Then, how could a programming language help me prevent high-level security bugs like Shellshock? Is it even possible/practical?
Yes. As trivial example, consider a UDP packet format { u16 type,size; u8 data[size]; } fed the 4-byte packet {ECHO,0xFFFF}, which is incompetently parsed as {ECHO,"<65535 bytes of stack memory>"} because the 'shotgun parser' assumes its input is well-formed. Whereas a 'recognizer' (ie a non-buggy parser) would reject that packet as 65535 bytes too short.
Good programming language design can make it harder to write the buggy parser and easier to write the 'recognizer', especially if the language/standard library provides built-in parsing tools.
I'm not being snarky. Compare "LangSec" to memory safety, which also kills this class of bug dead. Which approach is more powerful and forecloses on more bug classes? Which approach requires more developer effort? Introduces more jargon?
I know multiple very smart, capable people who work under the rubric of "LangSec". But I just don't get it. Is it a real thing?
FWIW, I think LangSec is saying "Code that doesn't have remote code execution vulnerabilities[0], or limits them to a weak computational model[1], is better than code with RCE vulnerabilities." - which is also "Well, no shit." - and "Parsing a nontrivial data format is the same thing as executing a (not-necessarily-)very constrained programing language."[2] - which seems obvious to me, but could plausibly be a "superpositions don't collapse"-level epiphany for someone who doesn't think about parsing the right way.
0: such as javascript or stack execution
1: like FSMs or pushdown atomata
2: with the implication that you had better make sure it actually is very constrained
There are whole classes of errors related to programs that parse then validate input when it's already too late. And often the validation happens in the source code in a cloud of checks that happen at run time. It's rather difficult to verify these programs.
It is much easier to verify a parser that only produces valid values at the edge of a program, isolated from the main program.
I'm familiar with the LangSec lingo and the concept of a "weird machine", much as I hate the term itself, has value. But it's not a product of LangSec so much as a name for a concept we've had for decades.
In Heartbleed the parser parsed an incoming field specifying length, but failed to correlate it with the rest of the request. Had the parser been written to a stricter specification, that would not have happened. Typically that is what is meant by bounding computational complexity. Or at least that is my understanding of it.
https://stackoverflow.com/questions/25353753/python-can-i-sa...
Also, how can you design an efficient data format that accepts data of a length determined by the user (hopefully cooperatively with the server), that is immune to buffer overflow reads.
Once you say "I want between four and one thousand bytes", haven't you just stuffed yourself?