Firmware Software Bill of Materials (SBoM) Proposal
uefi.org
uefi.org
Knowing know which version of a file made it into a binary still doesn't really help you, though. The compiler used (if any), the version of the compiler and linker, and even the settings / flags used affect the output and -- in some cases -- could convert an otherwise secure program into something exploitable.
A Software BoM sounds like a "first step" towards documenting a supply chain, but I'm not sure it's in the right direction.
This feels like this might actually be a use-case for a blockchain or a Merkle Tree.
Consider: A file exists in a git repository under a hash, which theoretically (excluding hash collisions) uniquely identifies a file. Embed the file hashes in the executable along with a repository URL and you essentially know which files were used to build a file. Sign the executable to ensure it's not tampered with, then upload the hash of the executable to a block chain.
If your executable is a compiler, then when someone else builds an executable then they can embed the hash of the compiler into the executable to link the binary back to the specific compiler build that made the binary. The compiler could even include into the binary the flags used to modify the compiler behavior.
edit: or do you mean that your hypothetical tool would generate the repository URLs + file hashes for its dependencies as well and bundle those with its own?
We still do this today. We rely on some tooling-added information and add our own custom information on top (git hash, hash of direct dependencies, compiler version and flags, etc).
Merkle trees sure but you do not need a blockchain to store/manage hashes of sources and build products. Just have trustworthy parties attest that source commits yield particular binaries by their hashes and publish those signatures somewhere. Even Bitcoin Core does something like this using PGP.
But maybe it's not necessary.
DNS Root Zones, for example, are publicly accessible and well managed (maybe not the registrars, but the registry) so perhaps a well-funded third party could establish a trusted, global, registry for supply chain Merkle Trees.
A few years ago, a similar idea for firmware binary security[0] had been explored by Google as a possible application of their Trillian[1] distributed ledger, which is based on Merkle Trees.
I don't know if they've advanced adoption of Trillian for firmware, however, the website lists Go packaging[2], Certificate Transparency[3], and SigStore[4] as current applications.
[0] https://github.com/google/trillian-examples/tree/master/bina...
[2] https://go.googlesource.com/proposal/+/master/design/25530-s...
slsa-framework/slsa-github-generator > Generate [signed] provenance metadata : https://github.com/slsa-framework/slsa-github-generator#gene... :
> Supply chain Levels for Software Artifacts, or SLSA (salsa), is a security framework, a check-list of standards and controls to prevent tampering, improve integrity, and secure packages and infrastructure in your projects, businesses or enterprises.
> SLSA defines an incrementally-adoptable set of levels which are defined in terms of increasing compliance and assurance. SLSA levels are like a common language to talk about how secure software, supply chains and their component parts really are.
There may be multiple sets of compiler flags and even multiple compilers (can there even be multiple linkers?)
In addition, having to pick between multiple functions called foo, a linker may pick either of them, and may not always pick the same one (parallel linkers often aren’t deterministic), even if their implementations are different.
There also is a decent chance you’ll have to document your OS, for example if the compiler dynamically links to a system library, or if the OS has FPU support in software, or if the OS tweaks some obscure CPU flags (and of course, you’ll have to document the specific CPU and CPU revision, as those can affect things like constant folding.
> Consider: A file exists in a git repository under a hash, which theoretically (excluding hash collisions) uniquely identifies a file.
Git is of course incapable of including a file version number, as that’s not really well-defined in a distributed setting. But if you’re OK with a blob hash, put $Id$ in your source files, mark them with “ident” in .gitattributes, and you’ll see the hashes included and autoupdated on your next checkout[1]:
> When the attribute `ident` is set for a path, Git replaces $Id$ in the blob object with $Id:, followed by the 40-character hexadecimal blob object name, followed by a dollar sign $ upon checkout. Any byte sequence that begins with $Id: and ends with $ in the worktree file is replaced with $Id$ upon check-in.
However, there’s a pretty simple way to make sure the bill of materials can be checked by machine:
Require the firmware blob to be reproducibly buildable from source, and mandate distribution of the source code of the firmware.
Doing this doesn’t preclude signing the result of the build and then distributing the signature instead of the signing key.
I’d personally prefer it if vendors also had to distribute signing keys, but that’s a separate question.
This would completely and entirely defeat the purpose of signing. Once a private (signing) key is no longer private, it's useless.
You can disagree, and it is unlikely that you'll change my mind.
A nice compromise is what my old Acer laptop does. You can disable the windows public key, and put your own in place if you want to run a CA, or you can just tell it to trust whatever boot loader is currently installed (and lock it to that boot loader, so now it won't boot Windows or even Ubuntu, but it will boot the Linux or BSD you installed). You can also disable the security checks if you want to.
I'm not looking to change your mind, merely pointing out how absolutely pointless is signing if private keys are public. At that point just don't sign, and don't add the silicon/firmware that verifies the signatures - less cost designing/testing/manufacturing/supporting the hardware.
The point of signing the firmware/bootloader is to make it more difficult to employ boot chain exploits. If you don't care that anything that even briefly gains root on your machine (curl|sudo bash) can proceed to permanently backdoor your OS in a way that's undetectable from the inside, just disable signing. /shrug
> A nice compromise is what my old Acer laptop does. You can disable the windows public key, and put your own in place if you want to run a CA, or you can just tell it to trust whatever boot loader is currently installed [...].
Either Microsoft, or the EFI standards body (under MS's direction; can't remember the details here) mandates that users must be able to enroll their own signing keys, at least on x86. I haven't run into PC hardware that doesn't allow you to do this, but then I'm not dealing with a lot of PC hardware regularly. The story is a bit different with Arm EFI - Microsoft tried to ape Apple and make it mandatory to NOT be able to enroll new first-party keys; meanwhile Apple skipped EFI on Arm Macs, and just allowed third-party OS's like they did on Intel. I'm not sure what's the story with signing a third-party OS on the Arm Macs, but I wouldn't like to run one unsigned.
Their example entry basically just says (very simplified) "Intel both built and supplied the microcode for your Intel CPU". But it says nothing like "Intel used Jenkins version x.y.z which bundled log4j version a.b.c" so you still can't tell if your binary blob was built by a compromised system or not, even once you learn about the potential attack.
That said, my company is moving towards having SBOMs, but seems to be making them available on request only. Which in my mind defeats the point of having them.
I've spent the last six months in this field and people will tell you that this or that is an industry best practice or "a standard" but in my experience none of that is true. Everyone is still trying to figure out how best to protect the software supply chain security and things are still very much in flux.
Once the CI/CD environment is compromised, how would the SBoM be trustworthy anyway?
i’m pretty annoyed that stuff like these requirements just makes it into law without getting small companies or startups to have a say in whether they want this or not.
Half the things we are struggling with are due to upstream hardware vendors throwing a hot mess over the fence. Linux&friends being GPL doesn't help us as much as we hoped, because many important and low-level pieces are quite opaque, closed, under-documented, etc.
Anything to force the vendors to be more transparent is good for us. We already automate our builds, keep extensive notes on what works and what doesn't, etc. Our own SBOM is additive on top of the vendor's stack, so all we're really concerned with is the pieces we wrote ourselves. I don't see it as a burden any more than ensuring proper test coverage, or maintaining a lockfile of dependencies.
So yes, please make this a law.
you point out the problem we face today: people in open source typically work at big tech and regularly push ideas that end up costing small hw companies a lot of money. for example rust in the kernel. or over complicated frameworks for the sake of security while lots of hardware does not need security. or features that are important for servers and make those default while most hardware is not a server. and so on.
if open source was primarily managed by small company contributors it would look very different.
Here are some examples of people using Ghidra. https://github.com/evyatar9/GptHidra https://github.com/likvidera/GhidraChatGPT https://github.com/tenable/ghidra_tools/tree/main/g3po
I suspect there are better ones being worked on though.
Not LLMs specifically, but transformers more generally, ought to be extremely good at this task.
Another one I wonder about a lot is execution traces. Something that isn't mentioned often enough is that transformer models can do previous-token completion just as easily as next-token completion. So you can train the model on paired program/exectrace (training data is trivial to generate here -- just execute it on a CPU!) and then ask the model to work backwards from a desired machine-state you want to reach.
in this case the length is about "11–14 minutes" of text
If you want to use Open Source code in an environment, it is necessary to validate all of the code (or know who validated it), including all the dependencies. The maintainers don’t have any obligation to help you, and it is a tedious job that they might not be interested in (or maybe you are in the policing or defense sectors any the maintainers find your application objectionable).
but where does the software end and the hardware begin anyways?
Okay? Neither hardware nor license type has any relevance to an SBOM. On an SBOM, you list all of your software dependencies, regardless of license.
> but where does the software end and the hardware begin anyways?
Hardware is the physically tangible part of the computer. Software is the instructions that run on it. For those compiling an SBOM, I think that's probably an easily answered question.
I see them "bending over backwards" to protect their right to keep an advantage to use, if they deem it necessary, against any one else.
this is why they go to such lengths to avoid publishing what they consider to be "the secret sauce".
as I see things, that they even came up with the notion of a "software bill of materials" is a bit disingenious from my perspective. already the concept is obscure and lends itself (in my opinion) to shading, hiding, and occulting the software source code and (or) the hardware designs (as appropiate)
finally, I consider that the concept of "SBOM" (which TIL existed) is designed with the intentions I already mentioned: to occult information (anti-open) for the sake of keeping a perceived advantage (pro-centralization)
I'll give you an example of what they're intended to be useful for:
Let's say you're an organization and you use, say, 500 different pieces of software made by 100 different companies. If you learn that there's a vulnerability in a particular dependency, what do you do? In the past, there was no standard way that vendors communicated this information, so the answer is that you would go through your list of 500 software programs, and email 100 different vendors asking about each one. This is not a good process.
If, instead, everyone provided an SBOM with their software, all you need to do is run a query against whatever inventory management system you're using and you have the answer in seconds.
it seems to me that you're arguing back that SBOMs are just a list, so how can it be a big deal?
my point is an issue about the whole reality of needing SBOMs. clearly they have something to do with supply chain trust. but I think your perspective is focused too closely on the engineering aspects of the problem.
I find that the technical realities (all of which I understand to be, in the end, some kind of engineering decision) that motivate having to share vulnerabilities of all components without revealing their designs sources (of either software, hardware, both, or neither) to be morally dubious.
so keeping in mind that hardware, by this point, is stored as code so it's code, and software also has a source code, I find it necessary by this point to open source all the things!!1!
the technical problem is really essentially a matter of trust, blockchains solved this problem in the technical sense.
the opreations of large organization are typically private property... this is a touchy issue because technology corporations are essentially the government by this point
I guess I should be glad this isn't really my problem, I'm just worried about the public and political consequences of what I see as potentially dangerous mistakes being made on the idelogical level
The straw that broke the camel's back on this issue was CVE-2021-44228 -- which was a vulnerability in open source software. If you missed that debacle -- the problem wasn't that people distrust Apache or software that use Log4j. The problem was that people didn't know where all it was installed.
This was because it isn't currently a standard for software developers to provide a list of all of their dependencies, regardless of whether they're open source or not. This isn't because they are untrustworthy. It just simply isn't standard practice. SBOMs are an attempt to standardize such a list.
A blockchain isn't necessary here -- nobody is trying to lie about what version of Log4j they're depending on in a piece of software they're selling.
however, thanks to all this dialogue, do see where my own critizicism goes amiss; and for that, thanks a lot.
Yea, this is not true. There is NO legal obligation. Be careful with what you read on the internet.
Guidance and law are two very different things.
Or at least that was my experience while working on sovereign cloud stuff in big tech.
And EO 14028 does it: https://www.gsa.gov/technology/it-contract-vehicles-and-purc...
If you were a contractor then you would know all specifications are waived unless specifically designated by your MOL and documented in GSA/SOW.
So here is one of these internet know it all's that doesn't know anything.