Hidden dependencies in Linux binaries
thelittleengineerthatcould.blogspot.com
thelittleengineerthatcould.blogspot.com
Just because the complexity is hidden from you doesn't mean it's not there. You have no idea what is statically bundled into the CUDA libs.
But we do know what it can not static link to, any GPL library, which many indirect dependencies are.
What do you mean by interface?
A dynamic library is handled very different from a static one. A dynamic library is loaded into the process virtual memory address space. There will be a tree trace there of loaded libraries. (I would guess this program walks this tree. But there may be better ways i do not know of that this program utilize)
In the world of gnu/linux a static library is more or less a collection of object files. The linker, to my best knowledge, will not treat the content of the static libraries different than from your own code. LTO can take place. In the final elf the static library will be indistinguishable from your own code.
My experience of the symbole table in elf files is limited and I do not know if they could help to unwrap static library dependencies. (A debug symbol table would of course help).
The other side of not accidentally loading more into your process than you thought is breaking down shared libraries into increasingly smaller sizes. In its limit I imagine it would be akin to a function per shared library, which probably defeats the point a bit.
(Btw, I'm pretty sure dlopen itself can't be lazy, due to needing to run constructors; the root comment is a bit vaguely worded... but ofc that only matters after dlopen is called.)
The bug here is much more a changing hardware paradigm than it is an issue with shared library dependencies that recapitulate it. Things moved and the software layers kludged along instead of reworking from scratch.
Obviously what's needed is a layer somewhere in the device stack that "owns" the GPU resources and doles them out to desktop rendering and compute clients as needed, without the two needing to know about each other. But that's a ton more work than just untangling some symbol dependencies!
https://www.dependencywalker.com/
https://learn.microsoft.com/en-us/sysinternals/downloads/pro...
However, that "requirement" doesn't prevent you from shipping an empty libm (or other libs listed there.)
(The actual reason is probably that glibc is old enough to have lived in a time where you cared about saving time and space by not linking the math functions when you didn't need them...)
Consider liblzma: would liblzma-as-a-service really be that bad, especially if the service client and service could share memory pages for zero-copy data transfer, just as we already do for, e.g. video decode?
Or consider React Native: RN works by having an application thread send a GUI scene to a renderer thread, which then adjusts a native widget tree to match what the GUI thread wants. Why do these threads have to be in the same process? You're doing a thread switch anyway to jump from the GUI thread to the renderer thread: is switching address spaces at the same time going to kill you? Especially if the two threads live on different cores and nothing has to "switch"?
Both dynamic linking and static linking should be rare in modern software ecosystems. We need to instead reinvigorate the idea of agent-based component systems with strongly isolated components.
Wouldn't it open up for a new attack vector where process could read each other data?
To some extent we are, if what you do is work on backend RPC or web app frameworks.
But the better answer is because sometimes what you actually want is the ability to put a C function in a separate file that can be versioned and updated on its own, which is what a shared library captures. Trying to replace a function call of 2-3 instructions with your io_uring monstrosity is... suboptimal for a lot of applications.
And in any case, the protocol parsing you'd need to provide to enable all that RPC is going to need to live somewhere, right? What is that going to be, other than a shared library or equivalent?
The answer is "yes".
I won't stop you, if you want to make React even slower, be my guest. I want off this ride.
And if you have one per core anyway so nothing "switches"? Computers aren't single-core 80486es anymore. We have highly parallel machines nowadays and old intuition about what's expensive and what's cheap decays by the year.
But, if you’re interested in this architecture, smalltalk did something similar. Fire up a smalltalk vm and play around!
Why? Relative to the in-process case, properly done multi-process data flow pipelines don't necessarily incur extra copies. Sure, switching to a different process is somewhat more expensive than switching to a different thread due to page table changes, but if you're doing bulk data processing, you amortize any process-separation-driven costs across lots of compute anyway --- and in a many-core world, you can run different parts of your system on different cores anyway and get away with not paying context-switch costs at all.
Also, 10% is actually a pretty modest price to pay for increased software robustness and modularity. We're paying more than that for speculative execution vulnerability anyway. Do you run your fancy low-level audio processing pipeline with "mitigations=off" in /proc/cmdline?
it's a completely crazy price to pay in a field where people routinely spend thousands of $$$ for <5% improvement
> Do you run your fancy low-level audio processing pipeline with "mitigations=off" in /proc/cmdline?
obviously yes! along with power saving CPU C-states or anything throttling-related disabled, specific real-time IRQ and threading configuration (e.g. making sure that the sound card interrupts aren't going to happen on a core handling network interrupts) and two dozen other optimizations (which do make a difference, I regularly set-up new machines from scratch for shows, art installations, etc. and always do this setup step-by-step to see if things are finally "good enough" and they always make a difference, in really a make-or-break sense).
[1] https://www.phoronix.com/news/Google-User-Thread-Futex-Swap
What is really needed, is sane memory model where you can easily call any function with buffers (pointer + size) and it is allowed to access only these buffers and nothing else(note). Not this mess coming from C where this is difficult by design.
(note)since HN likes to split hairs: except for its private storage and other well thought exceptions
I don't understand why this idea keeps failing to take hold even though it's constantly reintroduced in various forms. Surely now, 30 years after that paper was published, we can bear the "slightly increased execution time for distrusted modules" in return for (as the paper suggests) faster communication between isolated modules?
You could probably do it pretty decently in C via `pkey_mprotect` (probably with `dlmopen`).
Emacs wasn't Eight Megabytes and Constantly Swapping only due to Elisp.
No modern processor architecture has a proper message passing mechanism. All of them expect you to use interruptions; with it's inherent problems of losing cache, disrupting pipelines, and well, interrupting your process flow.
All the modern architectures are also so close to have a proper message passing mechanism that it's unsettling. You actually need this to have uniform memory in a multi-core CPU. They have all the mechanisms for zero copy sharing of memory, enforcing coherence, atomicity, etc. AFAIK, they just lack a userspace mechanism to signal other processes.
[1]: https://lore.kernel.org/lkml/20200722234538.166697-1-posk@po...
[2]: https://www.phoronix.com/news/Google-User-Thread-Futex-Swap
It does take a lot of unnecessary stuff out of the way. But isn't enough to change the picture at all.
What points did your comment make? You didn't define "signalling" specifically enough to discuss. Can you elaborate on precisely what kind of "signaling" primitive processors or operating systems should provide?
Process-level/address-space-level dependency sharing remains both easier to think about and simpler to implement (and capabilities are taking bites out of the security risks entailed by this model as time goes on).
Another approach would be to leverage a language like rust. I’d love it if rust provided a way to deny any sensitive access to part of my dependency tree. I want to pull a library but deny it the ability to run unsafe code or make any syscall (or maybe, make syscalls but I’ll whitelist what it’s allowed to call). Restrictions should be transitive to all of that library’s dependencies (optionally with further restrictions).
Both of these approaches would stop the library from doing untoward things. Way more so than you’d get running the library in a separate process.
It featured software-isolated processes that communicated via contract-based message passing, which allowed for zero-copy exchange of data.
As a research OS it never became a fully-fledged OS[2], but an interesting attempt IMHO.
[1]: https://www.microsoft.com/en-us/research/wp-content/uploads/...
[2]: https://en.wikipedia.org/wiki/Singularity_%28operating_syste...
I missed it, can you add a link or at least the post title I can search for?
Could you paste a link to the story? I haven't been able to find it through search engines, and I'd love to read the rationale of such idea...
What does it do?
This allows textures shaders and generally large amounts of data to skip being copied to and from the virtqueues, which is the usual method of virtio communication.
So to answer your question, if you use the Vulkan API on a guest to for example query the available Vulkan devices, if the correct mesa library is installed and virtio-gpu Venus is available, you will be able to use resources on the host with the Vulkan API.