Windows X86-64 System Call Table
j00ru.vexillium.org
j00ru.vexillium.org
For example, NT has no native 32-bit ABI like Linux does. Instead, userspace translates 32-bit to 64-bit system calls, handling the switch from long and back to long mode transparently. This elegant approach is possible because applications never make system calls directly.
>It sounds like bypassing nt.dll might be an exploit vector that is easily forgotten.
MSFT might be bad, but not that bad.
- the 64 bit binaries are stored in C:\Windows\System32
- the 32 bit binaries are stored in C:\Windows\SysWOW64
But even beyond that, I was thinking of one thing they added for performance reasons (see 3rd bullet): https://github.com/Microsoft/WSL/issues/873#issuecomment-425...
However I'm not entirely sure if it's a concrete "syscall" or if it's implemented as a new option through an existing one. (It's not clear to me how they could reasonably fit this into anything that already exists, but it might.) Though I would be mildly surprised if they didn't implement at least one syscall specifically for WSL...
Sure, but not in the NT system call table.
> But even beyond that, I was thinking of one thing they added for performance reasons (see 3rd bullet): https://github.com/Microsoft/WSL/issues/873#issuecomment-425....
> However I'm not entirely sure if it's a concrete "syscall" or if it's implemented as a new option through an existing one. (It's not clear to me how they could reasonably fit this into anything that already exists, but it might.) Though I would be mildly surprised if they didn't implement at least one syscall specifically for WSL...
It's neither, it's an API inside kernel space. They didn't expose it through the NT syscall table, but instead how the linux transtalion layer calls directly into NT kernel functions.
WSL is now a VM. So its really hypervisor that runs the linux syscalls on its virtual processor
By not calling syscalls based on their numbers on Windows. That might sound crazy, but actually to me coming from Windows, calling them by ordinals is what sounds crazy. :-) You're supposed to call exported functions in shared libraries (ntdll.dll etc.) that call the syscalls for you.
It's part of linux saying 'we don't break userspace' and the syscall numbers being how userspace talks to kernelspace.
It seems to me this also means Linux is more limited in what it can do in breaking syscalls too. On Windows, if you load any shared library (including ntdll) and an exported function isn't there, the loader can produce an error for you telling you that a function you need isn't there. On Linux... you're calling the syscall directly, so there's nothing that can stand between you and the syscalls to perform a check or anything. Overall it seems to me like coding against direct syscall numbers is a disadvantage on almost every front.
In the second case, both systems behave the same - if your program runs on an older kernel, the function entry point exists (either the syscall or ntdll symbol) but when you call the function to request the new feature, it returns a runtime error.
In the first case, Linux behaves identically to the first case - you get a runtime error (ENOSYS) - whereas Windows will produce an error at load time.
Typically what is done is that whatever is wrapping the syscall falls back to an alternate implementation for older kernels, if that's possible. If it's not, the program can either continue without that feature if it wants, or fail with a message that a newer kernel is required.
I don't feel that there's really a significant difference in the two approaches - wherever you draw the backwards-compatibility line you're going to have to do the same sort of work to maintain it.
Unless you want to use any other library on the OS, for example 3D graphics drivers which tend to have hardware-specific parts in userspace.
> Overall it seems to me like coding against direct syscall numbers is a disadvantage on almost every front.
Which is why noone does it unless they have a very good reason.
Not sure what you mean regarding this and glibc.
> Which is why noone does it unless they have a very good reason.
I think you changed the meaning of my sentence to be able to make this reply. I was referring to what Linux does re: syscalls vs. library exports, not hard-coding syscall numbers in the source code.
glibc is the GNU C library and wraps syscalls, among other things.
glib is a bunch of cross platform interfaces and data structures that grew out of Gtk+.
In the Linux world, since syscalls are exposed by libc, other non-C runtimes (for example, Go) end up needing to make system calls on their own without the libc wrappers. Is such a system just not a thing in Windows?
Yeah, it's not a thing. You just import functions from ntdll the same way you import from anything else. The OS program loader reads your program's import table and loads them for you.
> does ntdll expose any other functionality? Functionality that may want to be overridden by a different implementation, etc?
It exposes lots of other stuff (DLL loading, memory heap, etc.) that you may want to do differently, but you can hook those at runtime if you really want to. It's not as convenient as just linking another library, but on the other hand, it makes it easier to replace them at runtime instead of at compile-time -- say, if you want a shared library to override a syscall.
Linux is pretty much the only operating system that considers the actual system call ABI to be stable. The BSDs and OS X all want you to use the platform libc instead to access system calls, while Windows uses kernel32.dll et al. as its stable interface.
What about MS-DOS? :P
No, making system call interposition harder isn't a security feature. A user who can LD_PRELOAD you can already do arbitrary things to your program.
Can you elaborate on this? I'm aware of things like LD_PRELOAD, but if an adversary controls the environment, he could just change PATH to point to a rooted version of chrome anyway. That also has to do with the capabilities of the system linker, not glibc.
>the LGPL effectively forbids many projects from using static linking as an escape hatch
The LGPL explicitly allows static linking, that's the main difference versus the GPL.
LGPL allows dynamic linking. Only way it'll allow static is if your releases are accompanied by tools for decompiling and recompiling your binaries with the LGPL bits interchanged. But that actually might not be allowed either, since GCC 4.3+ headers and runtimes (e.g. libstdc++) kind of prohibit you from changing binaries on your own, after they've been compiled.
Can you elaborate on this?
The difference is that calls between dynamic libraries are on the same side of the "airtight hatchway" (as explained by Raymond Chen at https://devblogs.microsoft.com/oldnewthing/20060508-22/?p=31...), while there's a security boundary at the network connection.
What specific threat is bypassing libc supposed to protect against? I don't think you have an argument here. As the poster to whom you're replying mentioned, anyone who can do symbol interposition already has full control over your program.
I consider symbolic interposition to be one of those things. Linux users might not feel comfortable about the fact that glibc currently makes it so easy to intercept system calls to Linux Kernel's RNG, that your upstream dependencies might actually compromise your key generator unintentionally.
Don't you think that's worth discussing? We could also talk about the concerns surrounding Layered Service Providers, which is another great example of userspace libraries misrepresenting APIs that are generally believed to be talking to the operating system.
Users get to control how programs execute. They can interpose symbols. They can disassemble programs. They can run programs under a debugger. They can modify the kernel. They can run programs in a VM. A program can't detect this intermediation; nor does it have any right to do so. A program has no business breaking random OS visibility and control features. We generally call the ones that try malware.
Bypassing libc when making system calls doesn't give a program any insurance against environmental changes. It just inconveniences users while providing no "safety" guarantees. If you want full control over a program's execution environment, ship an appliance.
syscall(SYS_getcpu, &cpu, &node, NULL);AFAIK, there is: Go likes using tiny stacks, while libc expects normal-sized stacks. Switching to a reasonably-sized stack and back has a cost, which the Go developers want to avoid.
They didn't need to, they wanted to, and it can break at any moment on any platform other than Linux. In fact a few versions back they finally back-pedalled and started going through libSystem on OSX.
> Is such a system just not a thing in Windows?
It's not a thing anywhere other than Linux. Go's developers were told time and again that they were supposed to dynamically link to and go through the platform's libc on OSX and BSDs.
They finally relented on OSX after Go broke multiple times during the Sierra beta, but IIRC that's not the case yet for other BSDs (I was thinking it was also the reason why Go 1.11 is required on OpenBSD 6.4, it is but not entirely: OpenBSD requires that stack memory be mapped with MAP_STACK or syscalls will terminate the calling process, so that's an artefact of Go doing userland stacks / threads rather than doing raw syscalls).
But generally speaking, yeah, the stable interfaces are the C ones, exposed through system DLLs. Linux is unique in developing libc and kernel separately and the division between the two as a public abi boundary.
Also, Microsoft didn't necessarily document everything, but they did provide .lib files (and often headers) even for "undocumented" APIs in these libraries. Breaking these APIs would break programs they previously provided SDKs for.
The reality is, so many third-party applications depend directly on ntdll APIs (even Chrome) and some of them literally cannot work with kernel32 stuff (boot-time partition managers, for one), so as far as I'm concerned, it's about as much set in stone as any library could be.
In my experience a lot of the NT APIs are nicer, better thought out, more direct.
Of course this all ignores the fact that Win32 processes are heavier duty than Linux processes (though I don't know if that's due to Win32 subsystem overhead or not). Look at any benchmark and you see an order of magnitude difference in process creation times. You're much better off creating threads instead.
register char (*(*(*(*ram)[512])[512])[512])[4096] asm(cr3)
Your CPU resolves every pointer memory access through those tables under the hood. It's an extremely powerful data structure. It can let you allocate linear memory like a central banker prints money. But if you're a Windows user, then only Microsoft is authorized to access it, and they don't want you having the central banker privileges that are needed in order to implement fork(). That's the way the cookie crumbles.Not really. Only on Linux are direct syscalls an official ABI. On most unices (e.g. solaris, BSDs, OSX) you must dynamically link and go through the platform's libc.
That doesn't mean the interface isn't stable or at least not less stable than the kernel interface or literally any iterface for something that is still being developed.
It would be nice if glibc would provide an easy way to target the ABI of an older version (which is compatible with all newer versions) but you can do that the manual way by compiling and linking against that old version.
Lots of folks don't realize how important that is. If the Linux kernel space were to break as often as the Linux desktop, Linux' market share on the server and in the embedded space would likely barely exceed that of TempleOS.
There are many more lower level libraries that provide binary backwards compatibility: libasound, libX11, libGL, ...
Sure you're not going to be able to assume that /usr/lib/libpng.so.1 will provide the functionality that you want but that isn't any different from png1.dll some program installed in system32.
The real reason Linux's ABI boundary is the kernel is just that the kernel and glibc are sperate and frequently-antagonistic projects. Linux ships the org chart.
The system call number determines what table is used. Calls in the 0x0-0xFFF range are handled by the first table, 0x1000-0x1FFF by the second, and so on.
As far as I know, the ability to add another system call table was only ever used by the IIS web server kernel-mode component (spud.sys).
Starting IIS 4.0, Microsoft has added a kernel mode support driver (SPUD.SYS). This driver also calls KeAddSystemServiceTable function to add its own system services. This fills an entry in third array element of KeServiceDescriptorTableShadow. Hence, its services will start from 0x3000.
[0] https://community.osr.com/discussion/20626/system-service-di...
> As you’ll see in Chapter 6, each thread has a pointer to its system service table. Windows has two built-in system service tables, but up to four are supported. The system service dispatcher determines which table contains the requested service by interpreting a 2-bit field in the 32-bit system service number as a table index. The low 12 bits of the system service number serve as the index into the table specified by the table index.
[...]
> A primary default array table, KeServiceDescriptorTable, defines the core executive system services implemented in Ntosrknl.exe. The other table array, KeServiceDescriptorTableShadow, includes the Windows USER and GDI services implemented in the kernel-mode part of the Windows subsystem, Win32k.sys. The first time a Windows thread calls a Windows USER or GDI service, the address of the thread’s system service table is changed to point to a table that includes the Windows USER and GDI services. The KeAddSystemServiceTable function allows Win32k.sys and other device drivers to add system service tables. If you install Internet Information Services (IIS) on Windows 2000, its support driver (Spud.sys) upon loading defines an additional service table, leaving only one left for definition by third parties. With the exception of the Win32k.sys service table, a service table added with KeAddSystemServiceTable is copied into both the KeServiceDescriptorTable array and the KeServiceDescriptorTableShadow array. Windows supports the addition of only two system service tables beyond the core and Win32 tables.
Edit: I meant SP0, not SP1, sorry. It was this SysCall: NtListTransactions
Also you can't implement TxF in userspace. It has to detect conflicts with other applications and roll back in the case of an unsuccessful commit (power loss etc.) before the file system is used again. Any userspace implementation would leave stuff in a corrupted state until it's re-run.
I don't either, although the API is kind of overkill for most use cases so I'm not too surprised they discourage people from using it.
> And there's more to transactions than just the file system (TxF) so I'm not even sure it's related to this either.
True, I just assumed it was related to the txfs_list_transactions ioctl.
> Also you can't implement TxF in userspace. It has to detect conflicts with other applications and […]
I think the Kernel Transaction Manager already takes care of that. I think??? TxF could be implemented as a userspace library on top of KTM, but I'm not particularly familiar with either facility. Though if it was possible perhaps they would've done it that way in the first place, since TxF uses KTM regardless.
I wonder what did happen to it between SP0 and SP1.
Leading up to any Windows release, Microsoft has a whole bunch of teams working on different features. Some of those features make it into the release, other features don't make it and get cut. What sometimes happens, is that a feature makes it in, but then problems are discovered in testing, and it gets pulled out again at the last minute. When a feature is removed late in the game like that, it is desired to take the lowest risk removal mechanism as possible–one way of doing that is to actually leave the feature in the code, but leave it undocumented, and hope nobody discovers it and uses it (an approach avoided nowadays due to the risk of exploitable security bugs in the buggy API, but in the past it was more common). Another way is to leave its APIs in-place, and either stub out their implementations (to always return an error), or even hide the actual implementation behind a #define that is turned off in the shipping copy. That way, you avoid making any changes to DLL export tables, the system call table, etc., which you worry (even just out of an abundance of caution) might have some downstream negative effect, but also don't have to worry about anyone discovering the broken feature and trying to use it. Then, in the next release, you have a choice – sometimes they will fix the issues with the feature and put it back in, other times priorities have changed and the remnants of the feature (such as API stubs) get removed altogether. That is why sometimes Windows DLLs export undocumented APIs that don't do anything.
(I've never worked for Microsoft, so this is not based on any internal info, just an inference from observing how Microsoft and other vendors do things.)
man syscalls \
| grep -Po '\w+(?=\(2\))' \
| sort \
| uniq \
| wc -l
gives the same number of 481 as the number of rows in that table. Though, the grep includes man-pages that don't correspond with a syscall, like intro(2). It's just a curious coincidence.I can click "Show All" and "Hide All" over and over and there is hardly any delay. Maybe you should upgrade to the latest and see if it makes a difference.
UPDATE: At least 3 of the culprits seem to be MutationObservers from my own extensions (as in, from extensions that I have written for myself) that have {childList: true, subtree: true}! I guess I will have to optimize them. :-)
[1] https://queue.taskcluster.net/v1/task/Uht_zkkPThu384OAHNPwdA...
After doing a number of optimizations, our 2000 row table took 1.2 seconds for chrome to insert the elements (purely the appendChild call). That was down from around 2.6 before trimming down the HTML.
Once layout is involved even without looping to set innerHTML it becomes a lot slower.
Why is this even news? Doesn't Microsoft document this stuff? How does anybody use a system that doesn't have this documented?
I would love to read up on windows internals, but there are no real resources out there. A few books that at first glance appears to cover it, but then just deals with the application layer and up to point-and-click.
This is why I don't use windows anymore. I don't ever again want a job where I have to debug why old program P stopped working after an update. I prefer to be able to learn stuff.
https://www.amazon.co.uk/Windows-Internals-Part-architecture...
Certainly all systems have undocumented stuff that are left undocumented because the users are not supposed to use it. If they would, they couldn't change it.
They all have files that end up as /usr/include/sys/syscall.h and define well-known numerical constants for each (documented!) system call.
https://github.com/openbsd/src/blob/master/sys/sys/syscall.h
https://github.com/lattera/freebsd/blob/master/sys/sys/sysca...
http://cvsweb.netbsd.org/bsdweb.cgi/src/sys/sys/syscall.h?re...
Oh, look! They even share the exact same numerical values for the old school system calls. #1 is exit(), #2 is fork(), so on and so forth.
Gosh, on my linux laptop, here in /usr/include/asm/unistd_64.h is pretty much the same numbers, but objectively hidden a little bit more than in the direct Unix descendants. I interpret this as "less documented than in Unixes".
Wow, even Minix source code has the same numbers: https://minixnitc.github.io/posix.html
I wonder why that is? Maybe it's because the system call numbers are in fact documented, part of the POSIX API, and descend from Bell Labs Unix source code. Here's V7's list of system call numbers as proof:
https://minnie.tuhs.org/cgi-bin/utree.pl?file=V7/usr/include...
e.g.
OpenBSD deletes "relic from the past" syscall:
https://cvsweb.openbsd.org/src/sys/sys/syscall.h?rev=1.199&c...
Solaris documenting deleted syscalls (note however the libc interface is guaruaneed):
https://docs.oracle.com/cd/E23824_01/html/E22973/gkzlf.html
MacOS X is so unstable that even go developers had to go through libSystem:
Just to be clear to anyone else, the Mac OS X syscall interface is guaranteed to be unstable by Apple. Going through libSystem is the only way.
Please educate yourself about the Windows architecture.
Compare that with the man pages of some decent BSD. Or even Linux..
But yes. I have contemplated buying one of those windows internals. I will probably buy that book just because of your comment. It's not that expensive.