HNHacker News
TopNewBestAskShowJobs

markjdb

179 karma · joined August 15, 2013

submissionscomments
markjdb··on Gathering Linux Syscall Numbers in a C Table
At least FreeBSD's syscall ABI is guaranteed to be stable, one can run ancient binaries on a modern kernel. I believe the same is not true of OpenBSD and maybe NetBSD however.
markjdb··on Moving from OpenBSD to FreeBSD for firewalls
The pf maintainer in FreeBSD has been doing a ton of work to bring more recent improvements over from OpenBSD, trying to bring them in sync as much as possible without breaking compatibility: https://cgit.freebsd.org/src/log/sys/netpfil/pf

The state of affairs you described is much less the case now than in the past.

markjdb··on BSD kqueue is a mountain of technical debt (2021)
A design that works by default isn't automatically better either though. You have to look at the details.

> I guess one might be tempted to forgive a few warts in the interface layers

... well, yeah, that's exactly my sentiment about kqueue here. What you're talking about is basically a small wart that no one's bothered to address because it's inconsequential.

markjdb··on BSD kqueue is a mountain of technical debt (2021)
The article clearly isn't talking about technical debt within the kernel implementations of epoll and kqueue, and if one wanted, it'd be easy to define fallback EVFILT_READ/WRITE filters using a device's poll implementation.

I don't really understand what argument you're making. Is io_uring also a bad design because it requires new file_operations?

markjdb··on BSD kqueue is a mountain of technical debt (2021)
Well, no, it's "this interface works fine for you if you implement it."

The kernel doesn't magically know whether your device file has data available to read, your device file has to define what that means. That's all I'm referring to. Hooking that up to kqueue usually involves writing 5-10 lines of code.

markjdb··on BSD kqueue is a mountain of technical debt (2021)
I'm not sure. Maybe it's "wait for events that aren't tied to an fd."

For instance, FreeBSD (and I think other BSDs) also have EVFILT_PROC, which lets you monitor a PID (not an fd) for events. One such event is NOTE_FORK, i.e., the monitored process just forked. Can you wait for such events with epoll? I'm not sure.

More generally, suppose you wanted to automatically start watching all descendants of the process for events as well. If I was required to use a separate fd to monitor that child process, then upon getting the fork event I'd have to somehow obtain an fd for that child process and then tell epoll about it, and in that window I may have missed the child process forking off a grandchild.

I'm not sure how to solve this kind of problem in the epoll world. I guess you could introduce a new fd type which represents a process and all of its descendants, and define some semantics for how epoll reports events on that fd type. In FreeBSD we can just have a dedicated EVFILT_PROC filter, no need for a new fd type. I'm not really sure whether that's better or worse.

markjdb··on BSD kqueue is a mountain of technical debt (2021)
I don't think the article does a good job of arguing its premise, which I think is that kqueue is a less general interface than epoll.

When adding a new descriptor type, one can define semantics for existing filters (e.g., EVFILT_READ) as one pleases.

To give an example, FreeBSD has EVFILT_PROCDESC to watch for events on process descriptors, which are basically analogous to pidfds. Right now, using that filter kevent() can tell you that the process referenced by a procdesc has exited. That could have been defined using the EVFILT_READ filter instead of or in addition to adding EVFILT_PROCDESC. There was no specific need to introduce EVFILT_PROCDESC, except that the implementor presumably wanted to leave space to add additional event types, and it seemed cleaner to introduce a new EVFILT_PROCDESC filter. Process descriptors don't implement EVFILT_READ today, but there's no reason they can't.

So if one wants to define a new event type using kevent(), one has the option of adding a new definition (new filter, new note type for an existing filter, etc.), or adding a new type of file descriptor which implements EVFILT_READ and other "standard" filters. kqueue doesn't really constrain you either way.

In FreeBSD, most of the filter types correspond to non-fd-based events. But nothing stops one from adding new fd types for similar purposes. For instance, we have both EVFILT_TIMER (a non-fd event filter) and timerfd (which implements EVFILT_READ and in particular didn't need a new filter). Both are roughly equivalent; the latter behaves more like a regular file descriptor from kqueue's perspective, which might be better, but it'd be nice to see an example illustrating how.

One could argue that the simultaneous existence of timerfds and EVFILT_TIMER is technical debt, but that's not really kqueue's fault. EVFILT_TIMER has been around forever, and timerfd was added to improve Linux compatibility.

So, I think the article is misguided. In particular, the claim that "any time you want kqueue to do something new, you have to add a new type of event filter" is just wrong. I'm not arguing that there isn't technical debt here, but it's not really because of kqueue's design.

markjdb··on BSD kqueue is a mountain of technical debt (2021)
The same is true of kqueue/kevent though... the driver just needs to decide which filters it wants to implement. There's no need to extend kqueue when adding some custom driver or fd type. One just needs to define some semantics for the existing filters.
markjdb··on Why laptop support, why now: FreeBSD's strategic move toward broader adoption
For what it's worth, the default root shell is now /bin/sh instead of csh. I think that's true as of 14.0. /bin/sh is also a better interactive shell than it used to be, though yeah, I don't use it to do anything other than install my preferred shell.
markjdb··on Early performance results from the prototype CHERI ARM Morello microarchitecture
CHERI does more than help eliminate security vulnerabilities. Consider that today we rely on the MMU to provide memory isolation between Unix processes; CHERI enables isolation without switching page tables, at a smaller hardware cost (though it's not like you can drop unmodified software into such an architecture). So I don't think it's correct to consider this yet another layer of complexity. If anything it has the potential to lead to simpler system designs.
markjdb··on The Tower of Weakenings: Memory Models for Everyone
CHERI does permit tricks like storing flags in the low bits of a pointer, at least to some extent. Quite a lot of low level C code (including some in the CheriBSD kernel) needs that to work.
markjdb··on eBPF for tracing how Firefox uses page faults to load libraries
I use "ktrace -t f" once in a while for debugging and it's really handy. Output looks like

78436 cat PFLT 0x6c71f99cda8 0x2<VM_PROT_WRITE> 78436 cat PRET KERN_SUCCESS 78436 cat PFLT 0x3c6efd36c280 0x2<VM_PROT_WRITE> 78436 cat PRET KERN_SUCCESS 78436 cat PFLT 0x3c6efd36e158 0x2<VM_PROT_WRITE> 78436 cat PRET KERN_SUCCESS ...

Obviously not nearly as flexible as ebpf though. For instance it'll log all page faults happening in the context of the process, and so includes page faults that happen in the kernel due to copyin()/copyout() etc. Sometimes it's helpful and other times confusing.

markjdb··on Benchmarks: FreeBSD 13 vs. NetBSD 9.2 vs. OpenBSD 7 vs. DragonFlyBSD 6 vs. Linux
I'd be amazed if it isn't a configuration error of some kind.
markjdb··on Serving Netflix Video at 400Gb/s on FreeBSD [pdf]
How much data ends up being served from RAM? I had the impression that it was negligible and that the page cache was mostly used for file metadata and infrequently accessed data.
markjdb··on Fun with Unix domain sockets
You can even send a Unix socket over itself. The kernel has to be careful to handle that correctly. :)
markjdb··on Implement unprivileged chroot
This is mostly true on FreeBSD as well. The real problem is that capability mode also disallows openat(AT_FDCWD) - there has to be an explicit directory descriptor.
markjdb··on CVE-2021-22555: Turning \x00\x00 into 10000$
It depends on the bug. syzkaller does an excellent job finding race conditions, but it can be difficult to generate a reliable reproducer for them. It often succeeds nonetheless. In other cases there can be a wide gap between the proximate and root causes of a crash. For instance some system call bug might corrupt memory in a way that only results in a crash some time later, when some asynchronous task runs, in which case it's also difficult to find a reproducer. Sanitizers can help identify such bugs earlier and so reduce the amount of manual analysis needed in the absence of a reproducer.

I'm not sure what happened in this case. The linked report does indeed have an associated reproducer.

markjdb··on Buffer overruns, license violations, and bad code: FreeBSD 13’s close call
That's fair. At the time, though, it wasn't clear that the author's angle was that FreeBSD doesn't have a culture of doing code reviews. We do, and I don't think one has to look very hard to see that.
markjdb··on Buffer overruns, license violations, and bad code: FreeBSD 13’s close call
There's some discussion happening now and I do expect to see some process changes coming out of this. It's tricky. The review you link does nominally follow the process of creating a review and having some discussion, but there isn't much actual code review happening there.
markjdb··on Buffer overruns, license violations, and bad code: FreeBSD 13’s close call
> Or is this just how it is on FreeBSD?

It's not. We do a lot of code review, and it's done publicly. It's easy to look at the commit logs. I find it telling that the article doesn't spend even one word trying to delve into our code review practices. This case was an aberration.

markjdb··on Exploring Swap on FreeBSD
The notion there is that at some point in the past free memory was scarce, so the kernel swapped out some pages, and that swap space may still be in use long after the shortage is alleviated. FreeBSD won't swap anything out unless there's a shortage of free pages.
markjdb··on Running BSDs on AMD Ryzen 5000 Series – FreeBSD/Linux Benchmarks
> as many of these tests show

The tests appear to compare ZFS and ext4 and clang and gcc as much as FreeBSD and Linux.

markjdb··on OpenZFS Merged to FreeBSD
"paravirtualized Solaris kernel" isn't really accurate - it's a collection of kernel interface shims. The whole thing is quite small, about 4KLOC on FreeBSD.
markjdb··on I'm back into the grind of FreeBSD's wireless stack and 802.11ac
A driver is in ports while the author works on getting it ready to import. Just pkg install iichid.
markjdb··on FreeBSD 11.4
Also mentioned here: https://lwn.net/Articles/808733/
markjdb··on MMU gang wars: the TLB drive-by shootdown
I was wondering how this scheme might break applications that do clever things with a SIGBUS/SIGSEGV handler. Do you happen to know of any OSS that actually does something like this?
markjdb··on MMU gang wars: the TLB drive-by shootdown
> I do wonder why there isn't an API for "lazy munmap()"

You don't really need a separate API. The kernel can implement munmap() lazily, it just needs to also ensure that the address range isn't reused until a TLB shootdown is completed. LATR is a system that leverages this to implement lazy TLB shootdowns for munmap(): http://www.cs.yale.edu/homes/abhishek/kumar-asplos18.pdf

markjdb··on PID Without a PhD (2016) [pdf]
If anyone is interested in their application to system software, the FreeBSD kernel uses a PID loop to regulate memory reclamation: https://svnweb.freebsd.org/base/head/sys/kern/subr_pidctrl.c... https://svnweb.freebsd.org/base/head/sys/vm/vm_pageout.c?vie...

It does a pretty good job of maintaining system responsiveness and latency when there's sustained memory pressure, at least much better than the simpler hysteresis loops that are commonly used for this sort of thing.

markjdb··on ThinkPad T480 is my new main laptop which runs FreeBSD
Because there are not many people with the ability to add 802.11ac support to drivers, the resources to do so (datasheets, test hardware), and the time required. The 802.11 stack does support at least some 11ac features.
markjdb··on Linus: Don't Use ZFS
It implements the x86 and x86_64 Linux system call ABI. Linux ELF binaries get vectored to an alternate system call table implemented by the compatibility layer. There are some other components like an implementation of a Linux-compatible procfs. How well it works in practice really depends on how far off the beaten path you go. There are lots of non-essential pieces that are not implemented, but for example I know of people running Steam on FreeBSD.
Page 1 of 2Next →