2,175 karma · joined May 9, 2017
Also, that was an easily citable doc line, but I think a more comprehensive look at the work done in Fuchsia shows it has or had ambitions to run on personal computers.
e.g.: fancy graphics stack, Linux compatibility layer, ideas about OS integrated cloud syncing, the original work on the Xi text editor.
> Zircon targets modern phones and modern personal computers with fast processors, non-trivial amounts of ram with arbitrary peripherals doing open ended computation.
Fuchsia also had a Linux compatibility layer similar to WSL1 at some point. Might still be there?
[0]: https://fuchsia.dev/fuchsia-src/concepts/kernel/zx_and_lk
Also, a link to the rust root of Zircon in case anyone else was interested: https://fuchsia.googlesource.com/fuchsia/+/refs/heads/main/z...
Totally fair, and I agree that you should only have a few threads (or TPC) and shouldn’t be inundating the allocator with requests. But I’m only grinding this axe because I’ve personally had to deal with a system that went against most of that guidance.
We ran far too many threads in a memory-constrained environment. Thread count was many multiples of core count.
I totally agree that this is “questionable territory”, but honestly, any application that’s outgrown the basic glibc malloc has made a few mistakes.
maybe i give it too much weight, but:
> massive oversubscription of physical CPU cores (ie tens of thousands of threads) where TLS becomes utterly wasteful while also thrashing the cache
seems worth solving to me
[0]: https://github.com/google/tcmalloc/blob/master/docs/gperftoo...
cost for interruption in an rseq critical section is that the PC gets overwritten to the rseq abort entry point before the task is rescheduled. no management thread necessary.
should be fairly minimal cost, especially assuming interruptions in the critical section are rare.
i don’t believe rseq based cpu local caches require memory barriers on the fast path.
to confirm, you were using tcmalloc from https://github.com/google/tcmalloc and not from https://github.com/gperftools/gperftools, right?
the latter is a lot worse iiuc.
this bit seems a bit messy.
for us, tc was among the fastest in runtime while being very space efficient [0]. large rust application using far too many threads.
we’ve since also had great success with tc’s built in profiling tools.
modern GNOME is honestly a more consistently clean interface than even (modern) macOS.
allocators are so simple to just swap into your program. if you can put together a few representative workloads, you should just try out a few allocators and profile whatever metrics you care about.
not entirely sure why, but i’d guess there were durability concerns with the screen wrapping around the edge. might have also felt weird to hold. potentially the crease would be worse in unfolded mode as well.