HNHacker News
TopNewBestAskShowJobs

touisteur

1,792 karma · joined March 27, 2017

submissionscomments
touisteur··on What Zig felt like, coming from Rust
I miss the years of writing CLI tools, web servers and clients, in Ada (and of course, real-time complex distributed system...). Felt so simple and right and fast and robust. The code is still readable today and maintaining it is a zero effort today. Clean Java without the enterprise BS was a close second in ease of programming - boilerplate be damned.

I'm glad NVIDIA found a way to make GPUs programmable and got us out of the shaders tarpit, but did it have to be C++...

touisteur··on German Rheinmetall open-sources its Battlesuite connected weapon system protcol
Asterix is an actual standard for radar data, coming from civilian radar systems. I wish it was updated somehow for 3d beam-steered (and moving) sensors, but still it covers most of the industry's needs... I keep hearing so many grumblings of replacing it (it is binary, unforgiving, quirky, very 80s...) and of course fragmentation is back in.
touisteur··on JetKVM Mini
Just in case, and for people who actually do cable management...
touisteur··on JetKVM Mini
I was thinking the point was it was so light to have it hanging on the cables themselves, as I'm certain absolutely no one does in any of their installations...
touisteur··on Use Vsock with Libzmq
Building micromvs with no network support helps with isolation and (more) kernel surface reduction.
touisteur··on Use Vsock with Libzmq
Been hoping for some time someone would write it better than my local ugly hack. This immediately useful. Thanks a bunch.
touisteur··on Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics
Thanks. I was just wondering, since I think I saw NXP had Arm SKUs with Mali GPU and had some hope someone here would Cunningham's Law me.
touisteur··on Arm Mali G2-Ultra NX GPU: desktop-class mobile gameplay with AI-native graphics
Any foundry in Europe able or going to manufacture a chip with this on ?
touisteur··on A practical guide to running 8x RTX PRO 6000's
Yes that was the second part of my comment. I've been benchmarking many things AI or not on those boards and while one can sometimes get to 600W I haven't seen more than 10% gain for those last 250W. If someone has a workload that gets more from those watts I'd be interested. For now on anything I run (full-capacity, continuous, batched...) capping at 350W seems better in bang-for-bucks when including energy costs (from the GPU + added cooling).
touisteur··on AI handles incidents, engineers lose touch with their systems
For more in this vein, look up "Automation should be like Iron Man, not like Ultron". Sad to see so many people let go of their agency.
touisteur··on A practical guide to running 8x RTX PRO 6000's
I'm curious whether actual inference workloads actually push to 600W (and not 350W) and what the last 250W get you. Rare is the (generic gpu) workload where I get >5%, some rare light inference benchmarks up to 10%...
touisteur··on Samsung's Processing-in-Memory (PIM)
Data movement and local operations are still bottlenecked today on memory bandwidth. Butterfly primitives, sorting/fft/1D-convolution, the whole cub library, could be ported there and have great performance wins. But the pain of programming and maintaining code using this...
touisteur··on Doctors are finally learning to manage antidepressant withdrawal
Venlafaxine has also been a godsend for me. My doctor started me on ut first and past the two weeks of rebalancing (big ideation swings and so much calling to act on it...) and all kinds of pains (digestive mostly), it literally allowed me to smile again, which my wife realized hadn't happened in over a year.

I can relate to the feeling when missing a dose. No tolerance for being forgetful. Lots of anxiety brought by this tether... and the permanent anorgasmia (in my case) is sad but I was dying...

Sadly I built tolerance over 2 years and had to stop, taper off (horrible experience) and try others (most that were cited in this comment section). All had painful and very annoying side effects and tapering was a pain after trying and each time feeling so disappointed and tired. Most worked on reducing ideation but the costs (inability to concentrate, brain fog, sometimes inability to work) were too high.

Turns out tolerance to venlafaxine ebbs out after some time, so after a moment I was able to restart for a year until I didn't need antidepressants anymore.

I've learned reading forums of the gamut of experiences with antidepressants and just want to add a voice to 'they can work, great even, not perfect' and between dying and trying them (with clear open eyes about the tradeoffs) I'd advise trying.

touisteur··on Bun 1.4 Rust rewrite is not looking good?
In a previous life I managed to deliver a system with (basic-block) coverage measurement compiled in (gcov) but instead of the mess of files gcov generates, the whole coverage structure was streamed to a remote server (it compressed very well). Once that was in place it was an amazing telemetry tool.

Later on used Intel Processor Trace in a similar fashion for even finer (mc-dc) coverage.

Coverage tools are very useful, if a bit hard to use...

touisteur··on Looking for Missed Alarm Bugs in a Formal Verification Tool
I wish John Regehr was more widely read. The constant grind that improving formal method tools is a somehow unrewarding but worthy calling.

Missing runtime checks are a pain and I remember one that struck me as hard: https://www.adacore.com/blog/the-most-obscure-arithmetic-run...

touisteur··on Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes
rsyslog is an incredible piece of software. Every time I'm looking for something to do with logs, opening the docs or googling finds the feature for me and myriads alternatives. I know it still exists and use it heavily on any system I'm in charge of, but there's some regret at having a dual system with journalctl...
touisteur··on LFM2.5 2.6B model competitive with 4x larger models
Really curious about people's workflows with these agentic-but-not-for-coding workflows. Are there some interesting people to follow there or just good testbeds/environments to get an idea ?
touisteur··on Exploiting System Management Mode with a very long interrupt
I have a reproducible way to have a pwrite syscall on a specific SSD on a specific machine take 15+ seconds and completely block any syscall related to that SSD by any other thread or core during that amount of time. I tried and couldn't preempt it either (sched_fifo and preempt kernel options). I should have a look soon with Intel PT to check whether it's on the same instruction every time :)
touisteur··on Wireblast a 100 Gbps packet generator in Go using AF_XDP
I half-wished I'd get Cunningham's law-ed here. I read a bit more and there seems to be some support for tcp and udp offload, https://netdevconf.info/0x17/sessions/talk/tcp-offload-via-a... but I haven't checked how easy it is to use.

A lot of the socket featureset of io_uring seems available in AF_XDP https://docs.kernel.org/networking/af_xdp.html which shows lots of progress since I looked last.

To get an idea of what DPDK gives low-level access to there is the overview https://doc.dpdk.org/guides/nics/features.html and my "favorite annual terabit read" https://doc.dpdk.org/guides/nics/mlx5.html#mlx5-net-features for NVIDIA NICs. Broadcom has some fun stuff too. The first time you hit top RX speed (2x400G my latest) with only one busy core (yay DMA engines) is always a thrill.

touisteur··on Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
Especially if it can be run as a read-only replica.
touisteur··on Wireblast a 100 Gbps packet generator in Go using AF_XDP
I think access to offload engines is a big part of the appeal of dpdk still, especially for me all the GPUdirect nvidia-only packet steerer.

I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.

Some silly stuff in DPDK are very useful for testing too (in-memory devices).

Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.

touisteur··on DeepSeek V4 Flash on a Single AMD MI300X
I thought MI350P wasn't available yet, curious where to source it right now.
touisteur··on Windows XP 2002 for the Itanium: Unbridled rage
I think the "Jim Keller" story around Zen is a bet on modularity, core-complexes, then chiplets. Smaller, less monolithic designs and a clean re-design of the x86 cores for compacity, ease of validation and scalability (in core count).
touisteur··on Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
Don't see how it could be cheaper with HBM than the RTX 6000 Blackwell Server Pro but in these times of relative shortage any additionnal supply should be slurped ?

I was remarking on the MI350P because I've had a hard time procuring "small" CDNAx systems (for e.g. development, experiments and lower-profile servers) and OAM seemed very niche (not if you're aiming for density and training/inference...).

I hope the MI350P fills a lower part of the spectrum and I can start massively porting CUDA stuff or at least work on HIP and ROCm and what I need to make most or some of our CUDA stuff run on AMD HW, then how make it run fast.

touisteur··on Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
Time for AMD to improve their directstorage story I guess. Is there a good writeup on how one would use it there ?

Same for gpudirect (more useful for scale-out or training).

Everyone is focused on AI but these are two interesting techs trickling down from the NVIDIA tree, relatively "easy" to use there, which maybe exist in AMD world but I somehow missed the docs and APIs and demos on how to use them...

touisteur··on Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
Hopefully the MI350P is available soon, at last a standard PCIe SKU, if a bit too-much previous-generation and gimped compared to the MI350X
touisteur··on Attention Decode on AMD MI450 GPUs: A Gluon Kernel Optimization Guide
As much as I enjoy these articles and for AMD to write more light technical articles, it really feels constrained, even strained, to be unable to cite the equivalent terms from the precursor here (NVIDIA). Another batch of jargon for very similar architectures and programming models... HIP and ROCm have actually made amazing strides in making CUDA developers' porting work easy, and I know playing catchup to a (monopolist) moving target you have no power over is bad... but I feel this is part of the thousand paper cuts.
touisteur··on Why don't people use formal methods? (2019)
Yes it gets hard really fast. We had a fun (if tongue-in-cheek) exploration of this (proving a sort implementation) with Yannick Moy of SPARK fame some time ago https://www.adacore.com/blog/i-cant-believe-that-i-can-prove...

I only regret not writing the obvious-but-buggy code that "forgot" or added some values and still passed proof...

touisteur··on Why don't people use formal methods? (2019)
GOTO is back ! So glad to see CBMC used. I used to write translators to GOTO for simple code checking and was wondering where the recent state of the art was. Thanks for the pointers.

Did you have a look at why3 and generating verification conditions from Rust or C code (as frama-c does) ?

touisteur··on The mean means nothing: data visualization to debug a latency problem
Cool that the Wikipedia page links to the "Mean Dinosaur" paper https://dl.acm.org/doi/10.1145/3025453.3025912 that I love to get out each time someone sends me mean, median, or stddev to measure processing latency. By all means use stats, but always eyeball the dataset to check assumptions extracted from statistics, I guess, especially in this world of matplotlib and notebooks and agents.
Page 1 of 34Next →