HNHacker News
TopNewBestAskShowJobs

raphlinus

13,722 karma · joined March 7, 2014

I do research on fundamental UI technology and 2D graphics, with a focus on Rust and fonts. @raph@mastodon.online
submissionscomments
raphlinus··on What TLA+ can and can't check
I disagree with this advice, and consider it outdated.

Arguably, seq_cst is helpful for informal reasoning, because it's not hard to imagine all permutations and interleavings. But in my opinion, nobody should be doing lock-free programming based on informal reasoning. Algorithms should be considered incorrect unless they've been rigorously validated, ideally with formal methods or at least model checking.

Very few lock-free algorithms require sequential consistency. There are exceptions, such as Chase-Lev queues, but they are rare.

The added confidence that seq_cst gives you if the algorithm hasn't been properly validated is IMHO worthless.

raphlinus··on The state of SIMD in Rust in 2026
You've got a point but are overstating it considerably. There is a big gap between just autovectorization and the portable primitives a library like Highway or Fearless SIMD will give you. For example, I haven't seen autovectorization do select or swizzle.

But there's another point in the tradeoff space. One of the explicit design decisions in Fearless SIMD is to support "downcasting," or specialization to a specific microarchitecture. At least for the kind of problems I've worked on, even when you're doing something fancy with arch-specific permutations or what not, the majority of the operations will be pretty vanilla, and can be expressed well in the portable subset.

So you can think of a library like Fearless SIMD as enabling your extreme optimization use case, just more ergonomically.

Of course, this depends on LLVM compiling intrinsics to assembly efficiently. That hasn't always been the case, and is not perfect now (a number of issues have been filed against rustc and LLVM while developing Fearless SIMD), but is pretty good.

As always, though, you do have to measure performance, and I frequently look at the assembler output to double-check that it's doing the right thing. The day of "fire and forget" portable SIMD has not yet arrived.

raphlinus··on Photon-Emission-Guided Laser Fault Injection Enables RP2350 Secure Debug
Yes, four cores in the chip. And yes, there's additional muxing, but I think that adds a fairly small amount of chip area compared with the crossbar. In addition to the two core slots, there are a lot of peripherals contending for single cycle bus access.
raphlinus··on The Farnese letter
I had a sense. Pangram identifies this as AI.
raphlinus··on C++26: Trivial infinite loops are no longer undefined behaviour
You will find the answer you seek not from an LLM, but from the talk Forward Progress Guarantees in C++ by Olivier Giroux at CppNow 2023. It's a long talk, with lots of details about forward progress, but I've set the timestamp[1] to the infinite loop bit.

[1]: https://youtu.be/g9Rgu6YEuqY?si=_l9JwKhjvIdFEDEX&t=3819

raphlinus··on I ran Photoshop on a £0.60 computer chip
I have a Fruit Jam and several other board (including the Olimex). It's great fun and I'm enjoying programming it. The biggest delight is that it's actually possible to reason about performance, extremely difficult on larger computers because of all the complexity.

That said, Photoshop doesn't feel like the sweet spot of what to run on this hardware. There's only 520kB of fast static RAM, and that's not enough to run a framebuffer with enough bit depth to hold a photo reasonably. I have invented a 4bpp compression scheme which opens up possibilities, but it's asymmetrical and in its current form compression is very slow, so it wouldn't be suitable for interactive work.

As a general matter, emulation makes a ton of sense. You can emulate all kinds of computers up to the late 80s, including video subsystems. Of course Mac and PC don't really do anything interesting like tiles and sprites, they're pretty much just dumb framebuffers.

raphlinus··on The mathematical beauty of hyperbezier curves
That was my first attempt at a PhD thesis. The one that actually got me over the line was the second attempt, on spirals and splines. Of course this was all years ago.
raphlinus··on The mathematical beauty of hyperbezier curves
Elastica curves have only three parameters, compared with four for cubic Béziers, Spiro curves, and the new hyperbezier. The easiest way to understand that limitation is to consider the parallel curve of an Euler spiral. That has a built-in asymmetry, one end has higher tension than the other. But the math for elastica locks in odd symmetry around the inflection point.

I think hyperbezier would be a natural fit for boat hulls, but it's not a mathematically precise approximation to elastica either. I'd say to carefully evaluate it, and I'd very much like to hear how that goes.

ETA: "parallel curve of elastica" is an intriguing curve family to consider for this application, as it has the correct number of parameters and ticks a lot of the other boxes. However, the math for this is hard mode.

raphlinus··on The mathematical beauty of hyperbezier curves
This is future work, but I expect it to go fairly well. One encouraging sign is that there are parts of the parameter space where you do get an exact solution: circular arcs and circle involute. Another encouraging sign is that it's easy to get bounds on curvature, which is important for figuring out whether there's a cusp in the offset curve.
raphlinus··on The mathematical beauty of hyperbezier curves
Super good questions.

First, for hair rendering you've got a 3D trajectory, which Béziers can handle just fine, but I'm not sure about these spirals. I do have a chapter in my thesis and there has been a bit of followup from others, but I can't say anything with confidence.

That said, because of the specular reflections, applications like hair really would benefit from higher degrees of continuity. Cem Yuksel has a recent SIGGRAPH paper on how to tweak Béziers to get more continuity. In fact, seeing that was one of the motivations to take this work off the back burner.

I'm not sure about maps. But I do know of applications representing the centerline for autonomous vehicles that use polynomial spiral representations, specifically because of the extremely high degrees of continuity you can attain. For that, you don't have to represent these high-tension curvature regions.

I do think there is a specific application for drawing smooth connecting lines in autogenerated diagrams. These tend to be S-shaped, and with Béziers you tend to get curvature peaks near the endpoints. You're much better off with monotonic curvature and with spirals that's easier to achieve. And I also believe being able to get those squircle/superellipse shapes would help as well.

raphlinus··on The mathematical beauty of hyperbezier curves
There are two separate questions here. One is how much moving control points creates expected changes in the same direction. Béziers nail this, as the position of a point at t is a linear combination of the control points, with the Bernstein polynomials as weighting functions. So it always feels like direct control. With my mapping, you get this for a nice big chunk of the parameter range – small to moderate angles and control point distances. But this property does fall apart when pushing to extremes.

The other question is whether the control is local. As others have remarked, this has much more to do with the way the curve is embedded into a spline than the curve family itself. In particular, it is deeply affected by the continuity constraints. As payment for the local support, cubic Béziers only give G1 continuity. Euler spiral splines, by contrast, are G2 but changes do cause those ripples.

The pen tool prototype linked in the blog post suggests giving designers more choices. If you specify all control points, you get G1 just like a cubic Bézier. But it also gives the choice of specifying one control point on a smooth endpoint, and solving the other for G2 continuity. In my experience, it feels more like local control than Euler spiral-like splines. When you want smoother curves, you do that, and when you're willing to sacrifice continuity for local control, that's also possible.

raphlinus··on The mathematical beauty of hyperbezier curves
I'll try to answer some of these questions.

Yes, there are regions of the parameter space where moving a control point has vanishing effect on the curve shape. I'm not thrilled about this and have explored alternatives, but this is the best I could come up with. Pretty much all approaches based on solving for optimal curve fit have discontinuous "flipping" behavior. Another approach would be to split

We haven't yet done the work wiring this up into a spline like the older pen tool draft. The "auto point" is not a form of subdivision but is a way of achieving G2 continuity.

Yes, hyperbeziers can be split and subdivided without much trouble at all, and being closed under subdivision is an important mathematical property that was lacking in the earlier draft. This follows directly from their formulation as polynomials. If you look into the code, you'll see that there's subdivision at the near-cusp to make the integral robust.

The P/C naming evokes Hermite but there wasn't a lot of thought into it.

Math is heavier, but this is more nuanced than you might expect. My feeling is that computers are fast, so if we can draw higher quality curves with less human effort, it's worth it. But the seeming simplicity of Béziers is deceptive; you often have to solve inverse arc length problems (for example to compute dashing), and curve fitting is really hard (I have a blog post in the pipeline on that). So I think overall the math is only slightly heavier.

The goal is human-driven interactive design, originally motivated by font design. Splitting into more Béziers obviously gives much finer control, but it's hard to avoid them getting lumpy. In most fonts you have a minimum number of Bézier segments and accept whatever fine details the Bézier math gives you. That's even more so when making variable fonts, where the control points are interpolated.

I haven't yet done a rigorous comparison with NURBS. While they're hugely popular in CAD, they're basically unused in things like font design.

Thanks for these questions, they obviously show a deeper understanding.

raphlinus··on The mathematical beauty of hyperbezier curves
As a clarification, the hyperbezier curve can do a loop, but I deliberately chose not to make that accessible in the parameter mapping. I couldn't figure out how to do it with any reasonable continuity. It's possible to flip from the normal shape to the loop, but I wanted to avoid that. And if you don't flip, then all trajectories have to go through an exact corner as an intermediate, which is very different behavior than the cubic Bézier. In the Bézier, the cusp can occur anywhere along the curve, but to make the corner shape it has to be exactly at the intersection of the endpoint tangents.

There are other Bézier shapes that the hyperbezier can't match, including two cusps on the interior. As you say, I consider these not very useful in actual 2D design, and am happy to give them up for smoother curves and the increased range of superellipse shapes.

raphlinus··on The mathematical beauty of hyperbezier curves
Not yet.
raphlinus··on Rust SIMD on the GPU
This is basically the goal of fearless_simd, but of course achieving the same level of maturity will take time.
raphlinus··on Os8088: A powerful Mac-like OS for the IBM XT, 286, 386
I was curious about this and had a look. The code doesn't look much better than what a good C compiler might produce, and I found quite a few opportunities to optimize. It doesn't look much at all like the 8088 code I wrote when I was a teen. Among other things, so much pushing and popping.

Take the irow loop in vga12.inc[1] (of course I'm going to look at the graphics code). Each iteration does push di; rep stosb; pop di; add di, ROW_BYTES (and some other stuff). Why not save the push/pop and add (ROW_BYTES - count), which could be stored in a register (dx is free here)? Just the push+pop is 15+12 cycles.

[1]: https://github.com/jggonz/os8088/blob/1f2fae44180fadaf85368c...

raphlinus··on Show HN: Pokémon Emerald Ported to Raspberry Pi Pico 2
If you're interested in HDMI audio, I have code for that lying around. The current best Rust source is in [1], but I also have a local branch with hand-tuned asm for the terc4 encoding. It's maybe 25% faster, which might be significant, as the per-scanline CPU load does go up a fair amount.

[1]: https://github.com/DusterTheFirst/pico-dvi-rs/blob/main/src/...

raphlinus··on Running a 28.9M parameter LLM on an $8 microcontroller
Entirely fair, $1 is just the chip, not the board. No question the ESP32-S3 is incredibly good value.
raphlinus··on Running a 28.9M parameter LLM on an $8 microcontroller
If you want to do this at the $1 price point, you can on RP2350, albeit with some limitations. In particular, it maxes out at full speed (12Mbps). The trick is to use the on-chip USB peripheral for one, and connect the other to GPIO pins backed by PIO.

This works today with tinyusb and pico-pio-usb, but I'm also playing with a Rust port which I'm hoping will have higher performance.

raphlinus··on A digestion of the Jacobian conjecture counterexample
Actually it's not true for the hyperreals either, though this is a common misconception. It's extremely easy for me to believe that an llm would produce a "proof" and that people would fall for it.
raphlinus··on Speech Recognition and TTS in less than 500kb
There are two Sams here, the Microsoft one and the C64 one. I don't believe there's any connection between the two other than the name.

According to [1], the weight of a modern runnable version is around 39k.

The ratio of how good it sounds compared to how much computing power it uses is ridiculous. The C64 has ballpark 3 orders of magnitude less CPU throughput as an RP2350, and the codebase uses an impressive array of tricks to do actual formant synthesis (barely) and a pretty refined form of Elovitz text to phoneme conversion. One of my favorite tricks is its up and down bouncy pitch, which is not random, but based on the opposite contour as the first formant. It's simplistic but enough to make it not sound like a robotic monotone.

I've been playing around with this some myself and SAM is an inspiration, along with other landmark systems like MITalk (predecessor to DECtalk), SP0256, and other. I believe it's possible to use modern techniques to get pretty good sounding speech in, say, 64k and 10% of the throughput of a RP2350. It's really cool to see projects like OP, especially under permissive license.

[1]: https://simulationcorner.net/index.php?page=sam

raphlinus··on Decoding the obfuscated bash script on a Uniqlo t-shirt
The font is Roboto Mono, not Consolas.

There's something else a lot stranger going on, though. It is a proper monospace font, but the typesetting on the shirt is not. There's some kerning going on (I noticed it especially in the 'Iy' pair), and also it appears that narrower characters such as 'i' take less horizontal space. If I had to guess, I would say that it was set with a tool such as "optical kerning" in InDesign.

raphlinus··on Mir Books – Books from the Soviet Era
I had a lot of these books as a kid. I'm not even 100% why, but my dad glommed on to the idea that they were of high quality, and also quite inexpensive. Probably my favorite was "Higher Mathematics for Beginners" by Yakov Zeldovich[1]. I was a voracious reader, and largely taught myself calculus from it. Another one on probability and statistics I also found quite accessible (can't find it immediately).

There were some other sketchier pop-sci books, I remember one that had the "water has structural memory" theory in it. But likely those didn't do any lasting damage.

[1]: https://mirtitles.org/2022/07/04/higher-mathematics-for-begi...

raphlinus··on Odin, Wikipedia and engagement farming
Greg Kroah-Hartman has given a talk entitled "Untrusted data in Linux — How Rust is going to save us"[1] which I think is fairly optimistic about the idea that Rust will have an increasingly large role in the kernel. The total number of lines of code is small today but that will change.

[1]: https://www.youtube.com/watch?v=Nzmj7K0FNRY

raphlinus··on Ante: A new way to blend borrow checking and reference counting
Jake has already mentioned how the type system enforces single-threading (using similar mechanisms as Rc in Rust), but it's not necessarily the case that read-write races from multiple threads are UB.

In particular, Java allows aggressive multithreaded access, and the memory model has some pretty strong guarantees. Informally stated, if a read and a write race, then the read is guaranteed to observe either the old or new value.

Go is something of a middle ground [1]. Races on simple scalar values are not UB, but "fat pointers" (slices and interfaces) can tear and can lead to "arbitrary memory corruption."

When I was reading about "stable shape" I was wondering if it might be something similar. You easily get the same UB problems when dealing with sum types, as a tear across tag and variant payload can cause all kinds of things to go wrong.

[1]: https://go.dev/ref/mem

raphlinus··on Examining circuit boards from the Space Shuttle's I/O Processor
The description of the BCE reminds me a lot of the PIO in the RP2040 and 2350 microcontrollers. From the article, BCE instructions include "Transmit Data, Receive Data, Load Timeout Register, Store Status, and Wait."

To me, these correspond more or less 1:1 with PIO instructions OUT, IN, SET, INT, and WAIT. These plus PUSH, PULL (which can considered auxiliaries of IN and OUT), MOV, and JMP are all the PIO instructions. Like the BCE, it runs with completely deterministic clocking, one instruction per clock, and like the BCE there are a bunch of them (a total of 12 state machines on the 2350), though they now run totally in parallel rather than being time-multiplexed.

As a hobby project, I've lately been implementing USB (aiming for higher performance than Pico-PIO-USB, which proves that it's possible), and that's been quite fun.

I wonder to what extent they were explicitly inspired, and to what extent you just get convergent solutions when there are similar goals and constraints.

raphlinus··on Parallel Parentheses Matching
Also see Fast GPU bounding boxes on tree-structured scenes[1] (unpublished paper) and notes toward a blog post[2]. This is a highly tuned GPU implementation of parentheses matching. It's actually used in Vello (the classic version in which we offload basically all the work to the GPU, not the newer CPU-GPU hybrid version in which tracking the blend stack is done on the CPU).

Earlier versions of the work were featured on HN [3][4], but this is much more sophisticated. (plus a few more zero-comment submissions)

The basic idea (bicyclic semigroup and binary search) is the same as the submission. I think earliest attribution is to Bar-On and Vishkin[5] from 1985. Another implementation of this idea is in pareas[6], an experimental GPU-accelerated compiler.

I believe this work is publishable and would love to work with a student to resubmit it. Especially if you're a student or prof in Sydney, please reach out.

[1]: https://arxiv.org/abs/2205.11659

[2]: https://github.com/raphlinus/raphlinus.github.io/issues/66

[3]: https://news.ycombinator.com/item?id=24385095

[4]: https://news.ycombinator.com/item?id=27164009

[5]: https://dl.acm.org/doi/10.1145/3318.3478

[6]: https://github.com/Snektron/pareas

raphlinus··on "Fix" MacBook Neo Cursor Lag: Record 1 Pixel of the Screen Every 10 Seconds
This is plausible to me as well. A couple years ago we were trying to make dynamic memory allocation in Vello more robust and explored using async readback of a status buffer. In that case, the async task doesn't wake until the command buffer completes and signals a fence back to the CPU.

Long story short, performance was disappointing and we abandoned the approach. It's easy to believe it's a real problem especially when there are other factors including GPU being clocked down to save power.

Same caveat as parent, I have no direct knowledge of MacBook Neo or this specific issue.

raphlinus··on Show HN: Are You in the Weights?
No, a normal person would say, "I don't know them."

Silicon Valley bros, on the other hand...

raphlinus··on U.S. science is in chaos
Quite a few Ozzies who were in the US (especially Bay Area) and have decided to move back, but so far no other people with no real connection.
Page 1 of 34Next →