HNHacker News
TopNewBestAskShowJobs

maxime_cb

448 karma · joined December 6, 2021

submissionscomments
maxime_cb··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
OP here: I think you're reading too much into the title. It's not about the failure of Rust. It's about where I stepped in with some custom optimizations for better performance. If you read the blog post, the solution implemented is very Rust-y as well. It uses a newtype with custom methods to try and make the code as nice and readable as possible.

FWIW this post was very well received on the Rust subreddit. They loved it and it got over 350 upvotes, so that community definitely didn't receive it as some sort of attack on Rust.

maxime_cb··on Replacing a Rust Enum with a 64-Bit Word Made My Interpreter 17% Faster
Author here. The disassembly for the old enum handling had many spills, simply because the old value enum can't fit in a single register. If you have an instruction that two operands with two of those big value enums, it needs 4 registers instead of 2. That, coupled with better cache-friendliness, explains a lot.
maxime_cb··on GPT 5.6 Sol is the best "vision" model OpenAI ever released
I'm assuming you mean that this tech became available in OpenCV 25 years ago, but as it turns out, the underlying tech can be traced back much further, at least as far as 1977! :)

https://ieeexplore.ieee.org/document/1674847 G. J. Vanderbrug and A. Rosenfeld, “Two-Stage Template Matching,” IEEE Transactions on Computers, Vol. C-26, No. 4, pp. 384–393, April 1977. DOI: 10.1109/TC.1977.1674847

maxime_cb··on Show HN: I built one of the most advanced web-based drum/beat sequencers
Thanks :)

I've made a couple of music apps over the years, including one of the oldest JS music apps on the web, that also does URL encoding: https://pointersgonewild.com/2012/04/23/musictoy-music-made-...

But also generally, I was looking at other web-based drum machines, and most of them are either super basic, full of ads, or they have bad cryptic UIs. I also wanted to design a more advanced sequencer with a timeline that would have an intuitive UI to work with.

I built a sequencer in NoiseCraft, a browser-based music programming language I built a few years ago (e.g. https://noisecraft.app/101), but you can't chain pattern in this sequencer, which makes it harder to create longer piece (though people have found workarounds).

maxime_cb··on ZJIT removes redundant object loads and stores
I gave a talk about ZJIT and the motivation for the change at RubyKaigi 2025 if people are curious. It's on YouTube.
maxime_cb··on ZJIT removes redundant object loads and stores
Max Bernstein is now leading the team. He's also an excellent compiler engineer.
maxime_cb··on Reflections on 2 years of CPython's JIT Compiler
Thanks Ken. Apologies if I misunderstood the situation. I wish you all the best.
maxime_cb··on Reflections on 2 years of CPython's JIT Compiler
Ruby has the same unfortunate problem.
maxime_cb··on Reflections on 2 years of CPython's JIT Compiler
Instigator of YJIT, the CRuby JIT here.

It's easy to dismiss our efforts, but Ruby is just as dynamic if not more than Python. It's also a very difficult language to optimize. I think we could have done the same for Python. In fact the Python JIT people reached out to me when they were starting this project. They probably felt encouraged seeing our success. However they decided to ignore my advice and go with their own unproven approach.

This is probably going to be an unpopular take but building a good JIT compiler is hard and leadership matters. I started the YJIT project with 10+ years of JIT compiler experience and a team of skilled engineers, whereas AFAIK the Python JIT project was lead by a student. It was an uphill battle getting YJIT to work well at first. We needed grit and I pushed for a very data-driven approach so we could learn from our early failures and make informed decisions. Make of that what you will.

Yes Python is hard to optimize. I Still believe that a good JIT for CPython is very possible but it needs to be done right. Hire me if you want that done :)

Several talks about YJIT on YouTube for those who want to know more: https://youtu.be/X0JRhh8w_4I

maxime_cb··on The Ruby on Rails Podcast Episode 508: YJIT with Maxime Chevalier-Boisvert
YJIT is optimized primarily for web workloads. We look at rails performance a lot, but also at various other libraries that are used in that context. If you look at the headline benchmarks at https://speed.yjit.org, it will give you an idea of what we're mostly focused on. This is in contrast with academic compiler project, which often focus on microbenchmarks and code that is very different from the code users actually run in practice.

YJIT or TruffleRuby: I'm biased being that I work on YJIT. The nice thing about YJIT is that it's likely to work out of the box, and just be a matter of calling `ruby --yjit` to turn it on. It will probably use a lot less memory than TR, and it's probably more likely to deliver the result you're looking for at this time (speed boost, no hassle).

That being said, for some small or specialized applications, TruffleRuby could deliver much higher peak performance. If you don't restart your server often and you have a lot of memory available, then maybe you don't care about warm-up time or memory usage, and TruffleRuby could be the right tool for you. Feel free to run your own benchmarks and also to blog about the results (though if you do, please share as much details about your setup as possible).

maxime_cb··on The Ruby on Rails Podcast Episode 508: YJIT with Maxime Chevalier-Boisvert
Yes, Marc Feeley was my PhD advisor. We came up with the original idea together. I also see it as a development of the work I did in my M.Sc. thesis on type-driven versioning of functions. Basic block versioning is lazy, type-driven tail splitting of code.
maxime_cb··on Ruby 3.3's YJIT: Faster While Using Less Memory
Hope you try again with 3.3. The improvements we've made to YJIT since Ruby 3.1 are massive.
maxime_cb··on YJIT is the most memory-efficient Ruby JIT
It is enough iterations for these VMs to warm up on the benchmarks we've looked at, but the warm-up time is still on the order of minutes on some benchmarks, which is impractical for many applications.
maxime_cb··on YJIT enabled by default, Active Model improvements and much more
The article doesn't go into super deep details but we do touch on it in the paper we've recently published: https://dl.acm.org/doi/10.1145/3617651.3622982

And I went into some more details the talk I gave at RubyKaigi 2023: https://www.youtube.com/watch?v=X0JRhh8w_4I&t=2404s

maxime_cb··on YJIT enabled by default, Active Model improvements and much more
Ruby 3.3 (coming this Christmas) will have a much faster and more memory efficient YJIT than 3.2. We've made major improvements this year.
maxime_cb··on YJIT enabled by default, Active Model improvements and much more
YJIT tech lead here.

On the flip side, YJIT is probably one of the most memory-efficient JIT compilers out there (for any language). I say this having spoken to other JIT implementers.

We've worked really hard to reduce the memory overhead and at Shopify it's now down to less than 10% in our flagship production deployment.

Regardless, if memory usage is a legitimate concern for you, you can very easily remove the Rails initializer that turns on YJIT. You can choose between memory usage and response time. The choice is yours.

maxime_cb··on Is frying food possible in space?
And very crispy.
maxime_cb··on Software bugs that cause real-world harm
Looks pretty cool :)
maxime_cb··on Software bugs that cause real-world harm
Author here. If you've grown up in a "normal", functional family, with two loving parents, and you enjoy talking to them on the phone, you should consider yourself very lucky.

I only take calls from my mother when I feel up to it. The reason for that is that she has no concept of boundaries. Her mental illness prevents her from grasping that concept. Her default behavior tends towards what a normal person would call harassment. If you can't relate or understand a situation like that, it might just be because you've had a relatively safe, coddled, privileged life.

maxime_cb··on Software bugs that cause real-world harm
> I was wondering what the author was smoking with the "ringtone & notifications" complaint. I don't remember seeing an Android phone that did not have separate volume slider for each.

This is why I included a screenshot of the volume sliders my Google Pixel displays. Samsung doesn't use stock android OS, and some Samsung users tend to assume every android phone works the same.

maxime_cb··on Software bugs that cause real-world harm
Author here.

The reason I tend to think it's a software bug is that it seems that the system to dispatch deliveries is automated. The subcontractors get their orders from some kind of computerized system, it seems. That system seems to systematically fail to specify when they are to carry items indoors/upstairs. Whether that's due to negligence or intentional malfeasance, don't know.

What I do know from experience is that there are numerous bugs on their website, besides the "unknown error" problem I've listed. It just seems like really shittily built software... So I would tend to think there's an issue with really poor software engineering practices at that company.

maxime_cb··on RJIT, a new JIT for Ruby
> However the Ruby community seem to want having their own JIT written in C more than they want performance.

YJIT is written in Rust, not C, but it's also not just a matter of wanting to write our own JIT for fun. There are a number of caveats with TruffleRuby which make a production deployment difficult:

1. The memory overhead is very large. Can be as much as 1-2GB IIRC.

2. The warm-up/compilation time is much too long (can be up to minutes for large applications). In practice this can mean that latency numbers go way up when you spin up your app. In the case of a server application, that can translate in lots of requests timing out.

3. It doesn't have 100% CRuby compatibility, so your code may not run out of the box.

There's a reason why you don't see that many TruffleRuby (or TrufflePython, TruffleJS, etc.) deployments in the wild. Peak execution speed after a lengthy warm-up is not the only metric that matters.

maxime_cb··on RJIT, a new JIT for Ruby
Yes. Prior to that point we used to allocate a large chunk of executable memory upfront. We switched to mapping that memory on demand, and that alone was a huge improvement.
maxime_cb··on RJIT, a new JIT for Ruby
For anyone curious, we've been working to reduce the memory overhead and have added some stats to keep track of memory usage over time. On this graph, you can see a comparison with the CRuby interpreter:

https://speed.yjit.org/memory_timeline#railsbench

maxime_cb··on RJIT, a new JIT for Ruby
You're right that the peak performance could be on par (or even better), but, and I acknowledge that I'm biased since I'm tech lead of the YJIT team, my takeaway is:

1. Kokubun, who works with us on the YJIT team, is leveraging insights he's learned while working on YJIT to build this. He has said so in his tweets, some of the code in RJIT is a direct port of the YJIT's code to Ruby. This is his side-project.

2. One of the challenges we face in YJIT is memory overhead. It's something we've been working hard to minimize. Programming RJIT in Ruby is going to make this problems worse. Not only will it be hard to be more memory-efficient, you're going to increase the workload of the Ruby GC (which is also working to GC your app's data).

3. Warm-up time will also be worse due to Ruby's performance not being on par with Rust. This doesn't matter much for smaller benchmarks, but for anything resembling a production deployment, it will.

On your second point, if you're not seeing perf gains with YJIT, we'd be curious to see a profile of your application, and the output of running with `--yjit-stats`. You can file an issue here to give us your feedback, and help us make YJIT better: https://github.com/Shopify/yjit/issues

maxime_cb··on Miller Puckette: Inside PureData – Lectures on pd/development of computer music
Shameless plug: for something more approachable, I've created NoiseCraft, which runs in a web browser, has fewer primitives, and is designed to be easier for beginners to grasp: https://noisecraft.app/1136 (just hit "Play" in the top-right corner).

More examples you can try on the browse page: https://noisecraft.app/browse

maxime_cb··on Building a Minimalistic Virtual Machine
If you want to compute like it's 1986, I think it would be fun to build a BASIC interpreter for UVM. Bonus points if you do your own text rendering on a blue background and you add simple primitives for 2D graphics.

https://github.com/maximecb/uvm/issues/7

maxime_cb··on Building a Minimalistic Virtual Machine
I was thinking something like actors, or independent processes sending messages would be nice. Just because it's very safe and predictable. Less error-prone than threads.

The thing that kind of gets me is it seems difficult to have safe shared memory with actors? You ideally want to be able to share memory if you want things to be efficient, but if you have shared memory, then you get into issues with atomic writes and things being observed in different orders, etc.

maxime_cb··on Building a Minimalistic Virtual Machine
> For concurrency it's a lot harder to make a good interface

Can you elaborate?

maxime_cb··on Building a Minimalistic Virtual Machine
> Getting around pointer events inconsistencies is a lot easier than building your own cross-platform VM

For sure, and I did. I wrote some code that makes an invisible fullscreen div appear just at the right time to prevent pointer events being triggered when they shouldn't be #cleancode. It's just frustrating that things like that need to be done, and how often they might need to be done.

> I imagine there will also be differences in the way macOS, Linux and Windows handle graphics, IO, audio, etc, that will eventually leak to UVM, it's just the nature of the challenge.

At the moment you can create a window with one function call, and you have another function call to copy one frame's worth of pixels into the window. The pixel format is in BGRA byte order, 32-bits per pixel, and that's the only option. I'm going with really basic, low-level APIs like that because they're harder to get wrong.

Audio is going to be equally simple. There could be cross-platform differences in things like the amount of latency to write audio, but I'll do my best to make the APIs extremely portable and hard to get wrong.

Page 1 of 3Next →