2,646 karma · joined September 22, 2013
andrew at $DOMAIN
vector conditionmask = <some computation...>; // E.g., 11111111 00000000 00000000 11111111
vector truebranch = <some computation...>;
vector falsebranch = <some computation...>;
vector result = (truebranch & conditionmask) | (falsebranch & ~conditionmask);
where each lane of the conditionmask has either all bits set or all bits clear, depending on the outcome of the conditional test for that lane.The processor obviously does execute both branches here, so there's going to be wasted work. But since it's just a linear sequence of operations it can often schedule them independently and run them out-of-order and in parallel. And of course, if there's any shared computation between the two branches, the compiler can do common subexpression elimination.
That said, that sort of approach where you go ahead and do both and then blend them was definitely the kind of optimization where you'd want to profile rather than doing it blindly. But it was a pretty common thing to do when hand-vectorizing code. (Thankfully, auto-vectorizers are pretty good at doing this sort of optimization for you these days. It's been a very long time now since I've had to hand-write vector intrinsics.)
I really don't get why people spend so much on flagship phones! I tend to aim for $200 Motos for me and my family, and it feels like it's a 90%+ solution at a tenth of the price. I can still browse the web, text with people, make calls, listen to music, and take pictures. They still last for years, and if they get dropped and the screen cracks, it just means its time to upgrade. What am I missing here?
All that said, there's definitely been research into samplers that combine low-discrepancy with blue noise properties (often including retaining those properties even in lower-dimensional projections produced by dropping axis).
[1] https://en.wikipedia.org/wiki/The_Cuckoo%27s_Egg_(book)
[2] https://news.ycombinator.com/user?id=CliffStoll
[3] https://news.ycombinator.com/item?id=21830277
I'm not logging in to my private account from a work machine that's MITMd for "security", just to browse search results or take a short break!
(I've been a user for nearly 21 years, but using it less and less with each stupid choice - these days I post a comment maybe once a week there. I think I may finally be done with Reddit, sadly. It was good while it lasted.)
> The study finds annual mileage thresholds below which the replacement EV’s manufacturing emissions are not recovered: about 7,054 kilometers for cars, 6,837 kilometers for SUVs and 10,794 kilometers for trucks. Those are well below the roughly 20,000 kilometers in the researchers’ average case, but very lightly driven vehicles clearly exist and can be better left in service.
("Kilometerage threshold"? Anyway, for those using miles, that's 4383 mi for cars, 4248 mi for SUVs, and 6706 mi for trucks.)
Still, I like my old ICE car partly because it's from the offline vehicle era and completely lacks any kind of remote telemetry. Perhaps if I didn't have to choose between EVs and lack of telemetry, I might be more inclined to upgrade sooner. (That, plus it's still nice to be able to just completely refuel in a couple of minutes on long road trips through areas where services are very sparse.)
Unlike the prior nag screens, this one was modal, blocked any scrolling, and had no way to close it. It seemed to pop up about 15 seconds after I first start scrolling any page - either a list of threads on a subreddit, or a thread page itself.
Man I miss the old iOS-looking Reddit mobile (m.reddit.com, IIRC).
My Reddit account is now just a few months short of US drinking age. I've been using the site since before subreddits were a thing (back when the only topic-specific alternative to the general front page was programming.reddit.com). I've posted a lot, modded some, and stuck with them through a lot of stupid decisions over the years. But I'm beginning to think it's finally time to move on; I'm just not sure where to.
[1] https://www.highperformancegraphics.org/2026/schedule/
[2] https://dl.acm.org/doi/10.1145/3820013
[3] https://www.youtube.com/live/vTUdO2A73i0?si=Hy6zle7EtWp7rL60...
I'm going to answer this in two parts.
First, you absolutely can update the BVH, keeping the tree topology the same, and just updating the internal bounding as the triangles and other primitives move around. It's essentially a bottom-up walk of the tree where you figure out where the triangles now are, update the bounding boxes to contain them tightly, then update the boxes of the parent nodes, then the parents of those parents, etc., all the way back up to the root node of the tree. The tree structure stays the same and the bounding boxes around each node just update. This is often called "refitting", as you just refit the bounding boxes around everything in the tree below them like a rubber band.
For simple animation where things only move slightly between frames, this works just fine. And if you're using the DirectX Ray tracing (DXR) API, you can do this via what it calls "acceleration structure updates" [1]. You have to tell the API that you plan to do this when you first build the BVH, and then pass another flag when you actually do it. And there are rules that you have to keep the triangles, the same and only move the vertices. But you can do it pretty easily. And in fact, it's usually nice and fast, often much faster than building a BVH from scratch.
Where it can fall apart, however, is in the performance of the ray tracing itself. As a super-simple example, let's say we have a little BVH of three triangles (A, B, and C) and three nodes (1, 2, and 3), something like this:
+-------------------------+
|1 +-----+|
| |3 * ||
| | /C\ || 1
| |*---*|| / \
| +-----+| 2 3
|+---------+ | / \ \
||2 * *---*| | A B C
|| /A\ \B/ | |
||*---* * | |
|+---------+ |
+-------------------------+
If, during animation, triangle B moves differently than A and C, e.g., B moves toward the right while A and C stay mostly in place, then after refitting/updating your BVH looks more like this: +-------------------------+
|1 +-----+|
| |3 * ||
| | /C\ || 1
| |*---*|| / \
| +-----+| 2 3
|+-----------------------+| / \ \
||2 * *---*|| A B C
|| /A\ \B/ ||
||*---* * ||
|+-----------------------+|
+-------------------------+
Still the same tree topology, but look how large the bounding box around node 2 has grown! If a ray happens to hit node 1 and we trace into this bit of the tree, there's a good chance that it will pass through all that empty space in node 2, without actually hitting triangles A or B. We'll pay the cost of intersection testing node 2, traversing into it and then doing intersection testing against triangles A and B, even though we'll likely miss them. That's going to slow down your tracing, and the more the boxes have to deform and bloat like this to fit your animated geometry, the worse the ray tracing performance will get.Instead, for something like the above, you'd want a tree that looks more like this:
+-------------------------+
|1 +-----+|
| |3 * ||
| | /C\ || 1
| |*---*|| / \
| | || 2 3
|+-----+ | || / / \
||2 * | |*---*|| A B C
|| /A\ | | \B/ ||
||*---*| | * ||
|+-----+ +-----+|
+-------------------------+
Now, triangles B and C which are nearer to each other are grouped together, while A is alone. And if the ray hits node 1 but just passes through the empty space between nodes 2 and 3 we won't have to do any triangle intersections. That's going to be faster. The downside is that the tree topology is different, which means this is no longer a simple refit. Instead, traditionally, we might have to build the tree from scratch. That's going to take a lot longer than doing a refit, but the trade-off is better performance for your ray tracing. So typically there's a bit of a balancing act that game developers do around deciding how long they can get away with just refitting the animation to the old BVH vs. when they need to throw the current BVH away and rebuild it from scratch. The name of the game is to optimize the sum of the BVH build and refit times plus the time to traverse and intersect them when rendering. (There are other tricks that can help, but I'm skipping in the interest of keeping this basic.)Now, the second part of my answer. There are 580 million triangles in the paper's demo video. Even if you only refit and never rebuild the BVHs, updating all of those vertices for animation is going to be prohibitively expensive. (And even just doing the bone animation for the vertices for 580 million triangles might be too costly, let alone the BVH updates.)
So what's happening here is a bit of cheating. A coarse grid of tetrahedra (2 million in this video) are built around the model in rest pose, and for each tetrahedra, the triangles contained in it or touching it are determined. Then just those tetrahedra are animated and the BVHs around them updated/rebuilt. When a ray intersects an animated tetrahedron, the ray is warped back to where it would be relative to the rest pose and intersected against against the triangles that tetrahedron touches, using a little micro BVH. (I should note that this sort of "warp the ray" idea is pretty common in ray tracing. It's how instancing works so cheaply, for example.) If you look at the paper itself [2], you can see an illustration of the tetragedral grid in Figure 2 and the ray warping in Figure 5.
But the point of the paper is that after all the initial setup, it's faster to update the BVH for those 2M tetrahedra than it is for 580M triangles, or whatever ratio you choose to use. You can choose the ratio depending on how approximate and fast or faithful and slow you want the animation cost to be (by how coarse or fine the tetrahedral mesh is, respectively), but the point is that it lets you decouple the BVH update costs from the triangle count.
[1] https://microsoft.github.io/DirectX-Specs/d3d/Raytracing.htm...
[2] https://gpuopen.com/download/TetrahedralMeshes_AuthorsVersio...
(At least that's how I've heard it phrased before.)
I myself am probably somewhere in between an over-writer and a side-writer lefty, edging slightly closer to an over-writer. Maybe a 20 degree angle to the line as my thumb wraps around my other fingers?
I've never found a fountain pen that I've liked because every time I've tried one, I ended up tearing the paper as I go as I'm essentially pushing the nib across the page and so the nib snags on the paper, digs in, and makes little holes. And I left pencils behind as soon as I could in school because every day I'd come home with the side of my left hand silvered with graphite.
These days, my preference is a Pentel EnerGel, a Uniball Signo, or a Pilot G-2 (i.e., gel pens like the article mentions). They can write easily without much force and dry quickly enough so that my hand doesn't smudge. For my daily pocket carry, I've also learned that I always want the capped variants rather the click kind, or else I'll end up with stained inner pockets and scribbles on my thigh. This makes the G-2s nice for home and desk use, but not for daily carry.
(And though the article suggests using cursive and not getting hung up on legibility, somewhere back in grade or middle school I switched to using either print style, or an engineering-hand-like small caps style, depending on mood. Both in the interest of legibility, despite my left-handedness.)
I don't think it's in the system prompt, but that the harnesses time-stamp each turn in the context.
And from what I've seen, they also include the current and max context, so that the model can decide whether to continue work, suggest compaction, or prefer actions that might reduce the growth of its context.
I'd also add `C-h b` to show you the key bindings. (And `C-h` after a prefix key will usually show you the bindings that start with that prefix.) `C-h a` for apropos to search commands by substring can also be useful.
The thing that makes it really "self documenting" is that these help commands reflect the live environment at the moment you use it. If you've added a new binding in your init.el, `C-h b` and `C-h k` will show it. If you've added a new function in your init.el, or loaded a custom package, all those functions can now be found via `C-h f`. The help system will show you the doc strings for them and provide hyperlinks right to the source.
Moreover, this works for anything that you define on the fly. Open an Emacs lisp buffer, type some elisp code to define a function or variable, execute the definition, and now it'll appear under the above in the help system the next time you invoke help.
As a user since '97, I've often felt that this philosophy is entirely, well, backwards. I know how to read the release notes to learn of such changes and how to edit my personal init.el file to revert a setting if I don't like the new default. As long as no one takes away the option, the default doesn't really matter too much to me. But newcomers who might not yet be comfortable with editing their init.el files could really benefit from a more optimal out-of-box experience.
(And besides that, often the newer option is something that I've already moved on to, so making it the new default means I can now remove it from my init.el. I always enjoy when I discover that I can cut something from my init.el because it's now in base Emacs.)
(Personally, I've also enjoyed unit origami, which involves folding the same module many times over and assembling them.)
(These days it's stored safely away with batteries removed, so I don't use it that much anymore. For convenience, I usually just use either Droid48 on my phone, or Emacs Calc at my computers.)
We were quiet, predictable, don't-rock-the-boat tenants, and the rando owner mentioned that they valued that enough that raising rents wasn't worth the potential risk of new tenants who might cause them more hassle.
(I know, I know... the answer is probably that they expect me to just move my software development to the cloud, too. Joy!)
> "Good Morning!" said Bilbo, and he meant it. The sun was shining, and the grass was very green. But Gandalf looked at him from under long bushy eyebrows that stuck out further than the brim of his shady hat.
> "What do you mean?" he said. "Do you wish me a good morning, or mean that it is a good morning whether I want it or not; or that you feel good this morning; or that it is a morning to be good on?"
> "All of them at once," said Bilbo. "And a very fine morning for a pipe of tobacco out of doors, into the bargain.
> The draft models seamlessly utilize the target model's activations and share its KV cache, meaning they don't have to waste time recalculating context the larger model has already figured out.
Or better yet, pitting Eliza vs. Parry (https://logic.stanford.edu/complaw/readings/elizaandparry.pd...), where Parry was meant to simulate a paranoid schizophrenic. That was 1973, more than 50 years ago.
Everything old is new again.
pal.r = pal.g = pal.b = (77 * pal.r + 150 * pal.g + 29 * pal.b) >> 8;
Hardware floating point was rare before the 486 DX and Pentiums. Not to mention that Integer<->FP conversion was slow. And division of any kind has always been slow. So you'd see a lot of fixed-point math approximations with power-of-two divisors so that you can shift-right.