The template error messages look great. I wonder if it’s worth writing a translator from clang/gcc messages to these ones, for users of clang and older gcc (to pipe one’s error messages to)
One of the big bottlenecks of plotting libraries is simply the time it takes to import the library. I’ve seen matplotlib being slow to import, and in Julia they even have a “time to first plot” metric. I’d be curious to see how this library compares.
The fastest thing I would think is to process 255 vectors at a time, using an accumulation vector with the k-th byte in the vector representing the # of evens seen for the k-th byte of the vectors in that 255-vector chunk. But, I don’t know how to make autovectorization do that. I could only do that with intrinsics.
Seems like you can do this sort of speed up even without the 256 constraint. Just run this sped up version, but after each set of 256 iterations, dump all the 8-bit counters to the 64-bit final result.
This article presents a readable overview of today’s NBA trends, but IMO is too absolute in its judgment. Basketball is not a solved sport. There is still innovation, for example with OKC’s historically good defense that relies on playing 5 smaller but faster players. There are still good all-around players. There are still people that hit a lot of mid range shots. We have trends going the other way, sure, but they have their own set of tradeoffs and are neither a total solution nor totally embraced in the NBA. Teams will continue to evolve based on the talents of people at their disposal and their own innovative ideas.
Would allocating a 640-byte string initially really be the right tradeoff? It seems like it could result in a lot of memory overhead if many small strings are created that don’t need that space. But it does save a copy at least
As for the int-to-string function, using the division result to do a faster modulus (eg with the div function) and possibly a lookup table seem like they’d help (there must be some good open source libraries focused on this to look at).
Do you expect that Apple’s bigger security initiatives, like pointer authentication and writing the OS in a memory safe language, will improve the situation?
All of the would be a big savings in code complexity and a win for reliability, compared to doing new untested optimizations. If memory usage is a concern, I’m sure there’s a fast C SAX parser out there (or maybe one within yyjson)
Very fun read! I’m curious though, when it comes to non-Ruby-specific optimizations, like the lookup table for escape characters, why not instead leverage an existing library like simdjson that’s already doing this sort of thing?
It could be improved even more, performance wise. The potentially expensive modulo could be avoided entirely with an if statement. Or, only use powers of 2 for the capacity, and then you can also use bit wise ops
I refuse to believe they didn’t notice a key difference in the phones - the iPhone 16 will feel much, much snappier. Games aside, just opening any old app will be much faster. I notice this every time I upgrade, even when it’s just going up 3 versions, not 8.
When I start a lot of projects and only get halfway through them, I feel overwhelmed and frustrated by all the loose ends. I like to see projects through to release if I think they’re worth it, but that also requires a bit of self-imposed discipline. I don’t think there’s any shame in having a lot of half-finished projects, but I find more happiness in pushing myself to finish at least some of them.
Moving Linux and its associated projects to GitHub would definitely help. I don't buy the arguments that GitHub isn't suitable for it, e.g. because it doesn't allow each sub-project to have its own area. Things like tags help a lot to separate issues and PRs based on which sub-project they're a part of.
When building something that I want to run on both CPU and GPU, depending, I’ve found it much easier to use PyTorch than some combination of NumPy and CuPy. I don’t have to fiddle around with some global replacing of numpy.* with cupy.*, and PyTorch has very nearly all the functions that those libraries have.
This looks similar to Triton, I wonder what it does differently. But in any case, for any of these libraries, it would be awesome if it could output object files from this, with PTX or SASS code. Then it can be linked into a binary instead of needing a Python environment to run it.
How uncommon is it for apps to store sensitive data in this way? It wouldn’t surprise me if this is a pretty common, albeit non-ideal, practice. For example, where does chrome store browsing history data?
Why wouldn’t they just replicate bash or some other UNIX shell, along with the basic UNIX tools like cp and find with matching APIs? Huge mistake there imo, even if they did add a few bells and whistles with powershell