HNHacker News
TopNewBestAskShowJobs

ashvardanian

4,506 karma · joined November 26, 2021

Bootstrapping <Unum.cloud> since 2015. Investing at <AAL.vc>. Hosting ML/Systems hackathons around the globe. Living between London, San Francisco, and Yerevan. Converting outrageous amounts of coffee into modest amounts of Assembly for Intel, Arm, and Nvidia chips. Maintaining FOSS libs deployed on ≫ 100M devices.

<ashvardanian.com>

<github.com/ashvardanian>

submissionscomments
ashvardanian··on Performance of WebAssembly Runtimes in 2026
One fairly revealing microbenchmark for WASM runtimes is `int8` dot products & angular/cosine distances.

(My) NumKong [1] has implementations targeting both vanilla AVX2/Haswell and AVX2-VNNI/Alder Lake, which makes it easy to see where runtimes and code generators leave performance on the table.

I started a few Wasmtime/Cranelift PRs around this, but didn’t manage to finish them :facepalm: Might be a fun weekend project for someone interested in backend/codegen work.

[1]: https://github.com/ashvardanian/NumKong

ashvardanian··on Incident with Github.com
If only we had a myriad of heavily funded, founder-led AI-for-coding startups that could afford to store a few source files and need all that code for training anyway :)

PS: Saw Cursor’s Origin announcement a second later.

ashvardanian··on Claude Sonnet 5
Got really excited for this model and asked my Opus planners in 3 pretty different projects to use Sonnets instead of Opus subagents to help me experiment on HPC kernels faster. Not one of them ended up writing a single line of code... Sonnets just kept spinning, wasting tokens. Can't remember the last time it happened with Opus in my codebases. Reverting back.
ashvardanian··on Cloudflare Flagship
I really like the speed at which Cloudflare is executing toward becoming a critical infrastructure player with all of those new product offerings. That said, not everything needs to be serverless. Their Gen 13 hardware looks impressive, and it’s a pity you can’t rent it by the hour like AWS EC2 Metal instances.
ashvardanian··on NumKong: 2'000 Mixed Precision Kernels for All
I'm not aware of that, but it would likely be a great application area for SME!
ashvardanian··on NumKong: 2'000 Mixed Precision Kernels for All
The README was written by a human. I’ve used models extensively to refine the content, but never accepted more than a couple of lines of edits at a time.
ashvardanian··on What makes Intel Optane stand out (2023)
I don't have the inside scoop on Intel's current mess, but they definitely have a habit of killing off their coolest projects.
ashvardanian··on Meta's Race to Scale AI Chips for Billions: Four Chips in Two Years
Would it be accurate to say that Meta currently produces more RISC-V chips than other vendors? The specs for those chips look extremely interesting and seem much more programmable than Google's TPUs. It would be cool to see Meta making them available to third parties.
ashvardanian··on Nebius to buy AI agent search company Tavily for 275M
https://www.bloomberg.com/news/articles/2026-02-10/nebius-ag...

https://www.tavily.com/blog/tavily-is-joining-nebius

ashvardanian··on Zvec: A lightweight, fast, in-process vector database
8K QPS is probably quite trivial on their setup and a 10M dataset. I rarely use comparably small instances & datasets in my benchmarks, but on 100M-1B datasets on a larger dual-socket server, 100K QPS was easily achievable in 2023: https://www.unum.cloud/blog/2023-11-07-scaling-vector-search... ;)

Typically, the recipe is to keep the hot parts of the data structure in SRAM in CPU caches and a lot of SIMD. At the time of those measurements, USearch used ~100 custom kernels for different data types, similarity metrics, and hardware platforms. The upcoming release of the underlying SimSIMD micro-kernels project will push this number beyond 1000. So we should be able to squeeze a lot more performance later this year.

ashvardanian··on Show HN: Similarity = cosine(your_GitHub_stars, Karpathy) Client-side
Cool project! And thanks for mentioning "unum-cloud/USearch" among repo examples :)
ashvardanian··on Nvidia Kicks Off the Next Generation of AI with Rubin – Six New Chips
PTX is on the GPU side and is already supported on available models. On the CPU side, it must be some form of an Arm ISA extension, I believe, like NEON-FHM or SVE-AES… I'm just not sure what the scope of those extensions would be and how they will coexist with ARM’s other extensions.
ashvardanian··on Nvidia Kicks Off the Next Generation of AI with Rubin – Six New Chips
Every founder probably dreams of a press release like this — complete with testimonials from the CEOs of OpenAI, Anthropic, Meta, xAI, Microsoft, CoreWeave, AWS, Google, Oracle, Dell, HPE, and Lenovo.

There aren’t many technical details about the new GPUs yet, but the notes on the Vera CPU caught my eye. NVIDIA Spatial Multithreading sounds like their take on SMT — something you don’t usually see on Arm-based designs. Native FP8 support is also notable, though it’s still unclear how it will be exposed to developers in practice.

Overall it looks like an interesting CPU, but it doesn’t feel like it’s in the same league as the rumored Apple M5 Ultra.

ashvardanian··on I switched from VSCode to Zed
My workflow isn't very common. I typically have 3-5 projects open on the local machines and 2 cloud instances - x86 and Arm. Each project has files in many programming languages (primarily C/C++/CUDA, Python, and Rust), and the average file is easily over 1'000 LOC, sometimes over 10'000 LOC.

VS Code glitches all the time, even when I keep most extensions disabled. A few times a day, I need to restart the program, as it just starts blinking/flickering. Diff views are also painfully slow. Zed handles my typical source files with ease, but lacks functionality. Sublime comes into play when I open huge codebases and multi-gigabyte dataset files.

ashvardanian··on I switched from VSCode to Zed
I’m currently using a mix of Zed, Sublime, and VS Code.

The biggest missing piece in Zed for my workflow right now is side-by-side diffs. There’s an open discussion about it, though it hasn’t seen much activity recently: https://github.com/zed-industries/zed/discussions/26770

Stronger support for GDB/LLDB and broader C/C++ tooling would also be a big win.

It’s pretty wild how bloated most software has become. Huge thanks to the people behind Zed and Sublime for actively pushing in the opposite direction!

ashvardanian··on Microsoft please get your tab to autocomplete shit together
Not a fan of Windows either, but playing devil’s advocate here: Apple’s Finder has steadily gotten worse over the last ~16 years, at least in my experience. It increasingly struggles with basic functionality.

There seems to be a pattern where higher market cap correlates with worse ~~tech~~ fundamentals.

ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
Yes, CaseFolding.txt. I'm considering using the collation rules for sorting. Now they only target lexicographic comparisons and seem to be 4x faster than Rust's standard quick-sort implementation, but few people use it: https://github.com/ashvardanian/StringWars?tab=readme-ov-fil...
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
Thanks a lot for the correction! I'll adjust the references in a bit.
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
I was just about to ask some friends about it. If I’m not mistaken, Postgres began using ICU for collation, but not string matching yet. Curious if someone here is working in that direction?
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
Levenshtein distance calculations are a pretty generic string operation, Genomics happens to be one of the domains where they are most used... and a passion of mine :)
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
This is a very good example! Still, “correct” needs context. You can be 100% “correct with respect to ICU”. It’s definitely not perfect, but it’s the best standard we have. And luckily for me, it also defines the locale-independent rules. I can expand to support locale-specific adjustments in the future, but waiting for the adoption to grow before investing even more engineering effort into this feature. Maybe worth opening a GitHub issue for that :)
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
The GoLang bindings – yes, they are based on cGo. I realize it's suboptimal, but seems like the only practical option at this point.
ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
This article is about the ugliest — but arguably the most important — piece of open-source software I’ve written this year. The write-up ended up long and dense, so here’s a short TL;DR:

I grouped all Unicode 17 case-folding rules and built ~3K lines of AVX-512 kernels around them to enable fully standards-compliant, case-insensitive substring search across the entire 1M+ Unicode range, operating directly on UTF-8 bytes. In practice, this is often ~50× faster than ICU, and also less wrong than most tools people rely on today—from grep-style utilities to products like Google Docs, Microsoft Excel, and VS Code.

StringZilla v4.5 is available for C99, C++11, Python 3, Rust, Swift, Go, and JavaScript. The article covers the algorithmic tradeoffs, benchmarks across 20+ Wikipedia dumps in different languages, and quick starts for each binding.

Thanks to everyone for feature requests and bug reports. I'll do my best to port this to Arm as well — but first, I'm trying to ship one more thing before year's end.

ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
Yes — fuzzy and phonetic matching across languages is part of the roadmap. That space is still poorly standardized, so I wanted to start with something widely understood and well-defined (ICU-style transforms) before layering on more advanced behavior.

Also, as shown in the later tables, the Armenian and Georgian fast paths still have room for improvement. Before introducing higher-level APIs, I need to tighten the existing Armenian kernel and add a dedicated one for Georgian. It’s not a true bicameral script, but some characters are folding fold targets for older scripts, which currently forces too many fallbacks to the serial path.

ashvardanian··on Full Unicode Search at 50× ICU Speed with AVX‑512
I get why it sounds that way, but it’s not actually true.

StringZilla added full Unicode case folding in an earlier release, and had a state-of-the-art exact case-sensitive substring search for years. However, doing a full fold of the entire haystack is significantly slower than the new case-insensitive search path.

The key point is that you don’t need to fully normalize the haystack to correctly answer most substring queries. The search algorithm can rule out the vast majority of positions using cheap, SIMD-friendly probes and only apply fold logic on a very small subset of candidates.

I go into the details in the “Ideation & Challenges in Substring Search” section of the article

ashvardanian··on Myths Programmers Believe about CPU Caches (2018)
Here's my favorite practically applicable cache-related fact: even on x86 on recent server CPUs, cache-coherency protocols may be operating at a different granularity than the cache line size. A typical case with new Intel server CPUs is operating at the granularity of 2 consecutive cache lines. Some thread-pool implementations like CrossBeam in Rust and my ForkUnion in Rust and C++, explicitly document that and align objects to 128 bytes [1]:

  /**
   *  @brief Defines variable alignment to avoid false sharing.
   *  @see https://en.cppreference.com/w/cpp/thread/hardware_destructive_interference_size
   *  @see https://docs.rs/crossbeam-utils/latest/crossbeam_utils/struct.CachePadded.html
   *
   *  The C++ STL way to do it is to use `std::hardware_destructive_interference_size` if available:
   *
   *  @code{.cpp}
   *  #if defined(__cpp_lib_hardware_interference_size)
   *  static constexpr std::size_t default_alignment_k = std::hardware_destructive_interference_size;
   *  #else
   *  static constexpr std::size_t default_alignment_k = alignof(std::max_align_t);
   *  #endif
   *  @endcode
   *
   *  That however results into all kinds of ABI warnings with GCC, and suboptimal alignment choice,
   *  unless you hard-code `--param hardware_destructive_interference_size=64` or disable the warning
   *  with `-Wno-interference-size`.
   */
  static constexpr std::size_t default_alignment_k = 128;
As mentioned in the docstring above, using STL's `std::hardware_destructive_interference_size` won't help you. On ARM, this issue becomes even more pronounced, so concurrency-heavy code should ideally be compiled multiple times for different coherence protocols and leverage "dynamic dispatch", similar to how I & others handle SIMD instructions in libraries that need to run on a very diverse set of platforms.

[1] https://github.com/ashvardanian/ForkUnion/blob/46666f6347ece...

ashvardanian··on An overengineered solution to `sort | uniq -c` with 25x throughput (hist)
Storage, strings, sorting, counting, bioinformatics... I got nerd-sniped! Can't resist a shameless plug here :)

Looking at the code, there are a few things I would consider optimizing. I'd start by trying (my) StringZilla for hashing and sorting.

HashBrown collections under the hood use aHash, which is an excellent hash function, but on both short and long inputs, on new CPUs, StringZilla seems faster [0]:

                               short               long
  aHash::hash_one         1.23 GiB/s         8.61 GiB/s
  stringzilla::hash       1.84 GiB/s        11.38 GiB/s

A similar story with sorting strings. Inner loops of arbitrary length string comparisons often dominate such workloads. Doing it in a more Radix-style fashion can 4x your performance [1]:

                                                    short                  long
  std::sort_unstable_by_key           ~54.35 M compares/s    57.70 M compares/s
  stringzilla::argsort_permutation   ~213.73 M compares/s    74.64 M compares/s
Bear in mind that "compares/s" is a made-up metric here; in reality, I'm comparing from the duration.

[0] https://github.com/ashvardanian/StringWars?tab=readme-ov-fil...

[1] https://github.com/ashvardanian/StringWars?tab=readme-ov-fil...

ashvardanian··on First convex polyhedron found that can't pass through itself
In case someone is searching for the computational part of the proof, its on GitHub, implemented using SageMath: https://github.com/Jakob256/Rupert
ashvardanian··on VectorWare – from creators of `rust-GPU` and `rust-CUDA`
My bad! "contributors" is more accurate, but HN doesn't allow editing titles, sadly :(
ashvardanian··on Show HN: Using LLMs and >1k 4090s to visualize 100k scientific research articles
Congrats on the release, Sam - the preview looks great!

I'm curious about the technical side: how are you handling the dimensionality reduction and visualization? Also noticed you mentioned "custom-trained LLMs" in the tweet - how large are those models, and what motivated using custom ones instead of existing open models?

Page 1 of 14Next →