HNHacker News
TopNewBestAskShowJobs

Archit3ch

513 karma · joined February 22, 2019

archit3ch@gmail.com
submissionscomments
Archit3ch··on GPT-Synopsys: Frontier Intelligence to Revolutionize Chip Design
You get your non-deterministic process (hired humans or LLMs) to create deterministic scripts (='generators' in EDA terms), which can then be audited, corrected, etc.

DRC/LVS/PEX/SPICE are deterministic, but the tools themselves are not without faults.

Archit3ch··on GPT-Synopsys: Frontier Intelligence to Revolutionize Chip Design
Astra is amazing with open PDKs. ;)
Archit3ch··on The state of SIMD in Rust in 2026
Note that ISPC does not handle symbolic math e.g. sparsity preserving linear solvers or equation simplification a la Mathematica. Manual ASM does. ;)
Archit3ch··on The state of SIMD in Rust in 2026
2x/4x/8x is still thinking in terms of autovec.

I'm getting 50x faster code with manual ASM. That's the difference between audio code that runs in realtime and code that does not.

The competing implementations use SIMD and native code, autovec works nicely there. I symbolically invert the LinAlg system at compile time.

Archit3ch··on The state of SIMD in Rust in 2026
Hot take: there is no portable SIMD.

You can either have performance (=write manual ASM for each platform), or portability, but not both.

What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)

Archit3ch··on Fearless SIMD v1.0
Wuthering SIMD
Archit3ch··on Fearless SIMD v1.0
Possible in Julia without patching the compiler with a compiler extension package.
Archit3ch··on More floating point alternatives
You can also use this in analog hardware! :D

https://en.wikipedia.org/wiki/Translinear_circuit

Archit3ch··on Apple M6 Pro Achieves the Highest Single-Core CPU Score in Geekbench 7
> AVX still wins though.

On what, microbenchmarks? I have realtime audio workloads where the hot loop is essentially linear algebra. Exactly the kind of work that suits AVX2/AVX512. Guess what, Apple Silicon still pulls ahead because real workloads are branchy, cache-hungry, full of dependencies and do not line up in 8 neat f64 operations per cycle.

Archit3ch··on Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?
Astra

+ hopefully reset?

Archit3ch··on I wanna live an NPC life
Gervais Principle strikes again.
Archit3ch··on Run macOS Software on Linux
> Is there some really amazing Mac only software that would justify the effort ?

As others said, the killer app is XCode. Even if you can technically get it to run, you'll still need the full support and test flows of a native solution, so you might as well follow the proper path and use a Mac.

I can get Visual Studio to run in Parallels today. I still test with a real Windows machine.

Archit3ch··on Run macOS Software on Linux
I assume this wouldn't work for software that requires e.g. the Secure Enclave.
Archit3ch··on Mojo is now open source
> But compared with Python, Julia, Matlab, R, Rust, C, C++ - Mojo feel relief for working with numerics.

Out of those, Julia is the only one that combines Multiple Dispatch and native code, both important for numerics.

Archit3ch··on Mojo is now open source
Julia compiles to native code, same as C++/Rust.
Archit3ch··on Mojo 1.0 vs. Jac, a Roadmap
Does Jac have Multiple Dispatch?
Archit3ch··on GPU Offload in Rust: Portable, Safe, and Fast
> multi-vendor GPU compilation framework

Technically true, since it supports NVIDIA and AMD.

But we have a different definition of portability, if I cannot bring a Metal device and expect it to work.

Archit3ch··on Apple announces changes for apps in the European Union
Wait, this is only for i(Pad)OS and not MacOS, right? The way it's written, it's a blanket 5% commission on desktop apps distributed on one's own website, which is insane!

> Web distribution, which is available only in the EU

I assume this settles it, since it was always possible to web-distribute desktop apps outside the EU.

Archit3ch··on Mojo 1.0
As someone who writes GPU kernels in Julia, yes it is. But Mojo is one of the few alternatives based on technical merit alone. Each language is better in different metrics, and it could be argued either way.

License-wise, Mojo reminds me of paid Borland compilers that were in vogue before I was born. Even if it's open-sourced, what incentive does Qualcomm have to maintain a happy path for Metal deployment? Mojo might actually be the best way to program Apple GPUs (on a technical level) and it would still lose due to politics.

Archit3ch··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
> 64GB Macbook pro with an M5 chip costs more than 4k Euros in Germany

Sure, if you want the latest and almost* greatest. You can pick up an M1 Max 64GB for ~1k.

* I guess 128GB also exists

Archit3ch··on Can Intel finally beat ARM on performance per Watt?
> in FP64, which is not something most people use

It's used in audio processing. But large DAW sessions need something beefier than both tested devices (Macbook Neo, XPS).

Archit3ch··on Can Intel finally beat ARM on performance per Watt?
> consumers basically never use fp64 any more

They do in audio processing. ;)

Archit3ch··on Helsinki Hacker News Meetup
Same!
Archit3ch··on We accidentally built an LLVM compiler for Jax
It proves the point. Numpy still doesn't have the semantics of a full programming language like Mojo/C++/Julia.
Archit3ch··on We accidentally built an LLVM compiler for Jax
Sure, but if you're writing in the JAX DSL you've already abandoned python.
Archit3ch··on Ask HN: Anyone still do work on Intel Macs?
Never had one, but I might pick one up on ebay if I ever need to support Intel Mac users.
Archit3ch··on We accidentally built an LLVM compiler for Jax
1. Why not Julia? IIRC, some/all of this already exists in that ecosystem.

> Something to note is that quantum algorithms are never purely quantum. There is a tonne of classical processing needed to get inputs ready, post-process outputs, and even construct and represent the quantum circuits themselves.

2. Did you face any problems for the 'surrounding' areas by JAX imposing the limitation of pure functions?

Archit3ch··on DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
If it matches GPT-5.4 on coding tasks (as benchmarks suggest) this could be my forever model. And with partial SSD streaming, I could run it locally today. :D
Archit3ch··on Lisp moving Forth moving Lisp
Take for instance Pandoric Macros that allow you to peek inside the hidden state of a closure:

  # Traditional Let over Lambda
  counter = let x = 100
    () -> x += 1
  end
  counter() # 101
  
  # Pandoric Macro
  captured_x_ref = Core.getfield(counter, :x)
  println("The hidden state is: ", captured_x_ref.contents) # Outputs: 101
Archit3ch··on Steel Bank Common Lisp version 2.6.7
> I much prefer it to writing intrinsics in C (and the results are just as good). It's an interactive, exploratory, coding: I write small modular functions, SBCL compiles them on the fly, I glue them together with high-level language constructs.

Same here, but in Julia instead. Sometimes I drop down to LLVM intrinsics (e.g. to force lop3.lut on GPUs).

Page 1 of 13Next →