HNHacker News
TopNewBestAskShowJobs

areddyyt

138 karma · joined January 17, 2023

@deepsilicon
submissionscomments
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Our CPU implementation for X86/AMD64 utilizes AVX-512 or AVX-2 instructions where possible. We're experimenting with support for ARM with NEON.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
100%.

When performing performance optimization on CPUs, I was impressed with Intel's suite of tools (like VTUNE). NVIDIA has some unbelievable tools, like Nsys and, of course, its container registry (NGC), which I think surpasses even Intel's software support.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
I should note that our linear layers are not the same as Microsoft's, in fact, we think Microsoft made a mistake in the code they uploaded. When I have time later today, I'll link to where I think they made a mistake.

I've been following TriLLM. They've achieved great results, and I'm really impressed with the llama.cpp contributors already getting the models integrated.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
I don't think I ever implied we started this for money. We started working on the technology because it was exciting and enabled us to run LLMs locally. We wouldn't have started this company if someone else came along and did it, but we waited a month or two and didn't see anyone making progress. It just so happens that hardware is capital intensive, so making hardware means you need access to a lot of capital through grants (which Dartmouth didn't have for chip hardware) or venture capital (which we're going for now). I'm not sure where you got the idea we're doing this solely for money when I explicitly said "We were essentially nerd-sniped into working on this problem"
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
I think there is a comment somewhere here where I comment on NVIDIA, but I think NVIDIA is the best hardware company for making good software. We had a very niche software issue for which NVIDIA maintained open-source repos. I don't think NVIDIA's main advantage is its hardware, though; I think it's the software and the flexibility it brings to its hardware.

Suppose that Transformers die tomorrow, and Mamba becomes all the rage. The released Mamba code already has CUDA kernels for inference and training. Any of the CSPs or other NVIDIA GPU users can switch their entire software stack to train and inference Mamba models. Meanwhile, we'll be completely dead in the water with similar companies that made the same bet, like Etched.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We were waiting for a Bitnet-based software and hardware stack, particularly from Microsoft, but it never did. We were essentially nerd-sniped into working on this problem, then we realized it was also monetizable.

On a side note, I deeply looked into every company in the space and was thoroughly unimpressed with how little they cared about the software stack to make their hardware seamlessly work. So, even if I did go to work at some other hardware company, I doubt a lot of customers would utilize the hardware.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We do quantization-aware training, so the model should minimize the loss w.r.t. the ternary weights, hence no degradation in performance.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
There was another founder that said this exact same thing. We'll definitely look into it especially as we train more ViTs.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Funnily enough, our ML engineer, Eddy, did a hackathon project working with Procyon to make a neural network with a photonic chip. Unfortunately, I think Lightmatter beat us to the punch.

Edit: I don't think the company exists in its current form anymore

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Have you sat in on my conversations with my cofounder?

The end plan is to have a single chip and flush all weights onto the chip at initialization. Because we are a single line of code that is Torch compatible (hence HF compatible), every other part of the codebase shouldn't change.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
This seems super cool. I'll have my cofounder look into it :)
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We don't achieve peak compression efficiency because more complex weight unpacking mechanisms kill throughput.

To be more explicit, the weight matrix's values belong to the set of -1, 0, and 1. When using two bits to encode these weights, we are not effectively utilizing one possible state:

10 => 1, 01 => 0, 00 =>-1, 11 => ?

I think selecting the optimal radix economy will have more of a play on custom silicon, where we can implement silicon and instructions to rapidly decompress weights or work with the compressed weights directly.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Thank you, and good catch.

We recently acquired deepsilicon.com, and it looks like the forwarding hasn't been registered yet. abhi@deepsilicon.net should work.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We actually were thinking about this to flush the weights in at initialization
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
It's always possible, but transformers have been around since 2017 and don't seem to be going anywhere. I was bullish on Mamba and researched extended context for structured state-space models at Dartmouth. However, no one cared. The bet we're taking is that Transformers will dominate for at least a few more years, but our bet could be wrong.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Video cropping issues should be fixed!
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We've spent a lot of time thinking about these things, in particular, the 3Ps.

Part of making the one line of code work is addressing programmability. If you're on Jetson, we should load the CUDA kernels for Jetson's. If you're using a CPU, we should load the CPU kernels. CPU with AVX512, load the appropriate kernels with AVX512 instruction, etc.

The end goal is that when we introduce our custom silicon, one line of code should make it far easier to bring customers over from Jetson/any other platform because we handle loading the correct backend for them.

We know this will be bordering impossible, but it's critical to ensure we take on that burden rather than shifting it to the ML engineer.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
You're absolutely right about mobile devices (Apple, Google, etc.). However, most companies, with the exception of Tesla, do use Nvidia for edge computing capabilities. We know for a fact that most of the automotive industry uses automotive rated Orins (the 32GB unified RAM SKU) [1] and Anduril also use Orins. Our primary GTM is with robotics companies, and we have not met a single robotics company not using Jetson, I'm not exaggerating.

[1] Particularly vehicles with advanced self driving capabilities. Qualcomm is another large vendor of hardware for vehicles (though they have even worse support)

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Oops, good catch. Will re upload shortly.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Great question. So a little bit of background about quantization (apologies if you are already familiar).

There are two types of quantization (generally), post training quantization (PTQ) and quantization aware training (QAT).

PTQ almost always suffers from some kind of accuracy degradation. This is because usually the loss is measured with respect to the FP16/BF16 parameters, and so the weights and distribution are selected to minimize the loss with those weights. Once the quantization function is applied, the weights and distribution change in some way (even if it's by a tiny amount), resulting in your model no longer being at minima.

We do QAT to get around the problem of PTQ. We actually quantize the weights during the forward pass of training, and measure the loss with respect to the quantized weights. As a result, once we converge the model, we have converged the ternary weights as well, and the accuracy it achieved at the end of training is the accuracy of the quantized model. At ~3B parameters the accuracy on downstream task performance between FP16 and ternary weights is identical.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Thank you!

CUDA and Nvidia are practically impenetrable on the server side. To be very concrete, we did training for our models on AWS with parallel cluster. We used P5 instances (8xH100) that were scheduled with SLURM. A problem we ran into however, was that our training jobs were containerized. Thankfully, pyxis and enroot exist to run containerized jobs on SLURM. And who else, but Nvidia, develop and maintain those plugins. For practically any weird niche use case, Nvidia seems to have some software solution - but only on x86.

Jetson is a whole other beast. There is no guarantee any pip package you install has an aarch64/arm64 wheel. For example, we could not use torch_tensorrt, to compile to TensorRT via Torch Inductor. Why? Because the Bazel build system was only configured to build for Jetpack 4.6 or Jetpack 5.1, and we were using Jetpack 6. While Nvidia provides docker images for x86 systems that come with torch_tensorrt installed, their L4T (Linux for Tegra) images do not. Instead we had to manually write out a new workspace file and compile for Jetpack6 to provide TensorRT compiling support.

tl;dr: Nvidia and CUDA have a great walled garden on x86, not so much on their edge computing devices

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We are not under the illusion these markets are easy to enter. Still, we believe providing an effortless and compatible experience for edge ML computing is a strong competitive advantage. We have not met anyone who likes using Jetsons yet, unlike A100/H100s in the server market.

Edit: I should note that if it weren't for Dusty and his docker image generating GitHub repo for Jetson, we would have spent weeks trying to get our kernels and optimized models shipped to customers.

areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
The non-linear layers, particularly the softmax(QK^T), will be crucial to getting ultra-low latency and high throughput. We're considering some custom silicon just for that portion of every transformer block
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
Agreed. We don't plan on making hardware until there is enough demand from customers to make it economically viable.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
In general, Jetson has quite a large market. Vehicle companies use automotive-rated Jetson Orins, and defense companies also use Jetson Orins to power ML applications on the edge (Anduril). Many of the companies we currently talk to are robotics companies that are forced to use Jetsons because they are both the least of the bad options and the only edge compute provider with enough juice to run larger transformer models.
areddyyt··on Launch HN: Deepsilicon (YC S24) – Software and hardware for ternary transformers
We're targeting the edge market first, such as NVIDIA's Jetson line, because it's far less supported/focussed on. In our experience, whenever we did training runs on H100 clusters with x86, any pip package would be easily installable, and a wide array of software just worked. This is not the case in Jetson, where we constantly have to rebuild packages from source, and in general, NVIDIA will only release a better board every five years. As for the second part of your question, we agree. Much of our work has been trying to make switching to our software layer straightforward (a single line of code). The ideal endgame is that, given an ONNX file, we can parse the generated node tree and determine if our hardware supports all the nodes. Of course, this is assuming we have a large enough share of the market using our software, so we know what operations we need to support on the hardware side of things.