(If you check my HN post history you'll see I post a lot about Haskell. That's right, part of Groq's compilation pipeline is written in Haskell!)
(If you check my HN post history you'll see I post a lot about Haskell. That's right, part of Groq's compilation pipeline is written in Haskell!)
It may be caching or it didn't change the model being queried or something else.
They allow you to configure chat participants (a model + params like context or temp) and then each AI answers each question independently in-line so you can compare and remix outputs.
We also use Haskell on our infra team. Most of our CI infra is written in Haskell and Nix. Some of the chip itself was designed in Haskell (or maybe Bluespec, a Haskell-like language for chip design, I'm not sure).
It'll answer something along the lines of self-attention being O(n^2) (where n is the sequence length) because you have to compute an attention matrix of size n^2.
There are other attention mechanisms with better computational complexity, but they usually result in worse large language models. To answer jart: We'll have to wait until someone finds a good linear attention mechanism and then wait some more until someone trains a huge model with it (not Groq, they only do inference).
import urllib.request, json, math
for i in range(1, 20):
url = f"https://huggingface.co/mistralai/Mixtral-8x7B-v0.1/resolve/main/model-{i:05d}-of-00019.safetensors?download=true"
with urllib.request.urlopen(url) as r:
header_size = int.from_bytes(r.read(8), byteorder="little")
header = json.loads(r.read(header_size).decode("utf-8"))
for name, value in header.items():
if name.endswith(".weight"):
shape = value["shape"]
mb = math.prod(shape) * 2e-6
print(mb, "MB for", shape, name)
tome's other comment mentions that they use 568 GroqChips in total, which should be enough to fit even Llama2-70B completely in SRAM. I did not do any math for the KV cache, but it probably fits in there as well. Their hardware can do matrix-matrix multiplications, so there should not be any issues with BLAS. I don't see why they'd need other hardware.Could you clarify your question about hardware support? Currently we build out our hardware to support our cloud offering, and we sell systems to enterprise customers.
1. How many GroqCards are you using to run the Demo?
2. Is there a newer version you're using which has more SRAM (since the one I see online only has 230MB)? Since this seems to be the number that will drive down your cost (to take advantage of batch processing, CMIIW!)
3. Can TTS pipelines be integrated with your stack? If so, we can truly have very low latency calls!
*Assuming you're using this: https://www.bittware.com/products/groq/
2. We're working on our second generation chip. I don't know how much SRAM it has exactly but we don't need to increase the SRAM to get efficient scaling. Our system is deterministic, which means no need for waiting or queuing anywhere, and we can have very low latency interconnect between cards.
3. Yeah absolutely, see this video of a live demo on CNN!
Follow up (noob) question: Are you using a KV cache? That would significantly increase your memory requirements. Or are you forwarding the whole prompt for each auto-regressive pass?
I think currently 1. Unlike with graphics processors, which really need data parallelism to get good throughput, our LPU architecture allows us to deliver good throughput even at batch size 1.
At that price 568 chips would be $11.7M
> Accelerator Cards GroqCard low latency AI/ML Inference PCIe accelerator card with single GroqChip
If they can get costs down and put more dies into each card then it'll be business/consumer friendly.
Let's see if they can scale production.
Also, where tf is the next coral chip, alphabet been slacking hard.
Fortunately LLMs and hard work of clever peeps running em on commodity hardware are starting to make this possible anyway.
Because Google Home/Assistant just seems to keep getting dumber and dumber...
We achieve low latency by basically being a software-defined architecture. Our functional units operate completely orthoganal to each other. We don't have to batch in order to achieve parallelism and the system behaviour is completely deterministic, so we can schedule all operations precisely.
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
it's really amazing! the first time I tried the demo, I had to try a few prompts to believe it wasn't just an animation :)
I read an article that indicated your Bill Of Materials compared to NVidia's is 10x to get 1/10 the latency, and 8x BOM for throughput if Nvidia optimizes for throughput? Does this seem accurate? That CAPEX is the primary drawback?
https://www.semianalysis.com/p/groq-inference-tokenomics-spe...
I'm coding to NVidia right now. That builds them a moat. The instant I can get other hardware working, the less of a moat they will have. The more open it is, the more likely I am to adopt it.
Best-case would be something I buy for <$2k (if out-of-pocket) or under $5k (if employer). Next best case would be a cloud service with a limited free tier. It's okay if it has barely enough quota that I can develop to it, but the quota should never expire.
(The mistake a lot of services make is to limit free tier to e.g. 30 day or 1 year, rather than hours/month; if I didn't get around to evaluating, switch employers, switch projects, etc. the free tier is gone).
I did sign up for your API service. I won't be able to use it in prod before your (very nice) privacy guarantees are turned into lawyer-compliant regulatory language. But it's an almost ideal fit for my application.
My only point was to, well, perhaps bump this up from #100 on your personal priority list perhaps to #87, to the limited extent that influences your business.
1) Privacy and security. I work with PII.
2) Low-level access and doing things the manufacturer did not intend, rather than just running inference on Mixtral.
3) Knowing it will be there tomorrow, and I'm not tied to you. I'm more than happy to pay for hosted services, so long as I know after your next pivot, I'm not left hanging.
Why free tier?
I'm only willing to subsidize my employer on rare occasions.
Paying $12 for a prototype means approvals and paperwork if employer does it. I won't do it out-of-pocket unless I'm very sure I'll use it. I've had free tier translate into millions of dollars of income for one cloud vendor about a decade ago. Ironically, it never happened again, since when I switched jobs, my free tier was gone.
This approach is very good if you want to spend several millions building a large inference server to achieve the lowest latency possible. But it doesn't make sense for a lone customer buying a single card, since you wouldn't really be able to run anything on it.
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
* https://alcf.anl.gov/news/researchers-accelerate-fusion-rese...
* https://wow.groq.com/groq-accelerates-covid-drug-discovery-3...
Are chips and models obsoleted on roughly the same timelines?
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...