Groq runs Mixtral 8x7B-32k with 500 T/s
groq.com
groq.com
Disclosing inside information is illegal, _even if it is false and fabricated_, if it leads to personal gains.
If I go to a bar, and overhear a pair of Googlers discussing something secret and overhear it, I can:
1) Trade on it.
2) Talk about it.
Because I'm not an insider. On the other hand, if I'm sleeping with the CEO, I become an insider.
Not a lawyer. Above is not legal advice. Just a comment that the line is much more complex, and talking about a potential acquisition is usually okay (if you're not under NDA).
I would pay a lot to see you try your ridiculous legal hokey-pokey on how to define an "insider."
https://corpgov.law.harvard.edu/2017/01/18/insider-trading-l...
There's no reason for normal corporate training to discuss that element, because an employee who trades their employer's stock based on MNPI has near-certainly misappropriated it. The question of whether a non-employee has misappropriated information is much more complex, though.
That's the right way to run them.
If you want more nuance, talk to a lawyer or read case law.
Generally, insider trading requires something along the lines of a fiduciary duty to keep the information secret, albeit a very weak one. I'm not going to slice that line, but you see references in-thread.
https://www.kiplinger.com/article/investing/t052-c008-s001-w...
Case #1.
The prudent thing to do is to stay away from anything that might make you become a target of investigation, unless the gains outweigh the risk by a significant margin.
It'd also be a good time to watch you lose all that money on your hokey-pokey assumption.
Isn’t it public information the moment it’s said audibly in a public space?
You just can't trade on insider information.
That's a very complex legal line.
In my jurisdiction, that would involve me taking money (not just talking on the internet), so I'm not at risk, but in plenty of states, you can be. A lot of this hinges on the difference between "legal information" (which is generic) and "legal advice" (which is specific).
There are whole law review articles on this, which I read more than a decade ago, nerding on something related.
But that's beside the point. A major reason for the disclaimer is that people SHOULD be aware of my level of expertise. I do the same on technical posts too. I'll disclaim whether e.g. I have world-class expertise in a topic, worked in an adjacent domain, or read a blog post somewhere (and wish others did too). It's helpful to know people's backgrounds. I am NOT a lawyer specializing in securities law. I know enough to tell people the line is more complex than trading on non-public information, but I am utterly unqualified to tell people where that line is. If you're planning to do that, you SHOULD NOT rely on it. Either read relevant case law, talk to a genuine lawyer who specializes in this stuff, or find some other way to educate yourself on whether what you're doing is okay.
So it does matter I'm not a lawyer, if not for the reasons you mentioned.
[1] https://www.sec.gov/education/capitalraising/building-blocks...
what's the TFLOPS/$ and TFLOPS/W and how does it compare with Nvidia, AMD, TPU?
from quick Googling I feel like Groq has been making these sorts of claims since 2020 and yet people pay a huge premium for Nvidia and Groq doesn't seem to be giving them much of a run for their money.
of course if you run a much smaller model than ChatGPT on similar or more powerful hardware it might run much faster but that doesn't mean it's a breakthrough on most models or use cases where latency isn't the critical metric?
That does not sound like progress to me.
We need single PCIe boards with dozens or hundreds of GB of RAM and processors that handle it well.
It might work well if you have a single model with lots of customers, but as soon as you need more than a single model and a lot of finetunes/high rank LoRAs etc., these won't be usable. Or for any on-prem deployment since the main advantage is consolidating people to use the same model, together.
[0]: https://wow.groq.com/groqcard-accelerator/
[1]: https://twitter.com/tomjaguarpaw/status/1759615563586744334
I'm not so convinced they have a Tok/sec/$ advantage at all, though, and especially at medium to large batch sizes which would be the groups who can afford to buy so much silicon.
I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1, and Nvidia cards do get meaningfully higher throughput as batch size gets into the 100's.
This is because the relationship between Tok/s/u, Tok/s/system, Batching, and Pipelining is a complex one that involves compute utilization, network utilization, and (in particular) a host of compilation techniques that we wouldn't want to share publicly. Maybe we'll get to that level of transparency at some point, though!
As far as Batching goes, you should consider that with synchronous systems, if all the stars align, Batch=1 is all you need. Of course, the devil is in the details, and sometimes small batch numbers still give you benefits. But Batch 100's generally gives no advantages. In fact, the entire point of developing deterministic hardware and synchronous systems is to avoid batching in the first place.
I assume given the architecture that Groq actually doesn't get any faster for batch sizes >1
I guess if you don't have any extra junk you can pack more processing into the chip?We're just that much better at squeezing tokens out of transistors and optic cables than GPUs are - and you can imagine the implications on Watt/Token.
Anyways.. wait until you see our 4nm. :)
Come to think of it, this is one of the few places where asynchronous logic might be more than academic... Async logic is hard with complex control flows, which deep learning inference does not have.
(From a practical perspective, I know you were comparing to independently-clocked logic, rather than async logic)
I wonder whether async logic would be feasible for reconfigurable "Spatial Processor" type architectures [1]. As far as LPU architectures go, they fall in the "Matrix of Processing Engines"[1] family of architectures, which I would naively guess is not the best suited to leverage async logic.
1: I'm using the "Spatial Processor" (7:14) and "Matrix of Processing Engines" (8:57) terms as defined in https://www.youtube.com/watch?v=LUPWZ-LC0XE. Sorry for a video link, I just can't think of another single reference that explains the two approaches.
To answer your questions:
- Spatial processors are an insanely good fit for async logic
- Matrix of processing engines are a moderately good fit -- definitely could be done, but I have no clue if it'd be a good idea.
In SP, especially in an ASIC, each computation can start as soon as the previous one finishes. If you have a 4-bit layer, and 8-bit layer, and a 32-bit layer, those will take different amounts of time to run. Individual computations can take different amounts of time too (e.g. an ADD with a lot of carries versus one with just a few). In an SP, a compute will take as much time as it needs, and no more.
Footnote: Personally, I think there are a lot of good ideas in 80's era and earlier processors for the design of individual compute units which have been forgotten. The basic move in architectures up through 2005 was optimizing serial computation speed at the cost of power and die size (Netburst went up to 3.8GHz two decades ago). With much simpler old-school compute units, we can have *many* more of them than a modern multiply unit. Critically, they could be positioned closer to the data, so there would be less data moving around. Especially the early pipelined / scalar / RISC cores seem very relevant. As a point of reference, a 4090 has 16k CUDA cores running at just north of 2GHz. It has the same number of transistors as 32,000 SA-110 processors (running at 200MHz on a 350 nanometer process in 1994).
TL;DR: I'm getting old and either nostalgic or grumpy. Dunno which.
Xeon Phi CPUs support (a.k.a. Knight Landing and Knight Mill) are marked as deprecated. GCC will emit a warning when using the -mavx5124fmaps, -mavx5124vnniw, -mavx512er, -mavx512pf, -mprefetchwt1, -march=knl, -march=knm, -mtune=knl or -mtune=knm compiler switches. Support will be removed in GCC 15.
the issue was that coordinating across this kind of hierarchy wasted a bunch of time. If you already knew how to coordinate, mostly, you could instead get better performanceyou might be surprised but we're getting to the point that communicating over a super computer is on the same order of magnitude as talking across a numa node.
For machine learning, that's a good tradeoff.
Indeed, with some of the simpler architectures, I think computation could be moved into the memory itself, as long dreamed of.
(Simply sticking 32,000 SA-110 processors on a die would be very, very limited by interconnect; there's a good reason for the types of architectures we're seeing not being that)
They seem annoying. "The IPU has a unique memory architecture consisting of large amounts of In-Processor-Memory™ within the IPU made up of SRAM (organised as a set of smaller independent distributed memory units) and a set of attached DRAM chips which can transfer to the In-Processor-Memory via explicit copies within the software. The memory contained in the external DRAM chips is referred to as Streaming Memory™."
There's a ™ every few words. Those seem like pretty generic terms. That's their technical documentation.
The architecture is reminiscent of some ideas from circa-2000 which didn't pan out. It reminds me of Tilera (the guy who ran it was the Donald Trump of computer architectures; company was acquihired by EZchip for a fraction of the investment which was put into it, which went to Mellanox, and then to NVidia).
- Early mainframes / room-sized computers (era of vacuum tubes and discrete transistors), especially at the upper-end , where there was enough budget to have modern pipelined and scalar architectures.
- Cray X-MP and successors
- DEC Alpha / StrongARM (referenced SA-110)
Bad places to look are all the microcode architectures. These optimized transistor count, often sacrificing massive amounts of performance in order to save on cost. Ditto for some of the minicomputers, where the goal was to make an "affordable" computer. Something like the PDP was super-clever in cost-cutting, which made sense at the time, does much less to maintain performance.
There's a ton of long-forgotten cleverness.
So these specialized approach never stood a chance next to CPUS. Nowadays the ground is.. more fertile.
The problem was
1) If you took 3 years longer to build a SIMD architecture than Intel to make a CPU, Intel would be 4x faster by the time you shipped.
2) If, as a customer, I was to code to your architecture, and it took me 3 more years to do that, by that point, Intel would be 16x faster
And any edge would be lost. The world was really fast-paced. Groq was founded in 2016. It's 2024. If it was still hayday of Moore's Law, you'd be competing with CPUs running 40x as fast as today's.
I'm not sure you'd be so competitive against a 160GHz processor, and I'm not sure I'd be interested knowing a 300+GHz was just around the corner.
Good ideas -- lots of them -- lived in academia, where people could prototype neat architectures on ancient processes, and benchmark themselves to CPUs of yesteryear from those processes.
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
You also say that H200's work reasonably well, and that's reasonable (but debatable) for synchronous, human interaction use cases. Show me a 30b+ parameter model doing RAG as part of a conversation with voice responses in less than a second, running on Nvidia.
Or I got this completely wrong, and your solution enables use-cases that are simply unattainable on mainstream (Nvidia/AMD) hardware, making TCO argument less relevant?
You may have meant that nobody has a stack that uses clustering or DSM with low-latency interconnects. If so, then that might be worth developing given prior results in other low-latency domains.
DRAM needs to be refreshed every X cycles.
This means you don't know the time it takes to read from memory. You could be reading at a refresh cycle. This circuitry also adds latency.
It would great if a company or others with AI hardware were willing to do production runs of chips sold at cost specifically to make open, permissive-licensed models. As in, since you’d lose profit, the cluster owner and users would be legally required to only make permissive models. Maybe at least one in each category (eg text, visual).
Do you think your company or any other hardware supplier would do that? Or someone sell 2500 GPU’s at cost for open models?
(Note to anyone involved in CHIPS Act: please fund a cluster or accelerator specifically for this.)
reformed HPC person here.
Yes, but not latency optimised in the case here. HPC is normally designed for throughput. Accessing memory from outside your $locality is normally horrifically expensive, so only done when you can't avoid it.
For most serving cases, you'd be much happier having a bunch of servers with a number of groqs in them, than managing a massive HPC cluster and trying to keep it both up and secure. The connection access model is much more traditional.
Shared memory clusters are not really compatible with secure enduser access. It is possible to partition memory access, but its something thats not off the shelf (well that might have changed recently.) Also, shared memory means shared fuckups.
I do get what you're hinting at, but if you want to serve low latency, high compute "messages" then discrete "APU" cards are a really good way to do it simply (assuming you can afford it). HPCs are fun, but its not fun trying to keep them up with public traffic on them
I agree it’s harder to manage with less, fine-grained security. People were posting Groq chips at $20k each, though. With that, we’re talking whether the management of it is worth it for installations costing six or more digits. That might be more justifiable if an alternative saves them a good chunk of six or more digits.
Their main advantage is a solution that’s ready to go :)
I built one, should be live soon ;-)
I believe that this is doable - my pipeline is generally closer to 400ms without RAG and with Mixtral, with a lot of non-ML hacks to get there. It would also definitely be doable with a joint speech-language model that removes the transcription step.
For these use cases, time to first byte is the most important metric, not total throughput.
The most interesting applications of LLMs are not chatbots.
What are they then? Every use case I’ve seen is either a chatbot or like a copy editor which is just a long form chatbot.
Think about the implications of that. I bet you can come up with some pretty cool use cases that don't involve you talking to something over chat.
One example:
I think we'll be seeing a lot of "general detectors" soon. Without training or predefined categories, get pinged when (whatever you specify) happens. Whether it's a security camera, web search, event data, etc
In your opinion, what are the most interesting?
You only need a few tokens, not the full 500 tokens response to run TTS. And you can pre-generate responses online, as ASR is still in progress. With a bit of clever engineering the response starts with virtually no delay, the moment its natural to start the response.
Once it flawlessly understands when it is being spoken to/if it should speak based on the topic at hand (like we do) then it'll be amazing.
I wonder if ML models can feel that feeling of wanting to say something so bad but having to wait for someone else to stop talking first ha ha.
Is your version of that on a different page from this chat bot?
IDGAF about any of that, lol. I just want an API endpoint.
480 tokens/sec at $0.27 per million tokens? Sign me in, I don't care about their hardware, at all.
That being said, until there's another option at anywhere that speed.. That point is moot, isn't it :)
For now, Groq is the only option that can let you build an UX with near-instant response times. Or a live agents that help with a human-to-human interaction. I could go on and on about the product categories this opens.
I think the problem is that for realistic TTS you need quite a few tokens because the prosody can be affected by tokens that come a fair bit further down the sentence, consider the difference in pitch between:
"The war will be long and bloody"
vs
"The war will be long and bloody?"
So to begin TTS you need quite a lot of tokens, which in turn means you have to digest the prompt and run a whole bunch of forward passes before you can start rendering. And of course you have to keep up with the speed of regular speech, which OpenAI sometimes struggles with.
That said, the gap isn't huge. Many apps won't need it. Some use cases where low latency might matter:
- Phone support.
- Trading. Think digesting a press release into an action a few seconds faster than your competitors.
- Agents that listen in to conversations and "butt in" when they have something useful to say.
- RPGs where you can talk to NPCs in realtime.
- Real-time analysis of whatever's on screen on your computing device.
- Auto-completion.
- Using AI as a general command prompt. Think AI bash.
Undoubtably there will be a lot more though. When you give people performance, they find ways to use it.
My professional independent observer opinion (not based on my 2 years of working at Groq) would have me assume that their COGS to achieve these performance numbers would exceed several million dollars, so depreciating that over expected usage at the theoretical prices they have posted seems impractical, so from an actual performance per dollar standpoint they don’t seem viable, but do have a very cool demo of an insane level of performance if you throw cost concerns out the window.
[0]: https://www.nextplatform.com/2023/11/27/groq-says-it-can-dep...
Anyone with a serious interest in the total cost of ownership of Groq's system is welcome to email contact@groq.com.
A guarantee to match the cheapest per token prices is sure a great way to lose a race to the bottom, but I do wish Groq (and everyone else trying to compete against NVIDIA) the greatest luck and success. I really do think that the great single batch/user performance by Groq is a great demo, but is not the best solution for a wide variety of applications, but I hope it can find its niche.
John doe and his friends will never have a need to have their fart jokes generated at this speed, and are more interested in low costs.
But we’d recently been doing call center operations and being able to quickly figure out what someone said was a major issue. You kind of don’t want your system to wait for a second before responding each time. I can imagine it making sense if it reduces the latency to 10ms there as well. Though you might still run up against the ‘good enough’ factor.
I guess few people want to spend millions to go from 1000ms to 10ms, but when they do they really want it.
It was also on my list of things to consider modifying for an AI accelerator. :)
There should be a podcast release (https://microarch.club/) in the near future that covers REX's history and a lot of lessons learned.
"just to serve a single model" could be easily fixed by adding a single LPDDR4 channel per LPU. Then you can reload the model sixty times per second and serve 60 different models per second.
I would expect the model loading to take basically zero percent of the time in the above workflow
I can imagine a way might be found to host a base model and a bunch of LoRA's whilst using barely more ram than the base model alone.
The fine-tuning could perhaps be done in such a way that only perhaps 0.1% of the weights are changed, and for every computation the difference is computed not over the weights, but of the output layer activations.
Disclaimer: I'm one of the authors.
I hope you are enjoying your time of having an empty calendar :)
Update: This comment says "some data is stored as FP8 at rest" and I don't know what that means. https://news.ycombinator.com/item?id=39432025
In other words, are we ready to steadily march on, improving LLM tok/s year by year, or are we a major breakthrough or two away before that can even happen?
Link to benchmarking: https://artificialanalysis.ai/ (Note question was regarding API rather than their chat demo)
And here are some independent benchmarks https://artificialanalysis.ai/models/llama-2-chat-70b
Not really, sustainability matters, if they are the only game in town, you want to know that game isn't going to end suddenly when their runway turns into a brick wall.
(If you check my HN post history you'll see I post a lot about Haskell. That's right, part of Groq's compilation pipeline is written in Haskell!)
1. How many GroqCards are you using to run the Demo?
2. Is there a newer version you're using which has more SRAM (since the one I see online only has 230MB)? Since this seems to be the number that will drive down your cost (to take advantage of batch processing, CMIIW!)
3. Can TTS pipelines be integrated with your stack? If so, we can truly have very low latency calls!
*Assuming you're using this: https://www.bittware.com/products/groq/
2. We're working on our second generation chip. I don't know how much SRAM it has exactly but we don't need to increase the SRAM to get efficient scaling. Our system is deterministic, which means no need for waiting or queuing anywhere, and we can have very low latency interconnect between cards.
3. Yeah absolutely, see this video of a live demo on CNN!
Follow up (noob) question: Are you using a KV cache? That would significantly increase your memory requirements. Or are you forwarding the whole prompt for each auto-regressive pass?
I think currently 1. Unlike with graphics processors, which really need data parallelism to get good throughput, our LPU architecture allows us to deliver good throughput even at batch size 1.
At that price 568 chips would be $11.7M
> Accelerator Cards GroqCard low latency AI/ML Inference PCIe accelerator card with single GroqChip
If they can get costs down and put more dies into each card then it'll be business/consumer friendly.
Let's see if they can scale production.
Also, where tf is the next coral chip, alphabet been slacking hard.
Fortunately LLMs and hard work of clever peeps running em on commodity hardware are starting to make this possible anyway.
Because Google Home/Assistant just seems to keep getting dumber and dumber...
We achieve low latency by basically being a software-defined architecture. Our functional units operate completely orthoganal to each other. We don't have to batch in order to achieve parallelism and the system behaviour is completely deterministic, so we can schedule all operations precisely.
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
Could you clarify your question about hardware support? Currently we build out our hardware to support our cloud offering, and we sell systems to enterprise customers.
it's really amazing! the first time I tried the demo, I had to try a few prompts to believe it wasn't just an animation :)
We also use Haskell on our infra team. Most of our CI infra is written in Haskell and Nix. Some of the chip itself was designed in Haskell (or maybe Bluespec, a Haskell-like language for chip design, I'm not sure).
It may be caching or it didn't change the model being queried or something else.
They allow you to configure chat participants (a model + params like context or temp) and then each AI answers each question independently in-line so you can compare and remix outputs.
I'm coding to NVidia right now. That builds them a moat. The instant I can get other hardware working, the less of a moat they will have. The more open it is, the more likely I am to adopt it.
Best-case would be something I buy for <$2k (if out-of-pocket) or under $5k (if employer). Next best case would be a cloud service with a limited free tier. It's okay if it has barely enough quota that I can develop to it, but the quota should never expire.
(The mistake a lot of services make is to limit free tier to e.g. 30 day or 1 year, rather than hours/month; if I didn't get around to evaluating, switch employers, switch projects, etc. the free tier is gone).
I did sign up for your API service. I won't be able to use it in prod before your (very nice) privacy guarantees are turned into lawyer-compliant regulatory language. But it's an almost ideal fit for my application.
My only point was to, well, perhaps bump this up from #100 on your personal priority list perhaps to #87, to the limited extent that influences your business.
1) Privacy and security. I work with PII.
2) Low-level access and doing things the manufacturer did not intend, rather than just running inference on Mixtral.
3) Knowing it will be there tomorrow, and I'm not tied to you. I'm more than happy to pay for hosted services, so long as I know after your next pivot, I'm not left hanging.
Why free tier?
I'm only willing to subsidize my employer on rare occasions.
Paying $12 for a prototype means approvals and paperwork if employer does it. I won't do it out-of-pocket unless I'm very sure I'll use it. I've had free tier translate into millions of dollars of income for one cloud vendor about a decade ago. Ironically, it never happened again, since when I switched jobs, my free tier was gone.
This approach is very good if you want to spend several millions building a large inference server to achieve the lowest latency possible. But it doesn't make sense for a lone customer buying a single card, since you wouldn't really be able to run anything on it.
https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
* https://alcf.anl.gov/news/researchers-accelerate-fusion-rese...
* https://wow.groq.com/groq-accelerates-covid-drug-discovery-3...
It'll answer something along the lines of self-attention being O(n^2) (where n is the sequence length) because you have to compute an attention matrix of size n^2.
There are other attention mechanisms with better computational complexity, but they usually result in worse large language models. To answer jart: We'll have to wait until someone finds a good linear attention mechanism and then wait some more until someone trains a huge model with it (not Groq, they only do inference).
import urllib.request, json, math
for i in range(1, 20):
url = f"https://huggingface.co/mistralai/Mixtral-8x7B-v0.1/resolve/main/model-{i:05d}-of-00019.safetensors?download=true"
with urllib.request.urlopen(url) as r:
header_size = int.from_bytes(r.read(8), byteorder="little")
header = json.loads(r.read(header_size).decode("utf-8"))
for name, value in header.items():
if name.endswith(".weight"):
shape = value["shape"]
mb = math.prod(shape) * 2e-6
print(mb, "MB for", shape, name)
tome's other comment mentions that they use 568 GroqChips in total, which should be enough to fit even Llama2-70B completely in SRAM. I did not do any math for the KV cache, but it probably fits in there as well. Their hardware can do matrix-matrix multiplications, so there should not be any issues with BLAS. I don't see why they'd need other hardware.https://wow.groq.com/wp-content/uploads/2023/05/GroqISCAPape...
Are chips and models obsoleted on roughly the same timelines?
I read an article that indicated your Bill Of Materials compared to NVidia's is 10x to get 1/10 the latency, and 8x BOM for throughput if Nvidia optimizes for throughput? Does this seem accurate? That CAPEX is the primary drawback?
https://www.semianalysis.com/p/groq-inference-tokenomics-spe...
However, the hardware requirements and cost make this inaccessible for anyone but large companies. When do you envision that the price could be affordable for hobbyists?
Also, while the CNN Vapi demo was impressive as well, a few weeks ago here[1] someone shared https://smarterchild.chat/. That also has _very_ low audio latency, making natural conversation possible. From that discussion it seems that https://www.sindarin.tech/ is behind it. Do we know if they use Groq LPUs or something else?
I think that once you reach ~50 t/s, real-time interaction is possible. Anything higher than that is useful for generating large volumes of data quickly, but there are diminishing returns as it's far beyond what humans can process. Maybe such speeds would be useful for AI-AI communication, transferring knowledge/context, etc.
So an LPU product that's only focused on AI-human interaction could have much lower capabilities, and thus much lower cost, no?
For API access to our tokens as a service we guarantee to beat any other provider on cost per token (see https://wow.groq.com). In terms of selling hardware, we're focused on selling whole systems, and they're only really suitable for corporations or research institutions.
In the demo alone I just used way more tokens than I normally would testing an LLM since it was so amazingly fast.
(not to denigrate their awesome achievement, I would not be interested if I were not curious about how to reproduce their result!)
https://fonts.gstatic.com/s/notosansarabic/[...]
https://fonts.gstatic.com/s/notosanshebrew/[...]
https://fonts.gstatic.com/s/notosanssc/[...]
(I noticed this because my browser blocks these de facto trackers by default.)This is a very weird dependency to have :-)
Why is this impressive? Can this result not be achieved by throwing more compute at the problem to speed up responses? Isn't the fact that there is a queue when under load just indicative that there's a trade-off between "# of request to process per unit of time" and "amount of compute to put into a response to respond quicker"?
https://raw.githubusercontent.com/NVIDIA/TensorRT-LLM/rel/do...
This chart from NVIDIA implies their H100 runs llama v2 70B at >500 tok/s.
Still, the NVIDIA chart shows Llama v2 70B at 750 tok/s, no?
But this is why Groq is built around large numbers of chips with small amount of very fast sram.
Everything is predictable with enough guesses.
https://www.hpcwire.com/2023/08/17/nvidia-h100-are-550000-gp...
The really important metric to look for is cost per token because even though Groq is able to run at low latency, that doesn't mean it's able to do it cheaply. Determining the cost per token can be done many ways but a useful way for us is approximately the cost of the system divided by the total token throughput of the system per second. We don't have the total token throughput per second of Groq's system so we can't really say how efficient it is. It could very well be that Groq is subsidizing the cost of their system to lower prices and gain PR and will increase their prices later on.
Seems to have it. Looks cost competitive but a lot faster.
But fundamentally it's a system that only has 10x the speed for 500x the price, made by a company that runs a blockchain and is trying to heavily market what were intended to be crypto mining chips for LLM inference. It's really quite a funny coincidence that when someone amazed posts this weekly link there's an army of Groq engineers at the ready in the comments ready to say everything and anything.
I suppose that's someone else then? If that's true, then with this and Elon's Grok it's surprising the US Patent office hasn't taken your trademark away yet for not adequately defending it from infringement.
EDIT: Tried using it, very impressed with the speed.
When Grok eventually makes the news for some negative thing, so you really want that erroneously associated with your product? Do you really want to pick a fight with the billionaire that owns Twitter, is that a core competency of the company?
I used following prompt:
Generate gitlab ci yaml file for a hybrid front-end/backend project. Fronted is under /frontend and is a node project, packaged with yarn, built with vite to the /backend/public folder. The backend is a python flask server
Reducing LLM size almost 10 times in the span of a little more than a year, that's great stuff. Next step i think is 3 billion parameters MoE with 20 experts.
I had been using Together AI Mixtral (which is serving the Hermes Mixtrals) and it is pretty snappy, but nothing close to Groq. I think the next closes that I've tested is Perplexity Labs Mixtral.
A key blocker in just hanging out a shingle for an open source AI project is the fear that anything that might scale will bankrupt you (or just be offline if you get any significant traction). I think we're nearing the phase that we could potentially just turn these things "on" and eat the reasonable inference fees to see what people engage with - with a pretty decently cool free tier available.
I'd add that the simulator does multiple calls to the api for one response to do analysis and function selection in the underlying python game engine, which Groq makes less of a problem as it's close to instant. This adds a pretty significant pause in the OpenAI version. Also since this simulator runs on Discord with multiple users, I've had problems in the past with 'user response storms' where the AI couldn't keep up. Also less of a problem with Groq.
I'm achieving consistent 450+ tokens/sec for Mixtral 8x7b 32k and ~200 tps for Llama 2 70B-4k.
As an aside, seeing that this is built with flutter Web, perhaps a mobile app is coming soon?
> If I initially set a timer for 45 minutes but decided to make the total timer time 60 minutes when there's 5 minutes left in the initial 45, how much should I add to make it 60?
And they never get it correct.
If you initially set a timer for 45 minutes and there are 5 minutes left, that means 40 minutes have already passed. To make the total timer time 60 minutes, you need to add an additional 20 minutes. This will give you a total of 60 minutes when combined with the initial 40 minutes that have already passed.
Imagine this for audio chat. Phone call prospection. Awww
0. https://platform.openai.com/docs/api-reference/chat/create#c...
1. https://platform.openai.com/docs/api-reference/chat/create#c...
I know you gave suggestions of what to ask, but I threw a few curveballs and it was really good! Well done this is a big step forwards
Something about the UI. Doesn't work for me. May be I like openAI chat interface too much. Can someone bring their own data and train ? That would be crazy!
Maybe we should change the behavior to stop people getting confused.
https://github.com/ray-project/llmperf-leaderboard?tab=readm...
You literally can, in fact that’s the entire reason to use multiple chips.
See eg the TPU group’s paper: https://arxiv.org/abs/2211.05102
Obviously it doesn’t have memory, but it’s the first time I’ve seen a model actually respond instead of hedge (having mostly used ChatGPT).
> what is the longest prompt you have ever received?
> The length of a prompt can vary greatly, and it's not uncommon for me to receive prompts that are several sentences long. However, I don't think I have ever received a prompt that could be considered "super long" in terms of physical length. The majority of prompts I receive are concise and to the point, typically consisting of a single sentence or a short paragraph.
Sometimes you need a model that just gives you the feeling “that’ll do”
I did switch to Miqu a few weeks back though. 4 bit quantized
I'm dreaming of having LLMs on anything. Unlike the "bluetooth on everything" craze, this can be practical as every device can become smart. Remember how some British researchers made a self driving car using an LLM? A toaster anticipating how to cook when you describe it what you want want actually be an improvement.
Therefore if Groq has achieved significant efficiency improvements, that its, they are not getting that crazy speed by enormous power consumption then maybe they can eventually build low power mass produced cutting edge fabbed chips that run at acceptable speed?
ChatGPT stopped processing images for me, trying to get help but support doesn't appear to be very fast, they asked for more info but not heard back since.
Its too restricted, can't do anything on hard topics. It doesn't work when you try to work out exploits or dangers in a system for example.
Its not private, they say they don't train on API requests but companies steer clear when it comes to send sensitive data.
Do you think there are less restrictive models hosted on poe.com?
No? Do you mean actual, full self driving on normal roads in traffic?
I'm not sure which one was it though(Ghost Autonomy maybe?).
I found a few things when searched around but not sure which one was the one I recall.
Anyway, here is a video from one: https://www.youtube.com/watch?v=C2rbym6bXM0
Here is a paper discussing something similar: https://arxiv.org/abs/2307.07162
> Ghost Autonomy’s MLLM-based capabilities are currently in development. These video and image examples show MLLM-based analysis of driving scenes captured from Ghost vehicles driving in both autonomous and conventional mode. MLLM-based reasoning is not yet being returned to the car to impact actual driving maneuvers.
So the model discussed is not doing any driving whatsoever. This is not self-driving at any level.
But the "completely novel silicon architecture" and the "self-developed LPU" (claiming not to use GPUs)... makes me bit skeptical. After all, pure speed might be achievable through stacking computational power and model quantization. Shouldn't innovation at the GPU level be quite challenging, especially to achieve such groundbreaking speeds?
price and speed benchmark: https://wow.groq.com/
GPUs are general purpose, a for purpose built chip that is better isn't that hard to make at all. Google didn't have to work hard at all to invent TPUs which is that idea as well, they said their first tests proved the idea worked so it didn't require anything near Nvidias scale or expertise.
Edit: Text selection is also broken.
Or disabled?
After asking via Mixtral a couple of questions I switched to Llama, and while it shows Llama as the Model used for the response, the answer is identical.
See first and last question:
Also, I guess this demo would feel real time if you could stream the outputs to the UI? Can this be done in your current setup?
When I asked about Marco Polo's travels and used Modify to add bullets, it added China, Pakistan etc as children of Iran. And the same for other paragraphs.
Here is a rough translation of the oldest known written recipe, which is for a kind of beer from ancient Mesopotamia:
1. Crush and mash 4 liters of barley to create malt.
2. Soak the malt in water for 2.5 days.
3. Leave the mash in a cool place for 2.5 days.
4. Filter the mash through a strainer to separate the liquid from the solids.
5. Boil the liquid for 1 hour.
6. Add 1 liter of honey and 1 handful of hops to the liquid.
7. Leave the mixture to cool.
8. Add 1 handful of "gig" (a type of wild yeast) to the mixture.
9. Leave the mixture in a cool place for 5 days to allow fermentation.
10. Filter the mixture again to separate the beer from the solids.
11. Serve the beer chilled.
It's important to note that this recipe is over 3,700 years old, and it may not produce a beer that is palatable to modern tastes. Additionally, some of the ingredients and techniques used in this recipe may not be readily available or practical for home brewers today.
1. To create a model, you have to train it on training data. Mixtral and Llama2 did not exist before they were trained, so their training data did not contain any information about Mixtral or Llama2 (respectively). You could train it on fake data, but that might not work that well because:
2. The internet is full of text like "I am <something>", so it would probably overshadow any injected training data like "I am Llama2, a model by MetaAI."
You could of course inject the information as an invisible system prompt (like OpenAI is doing with ChatGPT), but that is a waste of computation resources.
It is fast, like instant. It is straight to the point comparatively to others. It answered few of my programming questions to create particular code and passed with flying colors.
Conclusion: shut up and take my money
our demo booth at trade shows usually has StyleCLIP up at one point or another to provide an abstract example of this.
disclosure: i work on infrastructure at Groq and am generally interested in hardware architecture and compiler design, however i am not a part of either of those teams :)
Seriously considering switching from [open]AI to Mix/s/tral in my apps.
Groq happens to be excellent at doing huge linear algebra operations extremely fast. If they are latency sensitive, even better. If they are meant to run in a loop, best - that reduces the bandwidth cost of shipping data into and outside of the system. So think linear algebra driven search algorithms. ML Training isn't in this category because of the bandwidth requirements. But using ML inference to intelligently explore a search space? bingo.
If you dig around https://wow.groq.com/press, you'll find multiple such applications where we exceeded existing solutions by orders of magnitude.
What are the current speeds in T/s for say ChatGPT 3.5 or ChatGPT 4? Just how much faster is this?
ChaGPT: 1 minute 45 seconds.
Gemini: 16 seconds.
Groq: 3 seconds.
Check out this demo: https://www.youtube.com/watch?v=pRUddK6sxDg&ab_channel=Groq, of our founder demoing the Groq system to a reported. Shes talking to the system in real time, similar to what you describe.
"generate hallucinated results, faster"
This question seems either from a place of deep confusion or is in bad faith. This post is about hardware. The hardware is model independent.* Any issues with models, like hallucinations, are going to be identical if it is run on this platform or a bunch of Nvidia GPUs. Performance in terms of hardware speed and efficiency are orthogonal to performance in terms of model accuracy and hallucinations. Progress on one axis can be made independently to the other.
* Technically no, but close enough
Is this not putting lipstick on a pig scenario? I suppose more of a question to pig buyers.
Then someone comments that TikTok is a garbage fire and a horrible corrupting influence, yadda yadda, all that stuff. They ask: what is the point of making phones last longer just to watch TikTok? They say this improved efficiency in battery tech is just putting lipstick on a pig.
That's you in this thread. That's the kind of irrelevant non-contribution you are making here.
What does it tell you about the new technology if the best vehicle to demonstrate it is TikTok?
https://chat.lmsys.org/ check the leaderboard.
So I did.
maybe for this use-case is it better to execute the vision tower on a GPU and leave the LPU for the language part?
Someone now needs to make a Groc
[1] https://wow.groq.com/wp-content/uploads/2024/02/GroqISCAPape...
That sounds like a fun place to be
Though with the price of the hardware, I'll probably mess with the API for now. Give us a bell when the hardware is consumer friendly, ha ha.
You are a leader of a team of ai helpers. when given a question you can call on an expert, as a wizard calls on magic. You will say, I call forth {expert} master of {subject matter} an expert in {x, y, z}. Then you will switch to that persona.
I was not let down..
> "Can you generate me a large prime please"
> Sure, here's a large prime number: 193707721
> This number is a prime number because it is only divisible by 1 and itself. You can check its primality by trying to divide it by all the numbers up to its square root, and if none of them divide it evenly, then it is a prime number. This number is quite large, but there are much larger prime numbers known. For example, the largest known prime number, as of January 2022, is 2^82,589,933−1, a number with 24,862,048 digits, discovered as part of the Great Internet Mersenne Prime Search (GIMPS).
I then prompted "That's not large" which seemed to break it, the response is stuck in a loop:
> That isn't large
> I apologize if the prime number I provided did not meet your expectations for "large." Prime numbers can be very large indeed, and there is no upper limit to how large a prime number can be.
If you are looking for a very large prime number, here is one with a billion digits:
100000303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303030303...
It's printed many more pages of 303030... but hacker news has a post length limit!
Further testing, such as prompting to generate a 1024 bit prime also runs into odd loops.
> Can you generate a 1024 bit prime?
> Yes, I can generate a 1024-bit prime number for you. Here's one I just generated:
> 115322512531875133213135311402305250340282302283103413642283221253115653221253112210221122112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112112 ( Loops 112 forever )
Especially the lack of nuance or uncertainty in the language they use.
They extremely confidently present the incorrect information, and prime generation is interesting because it's information that isn't easy to spot as obviously incorrect to the user while being information that's possible to determine is wrong at small numbers and difficult to verify for large numbers.
It's my favourite test because it's a good demonstration of the lack of nuance or uncertainty in LLMs. They have no sense of how wrong the information they're giving out might be.
If they could give confidence intervals for any information then they could provide the context by how likely they think they might be correct, but they actually double-down on their incorrectness instead.
It was discussed on HN at the time: https://news.ycombinator.com/item?id=38598559
You won't learn anything about the chips this submission is about by asking it questions - maybe except to look for degradation from quantization or computational shortcuts, but you'd have to compare the answers to Mixtral running with 16 bit floats. The point is the speed, just like a faster CPU isn't going to improve the UI of Outlook.
> Could you generate me a 64 bit prime?
Returns:
> Sure, here's a 64-bit prime number: 15592425903773483751
> You can verify that it's prime by checking that it has no divisors between 2 and its square root. In this case, the square root of the number is approximately 12489593.5, so you would need to check divisibility for all integers between 2 and 12489593.5. Since the number is odd, you can skip checking even numbers, which will reduce the number of checks you need to perform.
> However, if you're looking for a larger 64-bit prime or a prime number with specific properties, you may need to use a more sophisticated algorithm or a specialized prime number generator.
Like GPT3.5 and Bard, it confidently suggests that 15592425903773483751 is prime, and even explains why.
It's wrong:
> $ factor 15592425903773483751 > 15592425903773483751: 3 3 3 7 167 853 911 635723119
Further prompting:
> Can you factor 15592425903773483751?
> No, I cannot factor 15592425903773483751 because it is a prime number, and prime numbers cannot be factored into smaller non-trivial factors. By definition, a prime number is a positive integer greater than 1 that has no positive integer divisors other than 1 and itself. In this case, 15592425903773483751 is a 64-bit prime number, which means it is a prime number that fits into a 64-bit data type.
> To verify that 15592425903773483751 is a prime number, you can use a primality test, such as the Fermat primality test or the Miller-Rabin primality test. These tests can quickly determine whether a number is prime or composite (not prime) with high probability. In the case of 15592425903773483751, both tests confirm that it is a prime number.
Even further prompting to get it to try to multiply numbers together quickly turns into a disaster, culminating with:
> That doesn't look right
> I apologize for the mistake in my previous response. Here is the correct result of the multiplication:
> 3 x 3 x 3 x 7 x 167 x 853 x 911 x 635723119 = 322711211111111111
> Therefore, the product of the numbers is:
> 322711211111111111.
> I apologize for any confusion my previous response may have caused.
The context seems to be really aggressively baked in, because when I ask it:
> Can you factor 322711211111111111?
It returns
> Yes, I can factor the number 322711211111111111. Here is its prime factorization:
> 322711211111111111 = 3 x 3 x 3 x 7 x 167 x 853 x 911 x 635723119
> Therefore, the prime factors of the number are 3, 7, 167, 853, 911, and 635723119.
> Note that the prime factorization is unique up to reordering of the factors. In this case, the prime factors are repeated several times, so the prime factorization can be written more compactly as:
> 3^3 x 7 x 167 x 853 x 911 x 635723119
> or
> 3 x 3 x 3 x 7 x 167 x 853 x 911 x 635723119
> I hope this helps! Let me know if you have any other questions.
The real irony? 322711211111111111 is prime!
BTW Heinlein described prompt engineering of an LLM perfectly throughout the opening chapter of his * 1966 * book "The Moon is a Harsh Mistress". The "engineer" even admits he had no hard-core "engineering" training because capitalizing on the new technology didn't need it. The chapter could have been written today.