The point of SRAM, especially at the L1/L2 level is having an extremely high BW and extremely low latency (a few clock cycles). So it is not really an option to put them somewhere else (although L3 and as mentioned other lower level layers) can and are already being put into either separate chiplets in the same PCB w/extremely fast ring OR directly on top of the die (3D stacking).
> One of the biggest bottlenecks is memory bandwidth. That is also not cheap or simple to do.
This is precisely why people are trying to put logic into memory instead of just making the logic chips simpler. Compute being 10x faster doesn't mean much when you want real-time, near-zero latency in the current day (and potentially, future) ML workloads. Memory bandwith for low batches are much more important, and even though this chip comes with HBM3E (which is cutting edge), that by itself won't make this faster than H200/MI300X.
This is amazing for finding cute collectibles from my favorite TV show that I would otherwise not noticed among random t-shirt and other "slap the picture and call it co-branded" products! I'm not super sure how long it is going to be around, but I think I'm gonna keep playing with it for a while.
Seems like especially for the last year or so, there have been a significant amount of interest in single-language oriented (instead of a single core library w/N language bindings, winking at a particular one) TUI libraries that are getting better and better (potentially because some of them were able to attract VC money). Two of them off top of my head is Textual (by textualize.io) for Python and BubbleTea (by charm.sh) for Go.
Any reason why HN source code is not published? The best I can think of is not to let people see the penalizing behavior, but having an open standard might actually help it improve rather than keeping it as a hidden secret that slowly gets discovered by independent malicious parties.
Excalidraw itself is also pretty cool as an embeddable widget, we recently built a cool real-time AI-accelerated drawing playground [1] (source [2]) and the experience was super fun!
One of the main bottlenecks for inference is memory bandwith (esp when dealing with huge models, like SD/SDXL) and for that, nothing I know of comes close to matching memory speeds on Apple Silicon (up to 400GB/s).
This is not a good comparison. Nvidia doesn't have a fab, but they are the lead player in the AI chip space. Intel had both and look where it got them. TSMC has a good model, and you can basically take any of your designs for the same node and manufacture it in any of their plants. Same strategy can be applied to Samsung, and they already help a lot on the memory segment. The new HBM3E memory chips for H200s might be even coming from Samsung.
It was just published a couple hours ago. Chips and cheese usually go back in time and post deep dives into old chips, not everything has to be cutting edge to get an insightful article.
Both announced and the expected inflation (especially after the changes in the central bank admin, since May) numbers have been globally recognized both by local independent groups (most known one is ENAG) and international financial institutions. Where did you end up with this idea of inflation number not meaning anything?
Interest rates are a mean to control the expected inflation, not the realized one. The expected inflation for 2024 is 36%, which means the current interest rates are sufficient to settle it.
Have you seen fal.ai/dynamic where you can perform image to image synthesis (basically editing an existing image with the help of diffusion process) using LCMs to provide a real time UI?
When the biggest chunk of your compensation is in the form of PPUs (profit participation units) which might be worthless under the new direction of the company (or worth 1/10th of what you think they were), it might be actually much more of an easier jump than people think to get some fresh $MSFT stock options which can be cashed regardless.
"Why Sam Altman (who can have the funding, talent, and the vision OpenAI has right now) can't just create OpenAI 2.0?" is an amazing question that also answers whats OpenAI's moat.
People speculated it was the funding, or attracting talent or having "access". Turns out it was none of them (obviously they all have a part, but having all three doesn't mean you can best OpenAI which gives you the fundemental reason why it is so hard to compete with them).
> Only a fraction of Microsoft’s $10 billion investment in OpenAI has been wired to the startup, while a significant portion of the funding, divided into tranches, is in the form of cloud compute purchases instead of cash, according to people familiar with their agreement.
VRAM is not the main constraint, is it? The computational power of any of the new graphical cards (beside from the highest end models, where the VRAM is the actual constraint, like RTX 4090s) is absurdly low on the stuff that actually matters (tensor cores, cuda cores, etc.). They are graphical cards, equipped with consumer grade VRAM (instead of HBM) and IMHO it will take a very big shift before we see them being used as real AI accelerators.
This is indeed a bit weird, one interesting update we released just now is allowying you to change the seed so see different variations of the same prompt+input combination. Also what I have noticed is, if you describe what you meant with a few words (two suns in the sky, etc.) it is actually pretty decent in terms of generational quality.
Seems like the underlying die is basically the same as H100, with just a wider memory bus (and potentially changed IMC?). Which is very nice to see, since with H100 especially for inference workloads, the main problem was always memory bandwith being the bottleneck for us. $/perf was never there compared to A100s. Assuming this replaces H100s and the price becomes somewhat similar, we might be finally able to utilize them for our own inference workloads.
> David Weigand, chief financial officer at Supermicro, said on the call that rack-scale and AI system sales accounted for 53 percent of revenues, which is $1.12 billion. Sales to large datacenter customers and OEM appliances to other vendors (of which we also think Nvidia is one) accounted for $1.17 billion, up 26.3 percent.
This explains pretty much the whole thing. People are building new racks at an unprecedented pace, it is like a gold rush, and anyone who is selling tools (whether it be TSMC's raw wafers, NVIDIA's chips, or even a company that only builds HVACs for datacenters) is probably going to see their peak.
I don't think L40/L40S are allowed to be exported when all the other AD102 (the underlying silicon) variants are banned including lower specced ones like 4090s.
> exceeding certain performance thresholds (including but not limited to the A100, A800, H100, H800, L40, L40S, and RTX 4090).
Would be happy to give you a hand at this and let you started using fal.ai with to run some of the more advanced diffusion models (with free credits if this is more of an hobby/art-style project, although SDXL with reasonable settings costs almost ~20X cheaper than OpenAI so you might not even need it)! Shoot me an e-mail at batuhan [at] fal.ai.
L40S might be a more appealing choice if you are switching from an A40, but then why the hell aren't you using A100 which is actually much cheaper and much more commonly available. The only perceivable reason I can see is fp8 support, but even with it, I don't think it is worth the price.
L40S sound good on paper, but the memory bandwith compared to even a 40G A100 is reduced in half which is crazy in the era when we are bottlenecked by it rather than the actual compute. It costs as much (or even more) than an 80G A100, but instead of getting ~2TB/s, you get ~800GB/s.
Especially for these kinds of rather artistic use cases, I'd rather see it using some of the open source models which might cost almost an order of magnitude less (disclaimer, I might be biased towards extremely good price/perf of the OS ones as someone who works and maintains a popular inference service for these kind of ML models).
Would be curious how the performance compares to DataFusion[0] as one of the top contenders to DuckDB on this area (albeit they being different in a lot of parts, I find it one of the closest compared to all others).
ClickBench (from ClickHouse) has some benchmarks[1] where it can be compared, but am not super sure how up to date it is. At least a while back, they were majorly out of date and haven't looked too closely on whether they are keeping it fair for everyone else :)
I had a project a couple years back implementing this and various other 'dead' (rejected) language-level proposals dynamically at the interpreter with just an install of a package, https://github.com/isidentical-archive/pepgrave. Was a fun experience looking back to the history of Python to see all these rejected proposals and getting them to play nice with the current language itself which has significantly evolved since they were proposed.
EdgeDB is by far one of my favorite startups in the developer tooling scene, and even though I don't get to use it actively on my day job am still watching it with awe from outside.
My only humble question is whether the lack of a "dynamic" query builder for Python will continue when the TS had it for so long? I understand the point of Python typing not being expressive enough to support this, but it's a language problem and I don't think EdgeDB should be the one trying to solve it. I'd love if I can just write my queries within my business logic directly and iterate without any sort of delays as opposed to writing them somewhere else (creating new files), generating the APIs, trying something, and repeating this whole process when I want to change anything.