HNHacker News
TopNewBestAskShowJobs

treesciencebot

1,949 karma · joined December 27, 2020

python, hot silicon and anything in between.
submissionscomments
treesciencebot··on Show HN: Open-Source Image Model Leaderboard with Public Preference Data
we considered which one adheres to the prompt more, which one has overall best aesthetics etc but ended up with a simple which one is overall better type question. it is easier for people to vote and decide one and still applicable as preference data at a larger scope (trading volume for simplicity).

the dataset is open source and we plan to train an aesthetics picker on it but obviously have to do proper evals (with at least 1M data) to come to a reasonable conclusion.

treesciencebot··on I'm Peter Roberts, immigration attorney who does work for YC and startups. AMA
How transferable the open source experience from major projects (like being a core developer of the python language itself) in terms of O1 to provide the criteria of "reviewing other's work"?
treesciencebot··on Show HN: Flash Attention in ~100 lines of CUDA
zero cost abstractions exist. doesn't mean all abstractions are zero-cost. or being zero-cost somehow invalidates their abstractness/genericness. but maybe we differ on the definition of abstractions.
treesciencebot··on Show HN: Flash Attention in ~100 lines of CUDA
triton the DSL.
treesciencebot··on Show HN: Flash Attention in ~100 lines of CUDA
Pretty neat implementation. In general, for these sort of exercises (and even if the intention is to go to prod with custom kernels) I lean towards Triton to write the kernels themselves. It is much more easier to integrate to the tool chain, and allows a level of abstraction that doesn't affect performance even a little bit while providing useful constructs.
treesciencebot··on Show HN: Hatchet – Open-source distributed task queue
Latency is really important and that is honestly why we re-wrote most of this stuck ourselves but the project with the gurantee of 25ms< looks interesting. I wish there was an "instant" mode where enough workers are available it could just do direct placement.
treesciencebot··on The hater's guide to Kubernetes
> Above I alluded to the fact that we briefly ran ephemeral, interactive, session-lived processes on Kubernetes. We quickly realized that Kubernetes is designed for robustness and modularity over container start times.

Is there a clear example of this? E.g. is kubernetes inherently unable to start a pod (assuming the same sequence of events, e.g. warm/cold image with streaming enabled) under 500ms, 1s etc?

I am asking this as someone who spent quite a bit of time and wasn't able to bring it down 2s< mark, which eventually led us to rewrite the latency sensitive parts to use Nomad. But we are currently in a state where we are re-considering kubernetes for its auxilary tooling benefits and would love to learn more if anyone had experiences with starting and stopping thousands of pods with the lowest possible latencies without caring for utilization or placement but just observable boot latencies.

treesciencebot··on Elon Musk sues Sam Altman, Greg Brockman, and OpenAI [pdf]
It allows for research to continue, which might eventually benefit everyone. The primary advantage in my mind is giving academy a chance to learn from it and community to build cool stuff on top of it.
treesciencebot··on Elon Musk sues Sam Altman, Greg Brockman, and OpenAI [pdf]
> (1) They can give away the model but sell an API - but they can’t serve a model as cheap as Goog/Msft/Amzn who have better unit economics on their cloud and better pricing on GPUs (plus custom inference chips).

Which has a simple solution, release the model weights with a license which doesn't let anyone to commercially host them (like AGPL-ish) without your permission. That is what Stability.ai does it.

treesciencebot··on Show HN: Real-time image generation with SDXL Lightning
pixel art is a particularly hard thing for these models to do, especially without further fine tunes or loras. but I'm pretty sure you should be able to get that quality with one of nerijs's loras [0]. But for now, i'd do some prompt templating and try some variations of this: 'pixel-art picture of a cat. low-res, blocky, pixel art style, 8-bit graphics'

[0]: https://huggingface.co/nerijs/pixel-art-xl

treesciencebot··on Show HN: Real-time image generation with SDXL Lightning
Spatial prompt adherence is a general missing piece is SDXL (or previous versions of the SD). Hoping that the SD will get it into a good shape as your examples!

Test the example on Stable Cascade as well (latest open-weight stability model), and yeah, even that is not great at it https://fal.ai/models/stable-cascade?share=eab44060-690b-497....

treesciencebot··on Show HN: Real-time image generation with SDXL Lightning
Yep, this is using SDXL Lightning underneath which is trained by ByteDance on top of Stable Diffusion XL and released as an open source model. In addition to that, it is using our inference engine and real-time infrastructure to provide a smooth experience compared to other UIs out there (which as far as I know, speed-wise, are not even comparable, ~370ms for 4 step here vs ~2-3 seconds in the replicate link you posted).
treesciencebot··on Stable Diffusion 3
Quite nice to see diffusion transformers [0] becoming the next dominant architecture on the generative media.

[0]: https://twitter.com/EMostaque/status/1760660709308846135

treesciencebot··on Groq runs Mixtral 8x7B-32k with 500 T/s
per-chip compute is not the main thing this chip innovates for fast inference, it is the extremely fast memory bandwith. when you do that, you'll loose all of that and will be much worse off than any off the shelf accelerators.
treesciencebot··on Groq runs Mixtral 8x7B-32k with 500 T/s
there are providers out there offering for $0 per million tokens, that doesn't mean it is sustainable and won't disappear as soon as the VC well runs dry. Am not saying this is the case for Groq, but in general you probably should care if you want to build something serious on top of anything.
treesciencebot··on Groq runs Mixtral 8x7B-32k with 500 T/s
The main problem with the Groq LPUs is, they don't have any HBM on them at all. Just a miniscule (230 MiB) [0] amount of ultra-fast SRAM (20x faster than HBM3, just to be clear). Which means you need ~256 LPUs (4 full server racks of compute, each unit on the rack contains 8x LPUs and there are 8x of those units on a single rack) just to serve a single model [1] where as you can get a single H200 (1/256 of the server rack density) and serve these models reasonably well.

It might work well if you have a single model with lots of customers, but as soon as you need more than a single model and a lot of finetunes/high rank LoRAs etc., these won't be usable. Or for any on-prem deployment since the main advantage is consolidating people to use the same model, together.

[0]: https://wow.groq.com/groqcard-accelerator/

[1]: https://twitter.com/tomjaguarpaw/status/1759615563586744334

treesciencebot··on Replicate vs. Fly GPU cold-start latency
Yep, we are currently in private beta for custom models. Hit us at hello@fal.ai for access!
treesciencebot··on Replicate vs. Fly GPU cold-start latency
Just as a top-level disclaimer, I'm working at one of the companies in "this" space (serverless GPU compute) so take anything I say with a grain of salt.

This is one of the things we (at https://fal.ai) working very hard to solve. Because of ML workloads and their multiple GB environments (torch, all those cuda/cudnn libraries, and anything else they pull) it is a real challange just to get the container to start in a reasonable time frame. We had to write our own shared Python virtual environment runtime using SquashFS distributed thru a peer-to-peer caching system to bring it down sub-second mark.

After the container boots, there is the aspect of storing model weights, which IMHO less challenging since it is just big blobs of data (compared to Python environments where there are thousands of smaller files where each might be sequentially read and incur a really major latency penalty). Distributing them once we had the system above was super easy since just like squashfs'd virtual environments, they are immutable data blobs.

We are also starting to play with GPUDirect on some of our bare metal clusters and hopefully planning to expose it to our customers, which is especially important if your models is 40GB or higher. At that point, you are technically operating at the PCIE/SXM speeds which is ~2-3 seconds for a model of that size.

treesciencebot··on Sora: Creating video from text
The examples are most certainly cherry-picked. But the problem is there are 50 of them. And even if you gave me 24 hour full access to SVD1.1/Pika/Runway (anything out there that I can use), I won't be able to get 5 examples that match these in quality (~temporal consistency/motions/prompt following) and more importantly in the length. Maybe I am overly optimistic, but this seems too good.
treesciencebot··on Sora: Creating video from text
If we go from DALL-E 3, it won't be nowhere near competitive while they have the superior ground. Generating a high quality 1024x1024 image with costs around ~$0.002, but $0.08 on DALL-E 3 (20x more expensive per-image). For videos with very high computational needs (since each frame needs to be temporally consistent, you need huge GPUs to serve this) I'm expecting this to be so much more expensive than its competitors (Pika or SVD1.1)
treesciencebot··on Sora: Creating video from text
This is leaps and bounds beyond anything out there, including both public models like SVD 1.1 and Pika Labs' / Runway's models. Incredible.
treesciencebot··on Fly.io has GPUs now
Just to correct the record, both $1.15 per A100 and $2.24 per H100 require a 3-year-commitment. On-demand prices are 2.5X that.
treesciencebot··on Stable Cascade
in my observation, it yields amazing perf at higher batch sizes (4 or better 8). i assume it is due to memory bandwith and the constrained latent space helping.
treesciencebot··on Stable Cascade
Uh, thanks for noticing it! We generally turn it off for popular models so people can see the underlying inference speed and the results but we forgot about it for this one, it should now be auth-less with a stricter rate limit just like other popular models in the gallery.
treesciencebot··on Stable Cascade
I think the model architecture (training code etc.) itself is still under MIT while the weights (which was the result of training in a huge GPU cluster as well as the dataset they have used [not sure if they publicly talked about it] is under this new license.
treesciencebot··on WiFi 7 is officially here, but routers are pricey. Do you need it yet
with MLO and the promised latency decreases, even at 300Mbps, it should significantly make a difference on how "snappy" everything is (and how reliable video calls are).
treesciencebot··on Power over fiber
I don't understand how people find 3V at 180mA usable, isn't it like 0.5 watts?
treesciencebot··on How big is YouTube?
Is 32,000 a good enough number to estimate the entirety of the Youtube’s video space? It felt to little for what they are trying to accomplish (especially when they started doing year by year analysis)
treesciencebot··on Most 16-year-olds don't have servers in their rooms
Something that miss is, unless you need tons of IO (in the form of SAS/SATA storage, or old generation PCIe cards), avoiding these huge, noisy and power hungry servers are a lot simpler than people may think. Mini/Micro form factor OEM PCs on the same price point generally come with much newer generation of hardware (instead of a 3rd gen i7, you might get a 8th gen i5) and overall performance is so much better. It’s also so much easier to host and maintain (just plug it near your ISP modem, and forget about it).
treesciencebot··on Apple to halt Apple Watch Series 9 and Ultra 2 sales in the US this week
Masimo, the company who is suing apple, has a market cap around ~$6B. Apple's wearable business (which might include the audio products as well, I guess), at least according to the article, "generated $13.48 billion in revenue". Don't think this is as big of a threat as someone like Samsung suing iPhone.
← PreviousPage 2 of 5Next →