HNHacker News
TopNewBestAskShowJobs

treesciencebot

1,949 karma · joined December 27, 2020

python, hot silicon and anything in between.
submissionscomments
treesciencebot··on Ask HN: Who is hiring? (May 2026)
fal | San Francisco, CA / Remote | Full-time

We are on a mission to build world’s first generative media platform for developers. We are running inference on tens of thousands of GPUs, and looking for people (in all functions) to help us scale it to hundreds of thousands.

Featured Roles:

- Distributed Systems Engineer, https://fal.ai/careers/4009192009

- Virtualization Engineer, https://fal.ai/careers/4146037009

- ML Performance & Systems Engineer, https://fal.ai/careers/4009191009

Remaining: https://fal.ai/careers

treesciencebot··on We ran over 600 image generations to compare AI image models
we build our sandbox just for this use case, fal.ai/sandbox. take the same image/prompt, and compare across tens of models.
treesciencebot··on FLUX.1 Kontext [Dev] – Open Weights for Image Editing
One interesting feature that gets enabled with open weights is adding new capabilities (tasks) to these editing models. They generalize quite well with low samples (30 ish). We talk about it here https://blog.fal.ai/announcing-flux-1-kontext-dev-inference-...
treesciencebot··on AMD's Freshly-Baked MI350: An Interview with the Chief Architect
the main question is going to be software stack. NVIDIA is already shipping NVFP4 kernels and perf is looking good. It took a really long time after MI300X's that the FP8 kernels were OK (not even good, compared to almost perfect FP8 support in NVIDIA side of things).

I will doubt that they will be able to reach %60-70 of the FLOPs in majority of the workloads (unless they hand craft and tune a specific GEMM kernel for their benchmark shape). But would be happy to be proven wrong, and go buy a bunch of them

treesciencebot··on Ask HN: Who is hiring? (June 2025)
fal | Growth Engineer | San Francisco (on site 5 days/wk)

Help us scale generative‑media infra: hack demos in the AM, pitch partners over coffee.

You’ll build quick client libs & microsites, run data A/Bs, write content that drives sign‑ups, and hand‑hold new devs.

Need: Python, JS/React/Next.js, SQL; speed, ownership, love for gen‑AI. Get: strong salary + equity, platinum health, unlimited “build‑something” stipend and most importantly a seat at a rocketship.

Shoot a link to something you’ve built to careers@fal.ai

treesciencebot··on World Emulation via Neural Network
author is: https://x.com/madebyollin
treesciencebot··on Apple M3 Ultra
GH200 is nowhere near $343,000 number. You can get a single server order around 45k (with inception discount). If you are buying bulk, it goes down to sub-30k ish. This comes with a H100's performance and insane amount of high bandwith memory.
treesciencebot··on Google Fiber is coming to Las Vegas
at a relatively new high-rise in rincon hill, AT&T still charges 80-90$ for 1 gig symmetrical (same with webpass/xfinity).
treesciencebot··on The AMD Radeon Instinct MI300A's Giant Memory Subsystem
For traditional LLMs this might be true (especially large MoEs at bs=1) but I highly disagree with "multi-modal models" phrase since most of the models that output in other modalities are generally compute bound. Which means less flops will make the experience so much worse (imagine waiting a couple minutes for an image and hours for videos).
treesciencebot··on FastVideo: a lightweight framework for accelerating large video diffusion models
For anyone that wants to test the original (non-distilled) HunyuanVideo (which is an amazing model) we have 580p version taking under a minute and 720p version taking around 2.5-3 minutes in our playground: https://fal.ai/models/fal-ai/hunyuan-video (it requires github login & and is pay-per-use but new accounts get some free credits).
treesciencebot··on Sora is here
Hunyuan at other providers like fal.ai is cheaper than SORA for the same resolution (720p 5 seconds gets you ~15 videos for $20 vs almost 50 videos at fal). It is slower than SORA (~3 minutes for a 720p video) but faster than replicate's hunyuan (by 6-7x for the same settings).

https://fal.ai/models/fal-ai/hunyuan-video

treesciencebot··on YouTube Premium Showing Ads
I think the main problem is with 2) since 1 is entirely up to the content creator's discretion.
treesciencebot··on Play Dialog: A contextual turn-taking TTS model like NotebookLM Playground
i don't think anyone has done real-time multi-speaker dialog generation before
treesciencebot··on Perplexity CEO offers AI company's services to replace striking NYT staff
Liability. Till we solve this, we cant really give AI any real responsibilities.
treesciencebot··on New SOTA text-to-image model by Recraft
It excels at particularly text and scene composition, as well as being able to generate vector graphics. You can use it through their website or through fal.ai https://fal.ai/models/fal-ai/recraft-v3/playground.
treesciencebot··on Play 3.0 mini – A lightweight, reliable, cost-efficient Multilingual TTS model
Much faster than OpenAI's real-time mode, wow! Quality seems to be on par if not better as well.
treesciencebot··on AMD's Turin: 5th Gen EPYC Launched
all high end "gaming" rigs are either using ~16 real cores or 8:24 performance/efficiency cores these days. threadripper/other HEDT options are not particularly good at gaming due to (relatively) lower clock speed / inter-CCD latencies.
treesciencebot··on Zed AI
running a much worse model at a higher latency (since local GPU power is limited) is a worse experience for Zed.
treesciencebot··on [dead]
FLUX.1 [schnell] (Apache 2.0): https://huggingface.co/black-forest-labs/FLUX.1-schnell

FLUX.1 [dev] (non-commercial, weights-available): https://huggingface.co/black-forest-labs/FLUX.1-dev (Mindtown has a commercial license so all the images generated through it can be used for whatever purpose)

treesciencebot··on Character.ai CEO Noam Shazeer Returns to Google
Was this an Inflection.ai style acquisition considering C.AI was profitable?
treesciencebot··on Flux: Open-source text-to-image model with 12B parameters
I think they are mainly -dev and -schnell. Both models are 12B. -pro is the most powerful and raw, -dev is guidance distilled version of it and -schnell is step distilled version (where you can get pretty good results with 2-8 steps).
treesciencebot··on Flux: Open-source text-to-image model with 12B parameters
i just updated the links to clarify which models require sign-in and which doesn't!
treesciencebot··on Flux: Open-source text-to-image model with 12B parameters
You can try the models here:

(available without sign-in) FLUX.1 [schnell] (Apache 2.0, open weights, step distilled): https://fal.ai/models/fal-ai/flux/schnell

(requires sign-in) FLUX.1 [dev] (non-commercial, open weights, guidance distilled): https://fal.ai/models/fal-ai/flux/dev

FLUX.1 [pro] (closed source [only available thru APIs], SOTA, raw): https://fal.ai/models/fal-ai/flux-pro

treesciencebot··on Black Forest Labs – FLUX.1 open weights SOTA text to image model
You can try the models here:

FLUX.1 [dev] (non-commercial, open weights, guidance distilled): https://fal.ai/models/fal-ai/flux/dev

FLUX.1 [schnell] (Apache 2.0, open weights, step distilled): https://fal.ai/models/fal-ai/flux/dev

FLUX.1 [pro] (closed source [only available thru APIs], SOTA, raw): https://fal.ai/models/fal-ai/flux-pro

treesciencebot··on Stable Audio Open
This looks like the one that got leaked a couple weeks ago, so i guess they decided its better to open source at this point after the leak [0].

[0]: https://x.com/cto_junior/status/1794632281593893326

treesciencebot··on Stability announces release date for Stable Diffusion 3 2B (small variant)
Just to make it clear, the actual Stable Diffusion 3 (8B variant, the one they are exclusively serving as an API) that they announced in march is still closed source and there are no indications on when or if it will be released.
treesciencebot··on Apple introduces M4 chip
~38 TOPS at fp16 is amazing, if the quoted number if fp16 (ANE is fp16 according to this [1] but that honestly seems like a bad choice when people are going smaller and smaller even at the higher level datacenter cards so not sure why apple would use it instead of fp8 natively)

[1]: https://github.com/hollance/neural-engine/blob/master/docs/1...

treesciencebot··on Apple introduces M4 chip
backlight is now the main bottleneck for consumption heavy uses. I wonder what are the main advancements that are happening there to optimize the wattage.
treesciencebot··on Show HN: Open-Source Image Model Leaderboard with Public Preference Data
i see about that case, and yeah you are right. we probably need realistic/artistic tags as you mentioned. thanks for the example! we'll probably include something like that in the next release and group models by ELO on different categories (can be considered like language analogue)
treesciencebot··on Show HN: Open-Source Image Model Leaderboard with Public Preference Data
comparing it to lmsys chatbot arena, what sort of an option would you expect? the prompts essentially come from public HF datasets like parti prompts where they test a bunch of stuff (prompt adherence, attention mapping [something in front of something else etc], aesthetics, photo-realism, etc.) so it is hard to ask about each category.
Page 1 of 5Next →