HNHacker News
TopNewBestAskShowJobs

philipkiely

1,210 karma · joined August 20, 2018

DevRel @ https://baseten.co

Email me: username at baseten.co

submissionscomments
philipkiely··on We got admin access to Baseten's production GitHub
Hey all Philip from Baseten here.

Posting this on behalf of our security team. I wanted to confirm that we collaborated with Strix on the remediation of the reported vulnerability. We thank Strix for their responsible disclosure. We took immediate steps to invalidate the leaked key and remove the public container image. Our logs confirm the vulnerability was never exploited and no customer data was exposed.

philipkiely··on The efficient frontier of LLM inference
I draw my diagrams on notecards and send them to our designer who brings them to life.

The images start out looking like this: https://philipkiely.com/images/blogs/how-to-write-a-book/des...

philipkiely··on The efficient frontier of LLM inference
I also wrote this as somewhat of a defense of the techniques that don't move the frontier -- there is a lot of value in being able to pick a point on the curve.
philipkiely··on The efficient frontier of LLM inference
These are both good points that I attempted to cover, quotes:

> In practice, the efficient frontier is very jagged. Rather than a smooth, continuous line between outcomes, small changes can have big impacts. These cutoff points are often unintuitive and must be discovered empirically through sweeps.

> However, quantization introduces a new set of tradeoffs between quality and serving efficiency. This is a particularly jagged frontier, where a large degree of improvement to serving efficiency is possible with little-to-no reduction in model quality, especially when using microscaling floating-point number formats like MXFP4 and NVFP4.

Would appreciate ideas on how to explain in greater depth

philipkiely··on The efficient frontier of LLM inference
I think the biggest net new recent technique is P/D disaggregation. And that spec dec is very different now especially post DSpark/DFlash.

But overall yes the fundamentals of LLM performance optimization have been remarkably stable over the last few years.

philipkiely··on Elevated error rate across multiple models
Good thing we have GLM-5.2
philipkiely··on The Math Behind TurboQuant
https://github.com/AliesTaha/polar_quant
philipkiely··on GLM-4.7: Advancing the Coding Capability
GLM 4.6 has been very popular from my perspective as an inference provider with a surprising number of people using it as a daily driver for coding. Excited to see the improvements 4.7 delivers, this model has great PMF so to speak.
philipkiely··on [dead]
The Information link, for those with a subscription: https://www.theinformation.com/articles/inference-provider-b...
philipkiely··on Why Fei-Fei Li and Yann LeCun are both betting on "world models"
You give it a text prompt and optional image.

What you get is a 3D room based on the prompt/image. It rewrites your prompt to a specific format. Overall the rooms tend to be detailed and imaginative.

Then you can fly around the room like in Minecraft creative mode. Really looking forward to more editing features/infill to augment this.

philipkiely··on Why Fei-Fei Li and Yann LeCun are both betting on "world models"
I played with Marble yesterday, Fei-Fei/World Labs' new product.

It is the most impressed I've been with an AI experience since the first time I saw a model one-shot material code.

Sure, its an early product. The visual output reminds me a lot of early SDXL. But just look at what's happened to video in the last year and image in the last three. The same thing is going to happen here, and fast, and I see the vision for generative worlds for everything from gaming/media to education to RL/simulation.

philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
We have built a ton of tooling on top of TRT-LLM and use it not just for LLMs but also for TTS models (Orpheus), STT models (Whisper), and embedding models.
philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Yeah the custom hardware providers are super good at TPS. Kudos to their teams for sure, and the demos of instant reasoning are incredibly impressive.

That said, we are serving the model at its full 131K context window, and they are serving 33K max, which could matter for some edge case prompts.

Additionally, NVIDIA hardware is much more widely available if you are scaling a high-traffic application.

philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Yeah we have tried to build calculators before it just depends so much.

Your equation is roughly correct, but I tend to multiply by a factor of 2 not 1.2 to allow for highly concurrent traffic.

philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
TRT-LLM has its challenges from a DX perspective and yeah for Multi-modal we still use vLLM pretty often.

But for the kind of traffic we are trying to serve -- high volume and latency sensitive -- it consistently wins head-to-head in our benchmarking and we have invested a ton of dev work in the tooling around it.

philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
This comment made my day ty! Yeah definitely speaking from a datacenter perspective -- fastest piece of hardware I have in the parts drawer is probably my old iPhone 8.
philipkiely··on Running GPT-OSS-120B at 500 tokens per second on Nvidia GPUs
Went to bed with 2 votes, woke up to this. Thank you so much HN!
philipkiely··on Qwen-Image: Crafting with native text rendering
For prod inference, 1xH100 is working well.
philipkiely··on Chatterbox TTS
WhisperX! https://github.com/basetenlabs/truss-examples/tree/main/whis...
philipkiely··on Chatterbox TTS
Example implementation with sample inference code + voice cloning example:

https://github.com/basetenlabs/truss-examples/tree/main/chat...

Still working on streaming

philipkiely··on An analysis of DeepSeek's R1-Zero and R1
Yeah so MoE doesn't really come into play for production serving -- once you are batching your requests you hit every expert at a large enough batch size so you have to think about running the models as a whole.

There are two ways we can run it:

- 8xH200 GPU == 8x141GB == 1128 GB VRAM

- 16xH100 GPU == 8x80GB == 1280 GB VRAM

Within a single node (up to 8 GPUs) you don't see any meaningful hit from GPU-to-GPU communication.

More than that (e.g. 16xH100) requires multi-node inference which very few places have solved at a production-ready level, but it's massive because there are way more H100s out there than H200s.

philipkiely··on An analysis of DeepSeek's R1-Zero and R1
They're using Llama.cpp which is an amazing tool for local inference but doesn't match fast inference frameworks like TensorRT-LLM/SGLang for production speeds and throughputs on Hopper GPUs.

The Unsloth quantizations are really cool, but if you want to experiment with the R1 models in a smaller form factor the R1 Distills like Llama 70B are great and should run a lot faster as they take advantage of existing optimizations around inferencing llama-architecture models.

philipkiely··on Llama-3.3-70B-Instruct
Just spent a few minutes this morning spinning up a H100 model server and trying an FP8 quantized version (including kv cache quantization) to fit it on 2 H100s -- speed and quality looking promising.

I'm excited to see if the better instruction following benchmarks improves function calling / agentic capabilities.

philipkiely··on Launch HN: Human Layer (YC F24) – Human-in-the-Loop API for AI Systems
Congrats Dex! Excited to see what people build with this + tools like Stripe's new agent payments SDK (issuing a payment seems like a great place to ask permission).
philipkiely··on Ask HN: Who is hiring? (October 2024)
Baseten | Tech Leads, Software Engineers, and more | SF, NY, Remote | Full-time

2 years and 8 months ago, I joined Baseten after seeing them in a "Who is Hiring" post on HN. Doing so was one of the three best decisions I've ever made in my life.

Baseten is now a fast-growing Series B AI infrastructure startup focused on inference. We have PMF and are growing fast in a highly competitive market.

We are actively hiring for 10+ roles at https://jobs.ashbyhq.com/baseten, and are especially focused on hiring:

- Tech lead, model performance (SF only): https://jobs.ashbyhq.com/baseten/ce8f46eb-40f9-4c46-98f7-c20...

- Tech lead, infrastructure (SF only): https://jobs.ashbyhq.com/baseten/ec31db6a-fe49-4b77-961d-c5e...

- Software engineer, ML inference and performance (SF, NY, Remote): https://jobs.ashbyhq.com/baseten/d29e748c-7209-460d-a024-8f7...

If you want to learn more about working at Baseten, check out https://www.baseten.co/blog/ten-reasons-to-join-baseten/

philipkiely··on Ask HN: How to transcribe a couple thousand calls per day?
This is a complete shameless plug but I just published some documentation on automatically building Whisper inference engines with TensorRT-LLM which has the batch inference that you're looking for: https://docs.baseten.co/performance/examples/whisper-trt
philipkiely··on Ask HN: Who is hiring? (July 2024)
Baseten | SF, NY, or REMOTE | Engineering | www.baseten.co

Baseten provides fast, scalable inference for AI/ML models. We're at Series B and growing fast with great customers like Descript, Bland, and Patreon. Cool tech, unbeatable team, huge market opportunity.

I joined Baseten 2.5 years ago after reading about it on an HN Who is Hiring thread!

We're actively hiring for:

* Support Engineer: https://jobs.ashbyhq.com/baseten/ff008b8e-b38d-4941-b24f-9a4...

* Forward Deployed Engineer: https://jobs.ashbyhq.com/baseten/84c1801c-1a65-49fb-aaaa-bee...

philipkiely··on Show HN: Baseten Chains – Framework and SDK for Multi-Model AI Products
I've been working with this for a couple of weeks ... there are real challenges in traditional setups going from prompt in -> prediction out to building the actual backends that blend business logic and inference for multiple models, large inputs, etc.

It's fun to work on solving these challenges with new tooling. Happy to answer any questions.

philipkiely··on Ask HN: Why would Google sell Google domains to Squarespace?
Does anyone know if this change is going to affect top-level domains operated by Google, for example .dev?
philipkiely··on Ask HN: Who is hiring? (March 2023)
Baseten | Software Engineers | Full-time | REMOTE (optional offices in SF, NYC)

I joined Baseten over a year ago after seeing a comment in "Who is hiring." Best decision of my career.

Baseten is a platform for deploying and serving ML models. We also just launched a platform for fine-tuning. Biggest parts of the stack are React, Django, Python, Kubernetes.

Baseten is a well-funded Series A with customers including Patreon, Pipe, Laurel, and Motive.

We're hiring for 2 senior-level technical roles right now:

Infrastructure: https://jobs.ashbyhq.com/baseten/5aef2919-5d0e-4a37-a04f-7d7...

Frontend: https://jobs.ashbyhq.com/baseten/fc6e5f2e-eb2d-4a6c-8a51-842...

Page 1 of 7Next →