HNHacker News
TopNewBestAskShowJobs

kraddypatties

27 karma · joined November 10, 2025

Turning voice agents into lifelike video calls @ https://www.keyframelabs.com
submissionscomments
kraddypatties··on Natural Language Autoencoders: Turning Claude's Thoughts into Text
I think the only thing that gives me pause is the fact that they SFT on Opus 4.5 explanations as a pertaining step. But, generally I agree, especially given the auto encoder is only seeing a single token activation!
kraddypatties··on Natural Language Autoencoders: Turning Claude's Thoughts into Text
I believe that’s _part_ of the point (or at least a side-effect) of the KL divergence loss term they have on the AV. That and training stability.
kraddypatties··on Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
I can believe that in the long run.

Does the agent have access to arxiv (a brief skim of the README didn't have an answer)? If not, it could be that the current approach of relying on the model's weights only is resulting in the perceived local optimum of hyperparameter tuning.

Anecdotally, we built a little MCP for arxiv to help with our internal research, noticed a significant boost in the diversity of methods (architecture or otherwise) Claude and friends were able to reference.

kraddypatties··on Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
Hm, that's fair. It does feel like there's low hanging fruit in combining "old school" methods for conducting a hyperparameter sweep efficiently _with_ the higher level architecture edit ability of Autoresearch.

Probably would cut the number of runs down by a significant number (as far as I can tell it's doing a grid search once it decides to mess with a knob or section of the architecture).

kraddypatties··on Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
I feel like most of this recent Autoresearch trend boils down to reinventing hyper-parameter tuning. Is the SOTA still Bayesian optimization when given a small cluster? It was ~3 years ago when I was doing this kind of work, haven't kept up since then.

Also, shoutout SkyPilot! It's been a huge help for going multi-cloud with our training and inference jobs (getting GPUs is still a nightmare...)!

kraddypatties··on Show HN: Emotional photoreal AI humans at $0.06 / min
Glad you liked it!

Currently the avatar does it based on the text, which maps the incoming audio to one of our emotion codes, biasing the generation to that emotion. It's not foolproof, but we've found it works pretty well in practice.

kraddypatties··on Go away Python
I'm interpreting this as "uv was built off of years of PEPs", which is true; that being said the UX of `uv` is their own, and to me has significantly reduced the amount of time I spend thinking about requirements, modules, etc.
kraddypatties··on Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention
Running into "no healthy upstream" when navigating to the link -- hug of death maybe?
kraddypatties··on Tell HN: Happy Thanksgiving
been lurking for most of my adult life (and it shows :-))

Thanks HN! You make me smarter every (other) day.

kraddypatties··on Show HN: Realtime, expressive AI personas that you can video call
Thanks for trying it out!

Yea that latency makes sense; "listening" includes turn detection and STT, "thinking" LLM + TTS _and then_ our model, so the pipeline latency stacks up pretty quick. The actual video model starts streaming out frames <500ms from the TTS generation, but we're still working on reducing latency from parts of the pipeline that we are using off the shelf.

We have a high level blog post here https://www.keyframelabs.com/blog/persona-1 about the architecture of the video model, the WebRTC "agent" stack is Livekit + a few backend components hosted in Modal.

kraddypatties··on Ask HN: What Are You Working On? (Nov 2025)
We've been tinkering with building realtime talking head models (avatar models, etc.) for a while now, and finally have something that works (well enough)! Operates at ~2x realtime on a 4090, significantly faster than that on enterprise grade GPUs.

You can try it yourself at https://playground.keyframelabs.com/playground/persona-1 and there's a (semi)technical blog post at https://www.keyframelabs.com/blog/persona-1

The main use case we designed for was language learning, particularly having a conversational partner -- generally we've found that adding a face to the voice really helps trigger the fight or flight response, which we've found to be the hardest part of speaking a new language with confidence.

But in building out the system around the model to enable that use case (tool use on a canvas for speaking prompts and images, memory to make conversations less stale, etc.), we think there's potential for other use cases too.