A nice (if visually inscrutable) game of chess? An interactive Zelda-inspired crawler? Maybe let's just launch some fireworks?
Sharing partly because I’m promoting the workshop, but mostly because I simply think they’re fun.
429 karma · joined October 15, 2008
A nice (if visually inscrutable) game of chess? An interactive Zelda-inspired crawler? Maybe let's just launch some fireworks?
Sharing partly because I’m promoting the workshop, but mostly because I simply think they’re fun.
I’m the Head of Engineering at Starlight, and we’re hiring a senior full-stack engineer.
We build software that helps financial institutions connect people with government benefits they’re eligible for. More than $100B in benefits go unclaimed every year because the underlying systems are fragmented and difficult to navigate. We’re trying to change that.
Our engineering team is intentionally small, so this is a high-ownership role. You’ll work across the stack, help shape our architecture and engineering culture, and have significant influence over how and what we build.
One area I’m especially excited about is building an “AI-first” engineering organization. I put that in quotes intentionally because I don’t think the industry has figured out exactly what it means yet—and neither have we. I want this person to help me figure it out.
I’ll encourage experimentation and use of AI, and we’ll have lots of conversations about what works and what doesn’t. But I’m not setting token quotas or mandating that people use AI for everything. I care much more about figuring out where it genuinely makes us better.
I’m actively involved in the hiring process and happy to answer questions about the team, role, or company.
Job description (salary, tech stack, benefits, etc.):
- https://get-starlight.notion.site/Senior-Engineer-Full-Stack...
Company:
Before I'd looked that up I was going to say: I feel like "don't allow an invalid Unicode string to exist all" feels like a separate/bigger problem to me from "handling it fine" when they do get created. To the extent I can hand JavaScript an invalid combination of code units in a variety of other scenarios, returning a � felt fine.
e.g. // valid String.fromCodePoint(0xd83e, 0xdd20) // invalid, but "�" is ... fine? String.fromCodePoint(0xdd20, 0xd83e)
Yeah, I think that's fair. I didn't really think this through as I was writing it.
I'm not even so sure "ending up with nonsense" here is the worst outcome. It might be unavoidable with this approach and if that had been the only problem this bug might have been less memorable.
The real problem—which I mention didn't articulate/emphasize particularly well—was that these invalid surrogate pairs were getting passed into `encodeURIComponent` somewhere deep in the stack and choking catastrophically on them. That was the "real" bug at the end of the day, but the invalid surrogate pairs and the way they were getting created on the way were a fun journey to untangle.
If I'm remembering correctly, we briefly explored a solution where we told Python "This is a UTF-16LE encoded string" so the count would match, but I think we learned/realized the endianness is actually dictated by the client's machine (Going from memory here). Ultimately we just changed the solution so the client was the source of truth about lengths and counts.
These threads are surfacing all kinds of things I forgot about and didn't add in that blog post. Maybe I need to write another, haha.
It was really `encodeURIComponent` that didn't handle it gracefully.
If you just type this into the console (surrogate pair for cowboy smiley face emoji), you see it encodes it ("%F0%9F%A4%A0"):
encodeURIComponent("\uD83E\uDD20")
If you give it an invalid surrogate pair, it will throw an actual error:
encodeURIComponent("\uDD20\uD83E")
- https://george.mand.is/invalid-surrogate-pairs/
I thought it was something that's easier to play with and feel than necessarily just read about.
Clearly the next thing we need to test is removing all the vowels from words, or something like that :)
One thing that I thought was fairly clear in my write-up but feels a little lost in the comments: I didn't just try this with whisper. I tried it with their newer gpt-4o-transcription model, which seems considerably faster. There's no way to run that one locally.
I'm actually curious, if I run transcriptions back-to-back-to-back on the exact same audio, how much variance should I expect?
Maybe I'll try three approaches:
- A straight diff comparison (I know a lot of people are calling for this, but I really think this is less useful than it sounds)
- A "variance within the modal" test running it multiple times against the same audio, tracking how much it varies between runs
- An LLM analysis assessing if the primary points from a talk were captured and summarized at 1x, 2x, 3x, 4x runs (I think this is far more useful and interesting)
LLMs as the operating system, the way you interface with vibe-coding (smaller chunks) and the idea that maybe we haven't found the "GUI for AI" yet are all things I've pondered and discussed with people. You articulated them well.
I think some formats, like a talk, don't lend themselves easily to meaningful summaries. It's about giving the audience things to think about, to your point. It's the sum of storytelling that's more than the whole and why we still do it.
My post is, at the end of the day, really more about a neat trick to optimize transcriptions. This particular video might be a great example of why you may not always want to do that :)
Anyway, thanks for the time and thanks for the talk!
(Thanks for your good sense of humor)
> We do this internally with our tool that automatically transcribes local government council meetings right when they get uploaded to YouTube
Doesn't YouTube do this for you automatically these days within a day or so?
I don't think a simple diff is the way to go, at least for what I'm interested in. What I care about more is the overall accuracy of the summary—not the word-for-word transcription.
The test I want to setup is using LLMs to evaluate the summarized output and see if the primary themes/topics persist. That's more interesting and useful to me for this exercise.
There is just so much content out there. And context is everything. If the person sharing it had led with some specific ideas or thoughts I might have taken the time to watch and looked for those ideas. But in the context it was received—a quick link with no additional context—I really just wanted the "gist" to know what I was even potentially responding to.
In this case, for me, it was worth it. I can go back and decide if I want to watch it. Your comment has intrigued me so I very well might!
++ to "Slower is usually better for thinking"
Felt like a fun trick worth sharing. There’s a full script and cost breakdown.