HNHacker News
TopNewBestAskShowJobs

jakevoytko

3,669 karma · joined May 13, 2009

Staff Backend Engineer @ Hinge, recommendations team

Ex-Google, ex-Etsy

Writing: https://www.bitlog.com, https://www.clientserver.dev Email: jakevoytko@gmail.com Twitter: @jakevoytko

submissionscomments
jakevoytko··on Ask HN: What are you working on? (September 2026)
The only thing I would open source is the spec file from the Q/A so people can implement it if they'd like.

I find it extremely unlikely that I'd be able to sell it since its two main competitors are (1) literally free to use and has been polished for 16 years, and (2) a best-in-class native application that has been the lingua franca of business writing for over 40 years

jakevoytko··on Astra for Law
> all relevant facts will be cited and checked easily by humans

I've talked to a lawyer about how they handle this. They do indeed double-check everything, since it'd be embarrassing (or worse) to send hallucinated statements to opposing council or to the court. They still find the assembly a huge time saver

But based on stories in the news on the subject, not everyone has this same level of diligence

jakevoytko··on Can we stop with the uptime percentages?
These numbers are useful proxies for how likely you are to have your work disrupted outside of your own control.

If you do something 100 times a day against a four-nines service, you can reasonably expect that everything will succeed.

If you do something 10,000 times a day against a two-nines service, you can expect to hit a substantial number of errors during that day, or even have long periods where your work cannot happen at all.

People aren't frustrated with Github because Github has 98% uptime or whatever the specific number is. They're frustrated because it regularly interferes with their ability to work. The 98% number is just a concise way to say it.

jakevoytko··on Salesforce Global Outage
In my experience it’s a safe way to do something useful while everyone is getting their bearings. It immediately partitions the situation space between being persisted or systemic vs local or caused by long-running processes. Plus everyone’s going to ask if you’ve tried that already, so you might as well get it out of the way if it makes any amount of sense
jakevoytko··on XCancel service is suspended until further notice
I know there are like 5 trillion Instagram viewers. But I've never actually needed to use any of them because there is absolutely nothing so important on Instagram that I must see it. The entire site is optional.
jakevoytko··on XCancel service is suspended until further notice
Actual details.
jakevoytko··on XCancel service is suspended until further notice
I hate when companies do this. Instagram is a black box for me. Whenever someone sends me a link the web app is completely nonfunctional; clicking links just reload the page, it constantly tries opening the app, etc

The last time I signed up for an Instagram account (2016) it got immediately flagged as a bot account. I think I was supposed to appeal it to have it reinstanted, but why bother if this is how badly they treat users?

Can you imagine going to work as Instagram's web developer, spending your day finding new ways that the app accidentally works and disabling them?

jakevoytko··on Ask HN: What are you working on? (September 2026)
I built my own personal Google Docs clone to use as the editor for my blog

I used to work on the Google Docs team, so I basically did an 8-hour Q/A with Sol spanning like 400 questions. Did my best to reach back 10 years into my memories to see how things were implemented, and discussing architecture and tradeoffs. Then went through 5 ChatGPT resets implementing it on Astra to get a feel for the model.

I have a long way to go. But I'm writing my first post on it this week, and then I'll do a writeup of the editor project.

jakevoytko··on Ask HN: What default model do you use and why?
My workhorse is Opus orchestrating Sonnet or Sol orchestrating Luna, depending on which ecosystem I'm using this week. I'm a backend engineer, so for tasks that either interact with the machine learning or client codebases, I have Fable (Medium) review across the product spec, tech spec, and all the codebases after the implementation is done (haven't tried this workflow with Astra yet, don't know if it's a good reviewer).

As a side project, I'm doing an experimental task now (having Astra implement my own personal Google Docs clone, OT and everything, for my blog), and it feels more capable than Fable but needs to be watched closer than Fable. It is obsessed with verification and evidence to the point that I've needed to make some absolute rules to stop it from repeatedly e.g. making unit tests with 10,000 test cases with 5+ hours of runtime

For well-defined coding tasks, I'd default to either Luna or Sonnet, or hell, just writing it by hand.

jakevoytko··on On the Navier–Stokes Millennium Prize Problem
For full context, here's the HN thread from the other side of the "Concurrent Work" section: https://news.ycombinator.com/item?id=49605915

Unlike the vanilla read of the OpenAI press release, it is much more unfiltered and outlines some particularly aggressive behavior by specific OpenAI employees

jakevoytko··on How well do agents use test/verification techniques?
As someone who also spent a bunch of time benchmarking various techniques, this is as good as you can do without publishing a formal versioned benchmark suite that you want to maintain and run forever at immense cost to yourself. If you actually go to benchmark your own flow as you mentioned, you will quickly run into like a dozen problems that discourage you from publishing.

- Are you sure that temperature and other nondeterminism isn't affecting your output?

- Are you sure you're not being routed through an A/B test at this moment?

- Are you sure there's not a bug affecting the model at this moment?

- Are you sure that you picked the right model and effort level?

- Are you sure that your result generalizes across providers?

- Are you sure that you set up the correct level of sandboxing and the agent can't e.g. look at a sister directory or git history in the current directory for answers?

- Are you sure that the agent isn't leaking answers in memory or its conversation history?

- Are you sure that tool calls aren't somehow affecting results?

- Are you sure that your results are robust, i.e. you see the same results with mild tweaks to the prompt?

- Are you comfortable keeping your blog post live when your results are invalidated next week with the next model launch?

And that's just a quick list off the top of my head.

I personally decided that it wasn't worth it, I'm glad that Dan decided to publish his. Frankly I think we could use a lot more of these "I ran these 2 techniques side by side and here's what I saw" anecdata, because most people who promote prompt techniques can't produce a single prompt they ran twice because they never actually tested it per se.

Edited to add: formatting + the word "promote"

jakevoytko··on Finder is so frustrating and has been since day one
This was my instant reaction when I started using a Mac. But 15 years later, even knowing more tricks for interacting with it, it still baffles me. Once you've used Windows Explorer and a few Linux distro navigators, there's no denying that Finder is unquestionably worst in class.
jakevoytko··on GLP-1s are being linked to fewer serious infections, including TB
Yeah this is what I pay for Zepbound through LillyDirect. I also eat less food and drink less alcohol than I used to. So the net loss is probably smaller, maybe $200/mo
jakevoytko··on FBI Probes Service Selling 153M+ Drivers Licenses
As always, friendly reminder to lock your credit and enable your mobile carrier's protections against SIM swapping
jakevoytko··on Claude Fable 5.1 and Claude Mythos 5.1
My experience with output styles for long-running sessions is that Claude starts to forget the terse output style by the middle of the context window. Obviously I don't know if 5.1 suffers the same fate but I ran into this issue with both Opus and Fable 5
jakevoytko··on Bug Blindness
This reminds me of my first job, where we did a lot of 3d modeling and mobile robotics research. When we were trying to reproduce motion bugs, a coworker of mine would track down our manager and put him at the controls. More times than not, the bug would surface and it'd trip our logging and we tracked it down.

I asked him why he does this. His explanation was really built on this operant conditioning idea: "we use this stuff for 8 hours a day and we train ourselves to avoid all of its little pitfalls. So I get the most available person who hasn't used it all day, which is our manager. He uses it differently than we do because he doesn't avoid all of its little problems. But if we have a really tricky problem, I get our manager's manager. I don't know if you've ever seen him try to use an xbox controller, but he has the spatial reasoning abilities of a goldfish. He's never failed to reproduce a really hard bug. If I ever needed an Einstein-level bug reproduction, I'd track down the head of the department and put him in front of it, but it's never come to that.

jakevoytko··on Slack Code
They see the entire universe building their own version of Claude Tag and custom internal bots connected to your internal ecosystem, and realizing that they can fight for some enterprise revenue within their own product that everyone else is currently extracting
jakevoytko··on A 25-year-old video patent just expired, ending a legal headache for Linux
I worked at Etsy when it expired and the execs got this question a lot. The TL;DR is that the expected outcome is that it massively increases the support burden (wait I didn't mean to click that; wait I actually need to send it to another address; wait I didn't realize shipping was a hundred dollars) without really enabling more sales. So it was a neat idea and worth trying when ecommerce was new, but now we know enough about ecommerce to know that the user has to confirm details of their purchase.
jakevoytko··on "Solving a largely imaginary user goal"
Their argument is that the user almost always experiences the moment of interaction with the controls as a single decision with 2 branches, (a) the OS behavior matches my expectations and I won't change it, and (b) the OS behavior violates my expectations and I want it to be the other one.

I'd also rather just select from the full tristate diagram, but their framing also makes sense to me

jakevoytko··on Why does Opus 5 feel worse to work with?
It's clear that we're not the audience; it writes to be read by its training evaluator, not a professional software engineer. Professional software engineers can't read this word soup and are desperately trying to find ways to fix it.

It feels like it found a register that games the evaluator, where it can ramble forever and rarely be marked wrong while slowly racking up points as it talks more.

jakevoytko··on Writing by hand is good for your brain
I can take notes as fast as people can talk, but I'm much slower at writing so I needed to synthesize it into a condensed idea. It doesn't feel the same, to me at least.
jakevoytko··on Situational Awareness Down 67% in July in AI Stock Rout
If you’re gonna be a hater you at least gotta do it right! The pricing and the securities fraud were two separate things you can count against him.
jakevoytko··on Writing by hand is good for your brain
It kept me listening. When I was on minute 70 of a boring lecture I started daydreaming. If I was writing even minimal notes, I'd need to briefly consider each idea and decide whether it was worth recording.
jakevoytko··on Ask HN: Is it just me, or is software buggier across the board?
It's very simple; If you work 10x faster, you need to release 1/10th the bugs as before for the user's experience to have the same level of bugginess. And we're not incentivized for this; we're still exploring how to ship 10x faster and not exploring how to meet our users' needs 10x better (often conflating the two!)

And I mean, go tell your leadership "Those projects that used to take us a full quarter? Now we can do them in 6 weeks." You don't get the rest of the quarter to stabilize your codebase. Now you jam 2 releases into the quarter.

jakevoytko··on Ask HN: What Are You Working On? (July 2026)
This is great stuff; I've been prototyping with a few language-specific parsers like the Golang and the polyglot approach looks really helpful for me
jakevoytko··on Ask HN: What Are You Working On? (July 2026)
At the moment I'd check the sibling comment, which has a few links!
jakevoytko··on Ask HN: What Are You Working On? (July 2026)
My side project is now codebase explainability. I basically don't buy the premise that we just have to give up on comprehension as code generation scales; I just think that text is too limited by itself. So going a step deeper than asking Claudex "teach me this project", but having it produce a navigable snapshot of what's going on.

Big bang prototypes have been pretty awful, even after feeding the LLMs huge documents / wishlists / descriptions of how it should work, etc. Part of the experiment was giving LLMs some leeway to make product decisions with a lot of north star guidance, but AFAICT they are really bad at this. I also tried basic bottom-up efforts, which have been better but obviously more tedious. Now I'm trying to find a more scalable bottom-up approach that is more LLM-accelerated

jakevoytko··on GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps
Because for its entire existence, the top HN comment on articles is typically a contrarian take or pointing out flaws. This goes double for a study, where people just hunt for some aspect of the methodology they dislike. If you don't address the flaws, then it looks like you never considered them, and the top comment will say that your entire methodology is suspect. It's super predictable to the point that you can harness this kind of reaction to get stuff on the frontpage if you really want to.
jakevoytko··on FAANG Simulator
Out of curiosity, what has changed over time that has made it more toxic? I left Google in 2015.

It certainly had its share of toxic traits, but it was at the "what place doesn't?" level.

jakevoytko··on We built a persistent agent memory layer on Elasticsearch with 0.89 recall
Nah, "Any other vector DB" starts to fall apart once you need stuff like scripted scoring like OP uses. Then it starts to be a question of, "do you need ANN for performance?" since SQLite only does brute-force vector scoring. And granted, brute-force is performant for far more vectors than most people give it credit for, but it definitely hits a wall well below 1 million if you want it to have webpage-type latency.

Maintaining Elasticsearch isn't free, but picking an underpowered db and having to port to the right one is also quite time consuming.

Page 1 of 12Next →