HNHacker News
TopNewBestAskShowJobs

sethkim

441 karma · joined January 23, 2021

Founder of Sutro (https://sutro.sh/).

seth@sutro.sh

submissionscomments
sethkim··on LLM Ass Bench
Folks, we've reached the top.
sethkim··on The Analytical AI Handbook
I appreciate the classic HN sarcasm!
sethkim··on If DSPy is so great, why isn't anyone using it?
This is extremely true. In fact, from what we see many/most of the problems to be solved with LLMs do not have ground-truth values; even hand-labeled data tends to be mostly subjective.
sethkim··on If DSPy is so great, why isn't anyone using it?
Feel free to shoot me a note at seth@sutro.sh if you want to check it out!
sethkim··on If DSPy is so great, why isn't anyone using it?
We build a product that's somewhat similar in spirit to DSPy, but people come to us for different reasons than the OP listed here.

1) It's slow: you first have to get acquainted with DSPY and then get hand-labeled data for prompt optimization. This can be a slow process so it's important to just label cases that are ambiguous, not obvious.

2) They know that manual prompt engineering is brittle, and want a prompt that's optimized and robust against a model they're invoking, which DSPy offers. However, it's really the optimizer (ex. GEPA) doing the heavy-lifting.

3) They don't actually want a model or prompt at all. They want a task completed, reliably, and they want that task to not regress in performance. Ideally, the task keeps improving in production.

Curious if folks in this thread feel more of these pains than the ones in the article.

sethkim··on My trick for getting consistent classification from LLMs
Under-discussed superpower of LLMs is open-set labeling, which I sort of consider to be inverse classification. Instead of using a static set of pre-determined labels, you're using the LLM to find the semantic clusters within a corpus of unstructured data. It feels like "data mining" in the truest sense.
sethkim··on The Future (and Present) of AI Is Synthetic Data
The models you called out at the beginning were all released this year. What do you think is the difference between this generation of models and previous ones?
sethkim··on The End of Moore's Law for AI? Gemini Flash Offers a Warning
Yes! Both Llama 3 and Gemma 3 have 128k context windows.
sethkim··on The End of Moore's Law for AI? Gemini Flash Offers a Warning
Yes, we're a startup! And LLM inference is a major component of what we do - more importantly, we're working on making these models accessible as analytical processing tools, so we have a strong focus on making them cost-effective at scale.
sethkim··on The End of Moore's Law for AI? Gemini Flash Offers a Warning
My two cents here is the classic answer - it depends. If you need general "reasoning" capabilities, I see this being a strong possibility. If you need specific, factual information baked into the weights themselves, you'll need something large enough to store that data.

I think the best of both worlds is a sufficiently capable reasoning model with access to external tools and data that can perform CPU-based lookups for information that it doesn't possess.

sethkim··on The End of Moore's Law for AI? Gemini Flash Offers a Warning
No doubt prices will continue to drop! We just don't think it will be anything like the orders-of-magnitude YoY improvements we're used to seeing. Consequently, developers shouldn't expect the cost of building and scaling AI applications to be anything close to "free" in the near future as many suspect.
sethkim··on The End of Moore's Law for AI? Gemini Flash Offers a Warning
Both great points, but more or less speak to the same root cause - customer usage patterns are becoming more of a driver for pricing than underlying technology improvements. If so, we likely have hit a "soft" floor for now on pricing. Do you not see it this way?
sethkim··on Making 2.5 Flash and 2.5 Pro GA, and introducing Gemini 2.5 Flash-Lite
I run a batch inference/LLM data processing service and we do a lot of work around cost and performance profiling of (open-weight) models.

One odd disconnect that still exists in LLM pricing is the fact that providers charge linearly with respect to token consumption, but costs are actually quadratic with an increase in sequence length.

At this point, since a lot of models have converged around the same model architecture, inference algorithms, and hardware - the chosen costs are likely due to a historical, statistical analysis of the shape of customer requests. In other words, I'm not surprised to see costs increase as providers gather more data about real-world user consumption patterns.

sethkim··on Ask HN: Who is hiring? (June 2025)
Sutro.sh (fka Skysight) | Infrastructure/LLMs & Research Engineering | SF Bay Area | Full-time

We are building batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...

If you're interested in applying, please send an email to jobs@sutro.sh with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

sethkim··on Ask HN: Who is hiring? (May 2025)
Skysight | Infrastructure/LLMs & Research Engineering | SF Bay Area | Full-time

We are building large-scale batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...

If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

sethkim··on Gemini 2.5 Flash
How "huge" are these datasets? Did you build your own tooling to accomplish this?
sethkim··on Ask HN: Who is hiring? (March 2025)
Skysight | Infrastructure/LLMs & Product Engineering | SF Bay Area | Full-time

We are building large-scale, data-intensive inference tooling and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.

Our work involves fascinating distributed systems and ML research problems, newly-imagined and well-crafted user experiences, and a meaningful focus on mission and values.

Open Roles:

Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...

Product Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Product-...

If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.

sethkim··on TCAS Avoided Collision with Army Helicopter over Potomac River [video]
What's extremely confusing to me (as a private pilot) is that traffic is almost always routed directly over an airport (midfield), to safely avoid departing and landing traffic. The sense that I get is that it became routine for traffic to be routed directly through the glidepath in a staggered manner, likely because it's military.

Such unsafe habits (like driving without a seatbelt on) statistically will eventually result in a tragic outcome.

sethkim··on All text in Brooklyn
This is really cool, and hints at a near-future possibility of building a search engine on top of just about anything. It's clear we've moved past the ability to just search for website url's and webpage content. Anything that can be indexed - regardless of type of data or dimension (space, time, etc.) will be searchable.
sethkim··on Launch HN: Airhart Aeronautics (YC S22) – A modern personal airplane
I figured this comment would get me in trouble :)

I recommend doing some instrument lessons if you haven't already. When I got my instrument rating I questioned whether the private requirements are actually enough. The skills that the instrument rating teaches in terms of preparation, workload management, and emergency weather scenarios make me question whether I was really ready to fly before I had it.

With the private or sport license you'll be fine in the majority of cases. I think my comment comes more from the edge and corner cases that more skills and experience help with, not your ability to work in the system at all.

sethkim··on Launch HN: Airhart Aeronautics (YC S22) – A modern personal airplane
In single pilot IFR, an autopilot is often your best friend. It's exactly like you say - when you're busy with everything else you want the plane to fly itself. Isn't that problem already somewhat solved in a sense? Or are you referring to Garmin Autoland (or similar) in emergencies?

By the way - I can totally see how a great GA fly-by-wire system is an improvement to maintain positive control of an aircraft at all times. I'd personally love to give it a try and see how it reduces pilot effort while flying.

sethkim··on Launch HN: Airhart Aeronautics (YC S22) – A modern personal airplane
> What if this tech made the individual flyer safer?

I'd hope that's the case! That's why I put it in the "good" category".

> how much could be in the air at one time?

Hard to say, but there's a ton of congestion around busy airspace as is. I'd think an order-of-magnitude increase in GA traffic would require a major rework of the whole airspace system.

sethkim··on Launch HN: Airhart Aeronautics (YC S22) – A modern personal airplane
Instrument-rated pilot (and engineer) here.

First - congrats on the launch! I think you're working on an interesting set of components that will prove useful to GA aircraft technology. Bringing fly-by-wire, and lowering the cost of maintenance/manufacturing are both great efforts.

That being said, my personal view is that stick-and-rudder control is one of the less critical components to improving GA safety. Everything else - flight planning, comms, automation, navigation, weather, inspections, procedures, regs, and most importantly - working in the federal airspace system - are the "hard" parts of flying and where problems tend to occur. It's common belief that single-pilot IFR is the most challenging type of flying, because of how much you have to do all at once.

It may sound snobby - but I'm not super excited about the idea of lowering the barrier to entry for GA on a foundational skill basis. Like the light-sport rating, it encourages more people to be in the (already congested) airspace system who haven't really gained all the other skills necessary or experience to be there.

To be clear - I think improving technology and lowering costs = good. Lowering early-skill requirements for pilots and pushing more people without all the other skills into federal airspace = very bad. In general, I'd frame this effort more as an effort to raise the bar for system technology, not lower the bar to become a pilot in the first place.

sethkim··on OpenAI Acquires Multi
It became clear to me in the GPT-4o launch that OAI was interested in the "GitHub copilot for everything" route. With low-latency voice mode and the ability to take action on a user's behalf, this will basically feel like pair-programming for everything, or just having an employee who can do everything for you until they need your help or clarification.

To power that experience, an app will need to feel "multiplayer" like someone else is working with you. They'll probably bundle this in the API and have "agent mode" that developers can embed in any app or website, or just let consumers give OAI access to control their desktop. It'll also likely work async, so you can assign tasks, walk away or go to sleep for a few hours, and see the results.

This is speculation. But it feels like the interface that we'll look back on in ten years and say "that seemed obvious in hindsight."

sethkim··on Software Infrastructure 2.0: A Wishlist (2021)
What's cool is that Erik actually acted on these complaints. Modal is, by far, my favorite developer tool ever and makes me hopeful not just for the future of software engineering but the entire tech industry.

If you're a naysayer in the comments, I would encourage you to go give it an honest try, and consider again why you think infra has to be done in harder ways.

sethkim··on YC: Requests for Startups
Hey - can you shoot an email to seth@skysight.cloud? Curious what you've seen.
sethkim··on Ask HN: What have you built with LLMs?
I built https://tailgate.dev/ a few months ago. It can help with deployment of simple, client-facing generative web apps. There are a few simple demos on the home page!
sethkim··on HTML First
I built something recently with the same ideas in mind: https://github.com/sethkimmel3/tailgate. It allows people to build generative-AI applications without any of the complication of setting up a backend, and in the simplest cases only requires adding HTML attributes.

I originally learned to code with HTML, CSS, and JS, and I think it's still the easiest way to experience the magic of shipping a working application to others. We should keep encouraging more patterns and tooling that lower the barriers of entry to those just starting out.

sethkim··on Ask HN: Is GenAI heading for a crypto-esque bubble pop?
My two cents is that we will see a bubble pop, but it will rebound massively much like SaaS following the dotcom crash.

My primary concern is the ratio of AI-infrastructure products being built relative to the average utility delta of current end-user products. There are some shining counterexamples to this like Github copilot, but for the most part the utility gain from most text-based generative-AI seems incremental.

Current language models are bad to okay at most tasks - and great at a select few. The capabilities in unstructured ETL and production of labeled datasets fall in the latter category.

I think the next generation of really powerful capabilities to be unlocked with generative-AI will come with 1) an increase in compute availability (training even larger models) 2) advances in multimodal models, and 3) better interfaces to interact with them. Spatial computing will be a really big deal here, and people seem to have forgotten that Apple may have unlocked the next version of interactive computing earlier this year.

sethkim··on Hundreds of millions of stars turned into a map of GitHub projects
Very cool!

If you work on a dynamic version that allows users to understand changes in the open-source topography over time and detect/predict new clusters, this could be a very powerful tool for investment intelligence.

Page 1 of 3Next →