441 karma · joined January 23, 2021
seth@sutro.sh
1) It's slow: you first have to get acquainted with DSPY and then get hand-labeled data for prompt optimization. This can be a slow process so it's important to just label cases that are ambiguous, not obvious.
2) They know that manual prompt engineering is brittle, and want a prompt that's optimized and robust against a model they're invoking, which DSPy offers. However, it's really the optimizer (ex. GEPA) doing the heavy-lifting.
3) They don't actually want a model or prompt at all. They want a task completed, reliably, and they want that task to not regress in performance. Ideally, the task keeps improving in production.
Curious if folks in this thread feel more of these pains than the ones in the article.
I think the best of both worlds is a sufficiently capable reasoning model with access to external tools and data that can perform CPU-based lookups for information that it doesn't possess.
One odd disconnect that still exists in LLM pricing is the fact that providers charge linearly with respect to token consumption, but costs are actually quadratic with an increase in sequence length.
At this point, since a lot of models have converged around the same model architecture, inference algorithms, and hardware - the chosen costs are likely due to a historical, statistical analysis of the shape of customer requests. In other words, I'm not surprised to see costs increase as providers gather more data about real-world user consumption patterns.
We are building batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.
Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.
Open Roles:
Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...
Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...
If you're interested in applying, please send an email to jobs@sutro.sh with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.
We are building large-scale batch inference infrastructure and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.
Our work involves interesting distributed systems and LLM research problems, newly-imagined user experiences, and a meaningful focus on mission and values.
Open Roles:
Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...
Research Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Research...
If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.
We are building large-scale, data-intensive inference tooling and a great/user developer experience around it. We believe LLMs have not yet been meaningfully unlocked as data processing tools - we're changing that.
Our work involves fascinating distributed systems and ML research problems, newly-imagined and well-crafted user experiences, and a meaningful focus on mission and values.
Open Roles:
Infrastructure/LLM Engineer — https://jobs.skysight.inc/Member-of-Technical-Staff-Infrastr...
Product Engineer - https://jobs.skysight.inc/Member-of-Technical-Staff-Product-...
If you're interested in applying, please send an email to jobs@skysight.inc with a resume/LinkedIn Profile. For extra priority, please include [HN] in the subject line.
Such unsafe habits (like driving without a seatbelt on) statistically will eventually result in a tragic outcome.
I recommend doing some instrument lessons if you haven't already. When I got my instrument rating I questioned whether the private requirements are actually enough. The skills that the instrument rating teaches in terms of preparation, workload management, and emergency weather scenarios make me question whether I was really ready to fly before I had it.
With the private or sport license you'll be fine in the majority of cases. I think my comment comes more from the edge and corner cases that more skills and experience help with, not your ability to work in the system at all.
By the way - I can totally see how a great GA fly-by-wire system is an improvement to maintain positive control of an aircraft at all times. I'd personally love to give it a try and see how it reduces pilot effort while flying.
I'd hope that's the case! That's why I put it in the "good" category".
> how much could be in the air at one time?
Hard to say, but there's a ton of congestion around busy airspace as is. I'd think an order-of-magnitude increase in GA traffic would require a major rework of the whole airspace system.
First - congrats on the launch! I think you're working on an interesting set of components that will prove useful to GA aircraft technology. Bringing fly-by-wire, and lowering the cost of maintenance/manufacturing are both great efforts.
That being said, my personal view is that stick-and-rudder control is one of the less critical components to improving GA safety. Everything else - flight planning, comms, automation, navigation, weather, inspections, procedures, regs, and most importantly - working in the federal airspace system - are the "hard" parts of flying and where problems tend to occur. It's common belief that single-pilot IFR is the most challenging type of flying, because of how much you have to do all at once.
It may sound snobby - but I'm not super excited about the idea of lowering the barrier to entry for GA on a foundational skill basis. Like the light-sport rating, it encourages more people to be in the (already congested) airspace system who haven't really gained all the other skills necessary or experience to be there.
To be clear - I think improving technology and lowering costs = good. Lowering early-skill requirements for pilots and pushing more people without all the other skills into federal airspace = very bad. In general, I'd frame this effort more as an effort to raise the bar for system technology, not lower the bar to become a pilot in the first place.
To power that experience, an app will need to feel "multiplayer" like someone else is working with you. They'll probably bundle this in the API and have "agent mode" that developers can embed in any app or website, or just let consumers give OAI access to control their desktop. It'll also likely work async, so you can assign tasks, walk away or go to sleep for a few hours, and see the results.
This is speculation. But it feels like the interface that we'll look back on in ten years and say "that seemed obvious in hindsight."
If you're a naysayer in the comments, I would encourage you to go give it an honest try, and consider again why you think infra has to be done in harder ways.
I originally learned to code with HTML, CSS, and JS, and I think it's still the easiest way to experience the magic of shipping a working application to others. We should keep encouraging more patterns and tooling that lower the barriers of entry to those just starting out.
My primary concern is the ratio of AI-infrastructure products being built relative to the average utility delta of current end-user products. There are some shining counterexamples to this like Github copilot, but for the most part the utility gain from most text-based generative-AI seems incremental.
Current language models are bad to okay at most tasks - and great at a select few. The capabilities in unstructured ETL and production of labeled datasets fall in the latter category.
I think the next generation of really powerful capabilities to be unlocked with generative-AI will come with 1) an increase in compute availability (training even larger models) 2) advances in multimodal models, and 3) better interfaces to interact with them. Spatial computing will be a really big deal here, and people seem to have forgotten that Apple may have unlocked the next version of interactive computing earlier this year.
If you work on a dynamic version that allows users to understand changes in the open-source topography over time and detect/predict new clusters, this could be a very powerful tool for investment intelligence.