My 25-year adventure in AI and ML
austinhenley.com
austinhenley.com
This practical approach to AI feels refreshing in a field drowning in buzzwords. I’ve built tools where simple regression models outperformed neural networks, and convincing teams was an uphill battle. It's hard to not get pushback from teams for not going all-in on AI when it seems decisions and budgets are hype-driven.
- A custom/heuristic driven approach would perform better but will take much longer to build so would be lower ROI.
- There is a strategic angle to using AI here (building competency). We aren't sure what new use cases will open up in the medium term and we need to be fluent with building AI products.
- There is a perceptual/marketing angle to using AI here. We need to convince the market/investors we are on the bleeding edge (hype).
3 is widely mocked but is a completely rational allocation of resources when you need to compete in a market for funding.
This one is funny because my experience has been that ekeing out the issues in this sort of thing is enormously complicated and unreliable and takes an inordinate amount of time. Often the 'bugs' aren't trivially fixable. One we had was the LLM formatting URIs given in the prompt wrongly meaning they're no longer valid. Most of the time it works fine, but sometimes it doesn't, and it's not reproducible easily.
In the long run the latter is of course more valuable and has a larger market, so it's understandable large corps would try to "shoot for the moon" and unlock that value, but for now the former is far far more practical. It's just a more natural way for the tech to get integrated and come to market, in most large corp settings per-head productivity is already a measurable and well understood metric. "Hands off" LLM workflows are totally new and are a much less certain value proposition, there will be some hesitation at adoption until solutions are proven and mature.
- Assume LLMs will be more intelligent and cheaper, and the cost of switching to a new LLM model is non-existent. How does improving the custom/heuristic compare in that future?
Start with ifs, then svms then something else for example.
This has some technical benefits, like speed, and gives you a place to put important hard coded fixes for where a better model makes a small but key mistake. But the bigger benefit imo is getting something to solve the bulk of the problem quicker, and a organisationally it means not saying no to an approach - just where it fits and at what level of improvement it's worth it.
LLMs can have have great application where existing tech can't reach.
Too often, seeing LLMs doing something that's done better already by an existing tech or something it's not designed for seems to miss the impact being sought.
We were lucky enough to grow up with the industry and progressively learn more complexity. The kids out of school today are faced with decades worth of complexity on day one on the job.
Your comment made me curious so I looked at his posts and he has a one about leaving academia because he wasn't happy in 2022, and a more recent one about rejoining it some months ago.
I should say that I'm still in grad school (nearing the end), so the decision hasn't been made yet. The direction I'm thinking is away from academia.
I love the academic environment, access to university resources and close proximity to lots of domain experts. However my experience as of late has been pretty isolating, as my group is almost fully remote despite nearly everyone living in the same town making motivation difficult some times. I also sometimes miss exercising my practical engineering skills, as my current work is entirely analytical/simulation. Overall its been less rewarding than I had hoped.
But IMO it's pointless to hope for something else. AI at its core turns out to be pretty simple. No matter what the best intentioned scientist did, somebody else would think differently.
For example, Stable Diffusion was originally released with a filter that refused to generate porn. There's your scientist thinking of social consequences. But does anyone even still remember that was a thing? Because I'm pretty sure every SD UI in existence at this point has it disabled by default.
We would have to know their internal conversations to know whether that filter was being driven by scientific concern over social consequences or any number of business goals/concerns. We can't assume the reasoning behind it when we only have the end result.
And transformer architecture is too low level to care about things like that, there was no way for the people who made the guts of the modern AI systems to make it so that they only can ever make cute fluffy kittens or give accurate advice.
So what I'm saying is that there's no timeline in which socially conscious scientists would have succeeded in ensuring the the current gen AI landscape with its porn, deepfakes and propaganda didn't come to exist.
This is exactly an argument that supports technological determinism. We simply can't decide -- we have no ability for oversight to stop technology from evolving. That's precisely why I think AI is so dangerous.
Of course not all malevolent actors are dimwits, but there are also many things that even a highly intelligent individual couldn't do on their own, such as a Stuxnet level attack, that AI will eventually (how soon?) be able to facilitate.
I've never heard anyone raise serious concerns over fancier ML and generative algorithms - maybe concerns over job loss but I don't think that's what you had in mind (correct me if I'm wrong).
The more serious concerns I hear are related to actual artificial intelligence, something much smarter than humans acting on a time scale drastically different than humans.
I'm not vouching for those concerns here, but I would say its more fair to keep them in the context of AI rather than ML, LLMs, and generative tools.
This is just the continuation of the same old, just a bit fancier.
Speed is important, strength is important. There is no obvious qualitative difference, but qualitative differences emerge due to a massive increase in complexity, just like consciousness emerges in us but (probably) not in a bacteria due to the massive difference in complexity, even though we are just a scaling of the former.
Your text generation with Perl wouldn't be able to write an article, but ChatGPT can, and the magnitude difference is precisely what we cannot handle, just like I can't be hit by a speeding car at 100km/h and survive but I'd probably walk away from being hit at 2km/h (and once this actually happened to me, without injury). Would you say there's not much difference between the two?