Part of what we're doing is building a platform that captures the broader lifecycle of tasks beyond "inference alone" -- things like data import, index building & maintenance, drift detection, corpus query.
2,655 karma · joined March 5, 2009
Before: Founder @ Steamship. Head of ML @ Instabase, YC Alum (S15), PhD @ MIT CSAIL, Google Research
Twitter: @edwardbenson Email: edward.benson@gmail.com Website: http://edwardbenson.com
Part of what we're doing is building a platform that captures the broader lifecycle of tasks beyond "inference alone" -- things like data import, index building & maintenance, drift detection, corpus query.
File.import(url).transcribe('service').tag('service').query('..')
Except instead of manipulating a web page like jQuery, it orchestrates remote NLP workflow over your data. We haven't released our SDK yet, but we're working to make a bunch of awesome reference plugins to let folks mix and match different models out there.
Whisper will definitely be in the mix!
We debated this a lot internally -- specifically whether we should pick a less charged initial dataset to experiment with.
In the end, one of the reasons we decided to run with it was because we felt the controversy times listenership actually made it more needing of computer-assisted search.
There aren't a lot of potentially hazardous situations that could benefit from a deep analysis of Fresh Air, but there are quite a few situations related to Rogan's show that probably could have been engaged with more effectively if there was better access to the underlying data.
The hope is that easier search into "original/source data" will ultimately act as a net-positive societally. E.g.: best way to show the Rogan show was behaving irresponsibly during the pandemic is to make it really easy to get the receipts. But more generally -- so too on either side of any debate.
Totally agree on the "AI and nuance" problem. I think this is going to be perpetual (and good, and necessary) question that needs a lot of attention.
The similar thing this project has got me really wanting is the ability to find snippets of a topic across all the podcast archives I like. Sort of the podcast equivalent of falling into a Wikipedia hole and learning all about a topic from different angles.
Re: other tools -- We're a developer platform, so we're offering tooling from that angle: packages you can drop into a platform and just start using. (In this case, audio search). What's nice about the way it works is you can swap out components: use any transcription engine, any set of models, etc -- and then query across the results.
Some of the transcription-specific API companies (like Assembly) are starting to build in search capabilities, which will also be useful depending on workload and whether you want to add your models or endpoints to the mix.
And yes, the spaghetti code is a huge issue behind the scenes of production NLP deployments. We really hope this can make a huge dent in the problem.
Right now we're aiming for self-service signups in October and then the first SDK launch in November. If you've got particular needs/applications you're interested in, I'd love to hear! Feel free to reach out any time at ted@steamship.com.
We're built atop AWS at the moment, but we're designed in a way that's friendly to cloud-agnostic / private-VPC usage down the road if that becomes critical.
For now, we're focused on public API endpoints in which we shard data across different "workspaces" in the background. Each workspace is a combination of stateful data (relations, vectors, binary), infrastructure (models), and plugins (connections to other services).
When you import & use a Steamship Package, you pass it an identifier that binds it to the particular workspace you want to operate out of.
OP here. We're really excited about what we've been building and eager to get your feedback.
We think developer usability is one of the biggest bottlenecks in AI today. It's easy to make an AI demo, but 100x harder to scale that into production. You end up adopting a mess of spaghetti infrastructure.
We took a lot of inspiration from the way Heroku, Vercel, and Netlify simplified the prototype -> production pathway for web app development, and we're building a platform to do that for language AI.
Today we're launching three packages built on top of Steamship, but soon we'll have the whole SDK up and running for folks to build with.
Our dream is targeted, full-lifecycle AI that you can use with the easy of NPM or PyPi.
As a guy who found the Jargon File via that book, I'm grateful he did it. But I can see how publishing an edit of an online forum would ruffle a lot of feathers: 30 years later you get guys like me attributing it all to him.
I can't recall why it's necessary, but the first step in our developer onboarding README is "Install XCode, even though we don't use it"
FWIW our team has found JetBrains pretty fantastic for Swift. An important caveat of that being that we make a database product that runs on linux -- not a MacOS / iOS app that would benefit from all the special features of XCode. My guess is if that's your jam, XCode is the only high-productivity game in town.
We have written a few shortcut libraries ourselves, but we actually don't end up writing a great deal of string processing code in practice because most of our operations are expressed in a higher level query language that gets compiled down to lower-level operator implementations (which only get written once).
The vast majority of our code (maybe like a lot of systems?) ends up being more about the management of the broader data & processing environment that coordinates everything.
I assume folks are familiar with the downsides of Swift on Linux so I’ll focus on what we like about it:
- Fantastic type system
- Good language extensibility (we use Swift to build ourselves a framework to develop Steamship more easily)
- Good LLVM extensibility
- Mostly performant
- Compiled
All in all, we're very happy with the choice. My biggest gripe is slow build & test times in GitHub actions.
If Swift developers are out there and wanting to dig into more systems-style development, I'm happy to chat.
I used to read them over and over, and it really left an imprint on me. The early hacker ethos was such a strong flavor. Sort of a "one part gnostic, one part mechanic, one part counterculture" vibe.
Sometimes I wonder if that flavor always exists, but shifts from community of practice to community of practice, or if there was something specific about the early days of the net that caused it to arise uniquely.
NLUDB is a well-funded, pre-launch platform for full-stack natural language service hosting. We're motivated by the future Star Trek showed us we were kids: a future where humans and computers fluidly interact. That future requires more than just fancy new models. It requires a stack that makes it easy to build, ship, scale, and weave those models into our data and software products.
Our founding team is a group of veterans from the NLP & developer tooling spaces. We hiring smart, hungry, mission driven generalists above all. But specialization in infrastructure and NLP is a plus.
https://angel.co/company/nludb/jobs
Or drop an email to contact@nludb.com
UltOrg is roughly "spreadsheets re-built atop the RDBMS datamodel". The UI supports nested joins, aggregations, filtering, for both display and data update.
The result is essentially a general purpose app that can display just about any Microsoft Access UI that would been written in the 90s/2000s to provide editable views into relational data.
I'm not sure if it's a domain as universal as a grid of cells, but it's a very cool app and, as a friend of Eirik's, I wish him far more visibility than he's gotten for pushing longer on this particular niche than I think most people would have.
I suspect these methods of communication remain because of their robustness against attack and simplicity to implement. It's a guarantee that even if all the modern networks go down, commands can still be issued from anyone, to anyone using off-the-shelf hardware.
That doesn't mean it's necessarily true -- folks in any country, in aggregate, tend to talk about themselves as harder working than their neighbors. But I think this general idea/complaint is way more prevalent than a few Glassdoor comments, even if that's where they sourced the article.
I don't know anything about semiconductors or manufacturing in general, but I do know a friend at TSMC had to buy a second apartment to sleep in because he works such long hours and didn't want the additional commute back to Taipei at night. It's also worth mentioning that -- like the SF and NYC set who maintain "work" apartments -- this guy is compensated very, very well.
I believe Ant Financial has published an open source one but iirc the English language documentation is sparse.
I've been writing server-side / enterprise code in Swift for about a year now, and the language has really grown on me.
But it's a small community, and IBM's & Google's departures were a big blow to the feeling of inevitable growth to it.
I'd love to find a watering hole to chat with others in the same small niche.
If you’re interested in rolling your own, a good place to start is the sentence-transformers Python package along with a KNN search service like Spotify’s Annoy.
Pinecone, of course, looks awesome as well :-)
The show is called Office Girls. Friend in my Mandarin class recommended it as being fairly simple once you get past some office-related vocabulary. “Rich kid has to pose as poor kid and work his way up from the ground floor of dads business” story.
One of the best parts of LLwN as a way to study is the passive encouragement: you can click on a word to mark it as "known", and then it always shows up green in future subtitles.
Pretty soon, entire multi-line subtitles start showing up all green. And for me at least, that provides a huge confidence boost that helps me keep going.
Instead of seeing each new subtitle as a challenge ("<Deep breath> here we go..."), you think, "It's all green! I already know everything here! I just have to read it!" And that's made a huge difference in how it feels to study.
I wish Kindle had a similar feature for books.
Maybe I just haven't learned proper EventLoopFuture-style development style, but you really seem to get stuck callback hell, similar to the early days of NodeJS. It's a bear on code readability.
Re: the sibling comment about Combine -- I've heard great things about Combine but AFAIK it's still part of Apple's platform-specific layer atop Swift so the small group of us using Swift atop Linux can't use it (yet?).
One perspective is that, “knowledge generation wise,” the current system really does work from a long term perspective. Evolutionary pressure keeps the good work alive while bad work dies. Like that [Top Institution] paper: if nobody else could reproduce it, then the ideas within it die because nobody can extend the work.
But that comes at the heavy short term cost of good researchers getting duped into wasting time and bad researchers seeing incentives in lying. Which will make academia less attractive to the kind of people that ought to be there, dragging down the whole community.
My whole perception of academia and peer review changed that day.
Edit to elaborate: like many of our institutions, peer review is an effective system in many ways but was designed assuming good faith. Reviewers accept the author’s results on faith and largely just check to make sure you didn’t forget any obvious angles to cover and that the import of the work is worth flagging for the whole community to read. Since there’s no actual verification of results, it’s vulnerable to attack by dishonesty.