Right now there are a couple of core contributors, but it is a part-time effort for us, so I thought it would be great to solicit contributions from the Hacker News community. I'd love to hear what you think!
726 karma · joined October 19, 2016
Fmr: YC W24, AI researcher at FAIR, Autopilot ML engineer at Tesla FSD
Personal website: https://ben.bolte.cc/
Right now there are a couple of core contributors, but it is a part-time effort for us, so I thought it would be great to solicit contributions from the Hacker News community. I'd love to hear what you think!
Less tritely - I do find it fascinating how problems that would've been independent research endeavors can now be subsumed by large language models. Rather than building a big dataset of protein, calories, etc., just ask ChatGPT
But maybe this time price controls will be different.
Unrelated though, I don't know why Threads' recommender system is so bad. It's like, the minute it exhausts posts from people I follow, it starts aggressively showing me either 1) "Threads vs Twitter" posts 2) some random political hot take or 3) some combination of the first two, any of which is a huge turn-off for me from the entire platform. It's always been a bit of a mystery to me why Facebook seems to be so bad at showing me relevant content when some of the smartest people in Machine Learning work there (in fact, I used to work on machine learning there and it was a mystery to me then as well). Maybe I just have very atypical usage patterns, or maybe I need to turn off my ad blocker, or maybe (my personal pet theory) the quality of ranking systems is hidden from leadership by many layers of people padding their PSCs in various ways...
Citi Bike ridership indicates a lot of pent up demand for missing-middle transit, but the execution is pretty bad compared to on-demand bikes in Seattle (although I haven't been back there in a few years and I've heard it's gotten worse there too, so, grain of salt). The electric Citi Bikes are chronically unavailable, meaning you're stuck with huge clunkers that make you work up a sweat at the first hill. Also it can cost 2-3x as much as the bus or metro for a single trip. It makes way more sense to own your own bike, but that still doesn't fix the fact that half your trips invariably contain some section of "cross this four lane road without a marked bike lane".
[0] https://www.geekwire.com/2017/seattle-area-transit-ridership...
The MTA creates a lot of its own issues though. The bare minimum for a public transit system, in my opinion, should be reliable estimates of how long it will take you to get somewhere, but for non-rush hour trips estimates on the MTA are pretty hit or miss. I've heard this is because the switch operators are all running on some ancient system which is impossible to computerize, but for a system that has tens of billions of dollars in funding I find it hard to buy that they've exhausted every possible option.
I used to be much more sympathetic to these types of articles when I lived in Seattle. Buses in Seattle were generally quite nice and more predictable (Google maps would usually give you a good estimate on your arrival time). Biking was a much nicer experience.
I think cities looking to make better public transit should copy the Seattle model rather than the NYC model.
I mainly write about machine learning, and occasionally random thoughts.
What if I'm not a massive corporation with millions of lines of code to train on and I want to pay for an AI coding assistant? Doesn't this make it effectively illegal for me to purchase such a product for a reasonable price when big companies will presumably be able to use it without paying the tax?
Another situation - let's say you're a company that contributes heavily to open source, but also accepts external contributions. Could Facebook train a model on the React codebase, for example, without having to pay the AI tax?
Another situation - suppose I start an LLM coding assistant and sell it to my friend. Presumably I don't have to pay the tax as a "low revenue" company. Then I get acquired or get some huge seed round and suddenly my customers have to pay the AI tax. Doesn't this just nuke all my customers?
Anyway, as a software engineer, I personally want people to use my code for whatever they want to use it for, without having to pay me for it. I indicate that by using an MIT license. Why throw that precedent out the window?
2. Given that the internet is global, is every country supposed to make their own versions of this? Will I have to pay the EU tax to use models that might have been trained on data that Europeans posted online?
[1] https://www.kaggle.com/datasets/ehallmar/reddit-comment-scor...
Obtained from (check notes) public internet forums
> For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model.
You've got to be incredibly naive if you think public Reddit data isn't used to train ML models, not least by Reddit themselves
Here are some explanations from the Jamaican central bank, in the form of Reggae songs:
https://news.sky.com/video/jamaican-bank-releases-reggae-son...
https://www.npr.org/2019/02/08/692823677/how-jamaica-found-a...