HNHacker News
TopNewBestAskShowJobs

codekansas

726 karma · joined October 19, 2016

AI research @ OpenAI

Fmr: YC W24, AI researcher at FAIR, Autopilot ML engineer at Tesla FSD

Personal website: https://ben.bolte.cc/

submissionscomments
codekansas··on Humanoid Robots Wiki
Hello HN, I recently created an open wiki for humanoid robots. I think there's a good chance that general-purpose robots are the next compute platform, analogous to the personal computer or smartphone for an AI operating system, and I thought it would be cool to have a place to crowdsource information about what is happening in the space and how you can contribute.

Right now there are a couple of core contributors, but it is a part-time effort for us, so I thought it would be great to solicit contributions from the Hacker News community. I'd love to hear what you think!

codekansas··on Show HN: Real-time voice chat with AI, no transcription
Very cool! How is this differentiated from ChatGPT voice?
codekansas··on Humanoid robot startup Figure AI in funding talks with Microsoft, OpenAI
Exciting times for humanoid robots!
codekansas··on Bard Extensions
Written by Bard?
codekansas··on Google's advanced music generation model and two new AI experiments
It's really interesting for me to see AI voice cloning taking off. A few years ago I worked on a paper called HuBERT (https://arxiv.org/abs/2106.07447) which has since really taken off for doing this kind of stuff (it's the speech representation that SoViTs uses, for example). At the time the main research focus was low-resource ASR; it was sort of a fluke that it works so well for voice conversion. It makes me feel confident that open-source ML will drive more value over the next five years than closed-source ML, despite all the hype around big tech, simply because random people building weird, unexpected stuff on top of foundation models will make things that no one could have predicted.
codekansas··on GitHub 404 errors
Time for a break :)
codekansas··on Food and Generative AI
XKCD for this

https://xkcd.com/720/

Less tritely - I do find it fascinating how problems that would've been independent research endeavors can now be subsumed by large language models. Rather than building a big dataset of protein, calories, etc., just ask ChatGPT

codekansas··on U.S. GDP grew at a 4.9% annual pace in the third quarter, better than expected
As a US taxpayer, I would love if Europe started carrying it's own weight militarily rather than having our taxes pay for European security. The Ukraine conflict at least seems to have woken some Europeans up to the fact that they need to contribute equally to NATO instead of just relying on the US. But there doesn't seem to be much of an appetite in Europe for trading social spending for defense spending. It's more fun to dunk on American healthcare and pretend that Putin just really wants to be buddies.
codekansas··on The End of Cheap Europe Flights? France Proposes EU-Wide Minimum Price
Well prices for those customers were relatively inelastic (think, business flights being comped by the employer, airlines having quasi-monopolies on routes) so the airlines knew they could raise prices without impacting demand. Conversely lowering prices likely wouldn't impact demand either. The sensible thing to do under normal market conditions would be to fly fewer flights, but when a central planner is setting your prices it changes the logic.
codekansas··on The End of Cheap Europe Flights? France Proposes EU-Wide Minimum Price
The US used to have a fixed prices for flights, set by the Civil Aeronautics Board. There ended up being a lot of gimmickry with how airlines competed with each other - for example, airlines would do stuff like fly more flights than they needed to, so they could point to all the empty seats and ask the CAB to raise prices for their route to offset the money they were losing.

But maybe this time price controls will be different.

codekansas··on Europa the Tech Third World
It is a tightly integrated trading block though. Mississippi and New York are very different but it is much easier to move people and money between those two places than between New York and London.
codekansas··on Meta’s Threads App to Launch Web Version as Rivalry with X Enters New Stage
This is great - I have a suspicion that large chunks of high-quality computer science Twitter content is posted while people are in the middle of coding, or doing some other project. Mobile-only was a small to medium barrier to posting good quality stuff (but a low barrier to posting bad quality stuff).

Unrelated though, I don't know why Threads' recommender system is so bad. It's like, the minute it exhausts posts from people I follow, it starts aggressively showing me either 1) "Threads vs Twitter" posts 2) some random political hot take or 3) some combination of the first two, any of which is a huge turn-off for me from the entire platform. It's always been a bit of a mystery to me why Facebook seems to be so bad at showing me relevant content when some of the smartest people in Machine Learning work there (in fact, I used to work on machine learning there and it was a mystery to me then as well). Maybe I just have very atypical usage patterns, or maybe I need to turn off my ad blocker, or maybe (my personal pet theory) the quality of ranking systems is hidden from leadership by many layers of people padding their PSCs in various ways...

codekansas··on Maybe the problem is that Harvard exists
From my experience the best people are usually state school undergrad (e.g., Berkeley / UIUC / UNC / UW / GTech) plus selective grad school (Harvard / Stanford / MIT / CMU). At the graduate level those schools are much more meritocratic and it's usually a better indicator of ability. Plus doing well at one of those undergrads is actually really hard.
codekansas··on Driving is more expensive than you think (2020)
The "schedule at the mercy" issue has more to do with really bad estimates of commute times. The bus rapid transit in Seattle has good estimates so even though it might be slower it is at least predictable. My understanding is that MTA's ancient switching system can't be effectively computerized, hence the poor estimates.

Citi Bike ridership indicates a lot of pent up demand for missing-middle transit, but the execution is pretty bad compared to on-demand bikes in Seattle (although I haven't been back there in a few years and I've heard it's gotten worse there too, so, grain of salt). The electric Citi Bikes are chronically unavailable, meaning you're stuck with huge clunkers that make you work up a sweat at the first hill. Also it can cost 2-3x as much as the bus or metro for a single trip. It makes way more sense to own your own bike, but that still doesn't fix the fact that half your trips invariably contain some section of "cross this four lane road without a marked bike lane".

codekansas··on Driving is more expensive than you think (2020)
Seattle's growth in public transit ridership is the best in the nation (or at least was up to 2017) [0]. I remember one year Seattle and Minneapolis were the only two cities in the nation to have a drop in car commuters. NYC has way more density but it's a decaying system. NYC could start backing bus rapid transit to make it time-competitive with the subway and it would probably be both a better experience and save the city a lot of money.

[0] https://www.geekwire.com/2017/seattle-area-transit-ridership...

codekansas··on Driving is more expensive than you think (2020)
Well the article includes things like lost productivity due to sitting in traffic, so I think it's fair to point out that there's lost productivity due to train delays.

The MTA creates a lot of its own issues though. The bare minimum for a public transit system, in my opinion, should be reliable estimates of how long it will take you to get somewhere, but for non-rush hour trips estimates on the MTA are pretty hit or miss. I've heard this is because the switch operators are all running on some ancient system which is impossible to computerize, but for a system that has tens of billions of dollars in funding I find it hard to buy that they've exhausted every possible option.

codekansas··on Driving is more expensive than you think (2020)
MTA's issues aren't really funding issues, it's just not done very well in comparison to other large cities. For example, for legacy reasons MTA cars can't be automated like the cars in subways in most other places that have them, so you end up with a lot more humans being in the loop, which results in more delays due to human fallibility. It's a very old system that has been largely reluctant to upgrade and is falling behind the rest of the world in many ways (admittedly no where else in the US though).
codekansas··on Driving is more expensive than you think (2020)
I moved to NYC a year ago and I really don't like using the MTA. It's dirty, chronically delayed, and randomly alternates between swelteringly hot and freezing cold. Getting anywhere here takes half an hour. Outside of a few choice streets, bike infrastructure is abysmal. Riding the train feels like you're schedule is at the mercy of the whims of some random MTA employee or a weirdo who pull the emergency brake and it makes me wish I had a car again.

I used to be much more sympathetic to these types of articles when I lived in Seattle. Buses in Seattle were generally quite nice and more predictable (Google maps would usually give you a good estimate on your arrival time). Biking was a much nicer experience.

I think cities looking to make better public transit should copy the Seattle model rather than the NYC model.

codekansas··on I am drowning in mutes: The current Threads experience
It's very possible that ranking will be a lot better for threads though, simply because NLP models are currently so much better than what exists for video
codekansas··on Ask HN: Could you share your personal blog here?
https://ben.bolte.cc/

I mainly write about machine learning, and occasionally random thoughts.

codekansas··on Microsoft, OpenAI sued for ChatGPT 'privacy violations'
Also regarding international policy - good luck getting Chinese citizens to pay the US AI tax. Effectively you'd be nerfing anyone under US jurisdiction
codekansas··on Microsoft, OpenAI sued for ChatGPT 'privacy violations'
I don't think a model trained on a single company's data would be nearly as helpful as a model trained on all publicly licensed code on the internet. But suppose it were...

What if I'm not a massive corporation with millions of lines of code to train on and I want to pay for an AI coding assistant? Doesn't this make it effectively illegal for me to purchase such a product for a reasonable price when big companies will presumably be able to use it without paying the tax?

Another situation - let's say you're a company that contributes heavily to open source, but also accepts external contributions. Could Facebook train a model on the React codebase, for example, without having to pay the AI tax?

Another situation - suppose I start an LLM coding assistant and sell it to my friend. Presumably I don't have to pay the tax as a "low revenue" company. Then I get acquired or get some huge seed round and suddenly my customers have to pay the AI tax. Doesn't this just nuke all my customers?

Anyway, as a software engineer, I personally want people to use my code for whatever they want to use it for, without having to pay me for it. I indicate that by using an MIT license. Why throw that precedent out the window?

codekansas··on Microsoft, OpenAI sued for ChatGPT 'privacy violations'
1. How would this not make tools like Github Copilot exorbitantly expensive? Why should I have to pay a tax to everyone else in the United States to use something that was disproportionately trained on my own data?

2. Given that the internet is global, is every country supposed to make their own versions of this? Will I have to pay the EU tax to use models that might have been trained on data that Europeans posted online?

codekansas··on Microsoft, OpenAI sued for ChatGPT 'privacy violations'
People have been training ML models on data scraped from Reddit since at least 2015 [1], back when there were less than a million users

[1] https://www.kaggle.com/datasets/ehallmar/reddit-comment-scor...

codekansas··on Microsoft, OpenAI sued for ChatGPT 'privacy violations'
> including personal information obtained without consent

Obtained from (check notes) public internet forums

> For the 16 plaintiffs, the complaint indicates that they used ChatGPT, as well as other internet services like Reddit, and expected that their digital interactions would not be incorporated into an AI model.

You've got to be incredibly naive if you think public Reddit data isn't used to train ML models, not least by Reddit themselves

codekansas··on Governments should compete for residents, not businesses
Not OP, but one reason could be that if a lot of early infrastructure was debt financed, cities will have an increasingly harder time paying for infrastructure maintenance over time
codekansas··on “Forbidden-word” lists are a symptom of administrative bloat
By n-word do you mean "no-can-do" or "normal person"? Both on the Stanford Elimination of Harmful Language list
codekansas··on AI Won't Cause Unemployment
It is super hypocritical, but also, it's basically correct. I'd rather live in a world where random people showing up to Atherton city council meetings can block developments on property they don't even own (fortunately we seem to be moving that direction). This is very much an example of regulatory capture by vested interests
codekansas··on U.S. inflation stays high as housing costs bite
Some level of inflation is good because it encourages investment, too high or too unpredictable is distortionary.

Here are some explanations from the Jamaican central bank, in the form of Reggae songs:

https://news.sky.com/video/jamaican-bank-releases-reggae-son...

https://www.npr.org/2019/02/08/692823677/how-jamaica-found-a...

https://www.youtube.com/watch?v=ieRq4P1-yqY

codekansas··on ChatGPT get-rich-quick schemes are coming for magazines, Amazon, and YouTube
I feel like this was lost in the conversation about Meta / Twitter moving towards verified accounts. I think social media companies are starting to realize that they should prepare for a world where most online agents are AI or AI-assisted
← PreviousPage 3 of 7Next →