5,703 karma · joined February 11, 2019
Email: ian t butler 0 1 @ gmail.com
I'm more than capable of training a bert classifier in fact in 2019 I had trained many custom berts and was running them on hundreds of millions of documents a day.
I don't want to manage GPUs / CPUs now. I don't want to maintain my corpus and retrain as my product's data distribution shifts. The list of things I don't want to do goes on and on and on. And I'm happy for them to be someone else's problem.
I do just want a reasonably good general classifier served to me with a great devex and calibrated confidence scores to help me figure out when to fallback to another model.
This is very obviously not the type of situation and behavior marginalia_nu and then I were commenting about.
But ARE you doing anything about it? Because unless you're actually taking action it's basically just self flagellation to the benefit of no one. Like if you have some measure of power (community action, organized resistance, political lobbying) and are taking action that's great, reading articles that make you sad and angry and then going to bed isn't doing anything.
I think you have far more leverage than you think in most places still.
I often see this repeated, and it is not true task to task. I work on this daily and we have several tasks where long context is advantageous and our evals against a whole battery of models with different windows show it as being so.
This is why having good evals for the tasks you're working on is so important.
I do grant it's a good rule of thumb.
There are many applications that will benefit from the strength in audio here and until z.ai and co work in visual this could be very strong for general agentic applications, though I see there's a bit of weakness in the benches for areas that might make that less true.
Like all models need to slap it in your harness and do proper evals on the tasks you care about.
Config -> Switch models when a message is flagged -> false
That should stop it entirely when a message is flagged and then you can come back without Opus having potentially made a mess.
METR already redid the study at a later date and now finds a likely 18% speedup
"For the subset of the original developers who participated in the later study, we now estimate a speedup of -18% with a confidence interval between -38% and +9%" (note their use of - and + here could be slightly confusing but they do mean 18% faster per the post)
FAANG doesn't straightforwardly retain the best technical talent, and the median task there is routine.
Lol. I don't know I was talking with a guy yesterday who left a FAANG clearing 1.4 million / yr who's now running his own successful startup. My sample is successful founders backed by top tier VCs or exited founders who have done much better than FAANG. If you can't do more you stick around the FAANG, those who can go and do it.
Fwiw we both agree that LLMs should not design systems. I do the design, but otherwise I don't get how this is true, the success of a design is indicated by long term success in the system it built. You can measure this against success in the task it was deployed for via performance metrics for one. And then from a developer standpoint how easy it was to maintain later on. Success of a system is a measurement over time, but it's not some quality that can only be measured by those who built it.
> Off-topic but having worked in other companies as well, I can guarantee you that this is not the case. The skill of engineers in FAANGs and other "top tier" companies is much higher than average.
I have first hand knowledge of this so I agree to disagree. Being surrounded by google, aws, and meta folks my understanding is the best people leave faang when they get the itch to do something better with their time.
https://x.com/yaroslavvb/status/2067367657272422584 https://x.com/voratiq/status/2067667800643268928 https://arena.ai/leaderboard/agent
Edit: Mis remembered the timeline I saw not 2 weeks, 3 months, but still I think my point stands.
Every time I see an anecdote like this, A it reaffirms my belief that FAANG devs are fairly mediocre on the whole (not saying this is you, obviously there are good FAANG devs) and B it reaffirms my belief that the developers who kind of give up their thinking like this are really using the tool wrong or didn't really care about the work before AI either so its now just a quick means to an end.
The etiquette is to keep to yourself, airpods aren't creating that dynamic. The normal assumption if you don't know the person in most contexts them interacting with you in NYC is they're up to something and that's just normal US city dynamics with > 10 million people on a busy day.
The dynamic changes when you're hanging out in front of your apartment for say a cigarette or something or at a bar or sitting in the park but even then its not wrong to signal you're not open to interaction in those situations its simply more normal to chat with neighbors or other people hanging out.
Everyone is working on personal agents but their identity model is wrong. They act as you, risk your reputation, your data and more. Nym is a personal agent that has (and can make) all of its own accounts and only gets selective read only access to yours.
The goal is to make reliable agents that are able to operate safely in the world to help you do what you want, without exposing your accounts and personal identity to potential harms.
For instance nyms have their own e-mail addresses at nym-mail.com, you can CC them on chains and they can only respond to people on that chain with a lease of 5 days, or permanently for people you specifically add.
Everyone is working on personal agents but their identity model is wrong. They act as you, risk your reputation, your data and more. Nym is a personal agent that has (and can make) all of its own accounts and only gets selective read only access to yours.
The goal is to make reliable agents that are able to operate safely in the world to help you do what you want, without exposing your accounts and personal identity to potential harms.
For instance nyms have their own e-mail addresses at nym-mail.com, you can CC them on chains and they can only respond to people on that chain with a lease of 5 days, or permanently for people you specifically add.
Frankly I think this is faux outrage now that a lot of people are getting wise to the fact a substantial amount of modern programming is basically simple pipefitting.
It was always true that the amount of people providing the foundations the rest of us do work on was super tiny and that was just kind of accepted as fact at least from my > 11 years in industry and overall 20+ years of programming now. It's only now that the pipefitting may be devalued are people getting in a tizzy.
Don't get cute.
Now to answer substantially, no frankly unless I'm legally required to I don't credit things. I usually specifically go for licenses that let me do whatever I want. I don't think I've ever credited a library I didn't have to I just use them and make things with them. That's the point of them and no one would raise your point in a pre LLM world imo.
Edit: Like as the point of absurdity no one is thanking the creators of postgres for every project that happens to use postgres. You still made a thing even if you didn't write your database from scratch.
I don't take issue with their point. Just that we can use a stronger word.
But, I'll take one point in their article a step further you can just say "Humans are invaluable." instead.
I don't like defining humans in terms of valuable at all. Maybe because I feel like that word is very concrete and measured and to actually judge that on any one person requires perspective and capabilities none of us existing or have ever existed possess.
The complexity of the sum total of a human life is so great that I think its folly to try measure the value at all. Those who have tried are often reflected in history as the worst among us.