1,315 karma · joined March 19, 2014
Tables are the unit of abstraction.
You’re on the right track with your instinct that you should be able to pull repeated work into a reusable unit. In many DB systems you can just register these sub queries as views and treat them as real tables. In the ETL world that I’m typically working in, you go one step further and just write pipelines to make useful intermediate tables.
Reuse tables, not subqueries.
Sure, but you have no idea how much weight any given sample contributes to the model. Maybe the behavior is included. Maybe not. The weight assigned could even vary drastically between runs with different random seeds!
I probably picked a bad paper with CycleGAN. My main point is that the mapping from data to behavior is non-trivial and there are lots of examples of this in the lit. Medical models that realize that the output variable is encoded in the text in the upper right of the image rather than attempting to analyze the image, etc.
The conclusion is pretty wrong though.
> Regardless of how complicated your program’s behavior is, if you write it as a neural network, the program remains interpretable. To know what your neural network actually does, just read the dataset.
This is not true at all? That’s like saying “all binaries are open source, just read the machine code”. Most datasets contain a reasonable amount of pollution (even the author ran into this). And if you let the model train on its own output it’s pretty easy for it to cheat: https://techcrunch.com/2018/12/31/this-clever-ai-hid-data-fr...
Moreover current ML techniques are pretty sample inefficient. So your dataset is likely to be much, much larger than the equivalent program. Right now we haven’t developed much tooling to help you map a sample bad input or behavior back to training data that might be relevant. So “just reading the data” seems like lots more work than debugging a modern program.
I do think we’ll get better at debugging this stuff in time, but I don’t think it’s currently true that ML systems are simpler or more interpretable than a corresponding imperative program.
1. Talk through the requirements and write down some key bullets
2. Whiteboard some design options. Make sure you have at least 2 interesting ones.
3. Prototype them out. Decide on your favorite.
4. Delete the prototype.
5. Start writing tests now that you know where you want to be, and TDD from there.
The key point is that TDD has to start from a relatively defined sense of what you want. If you don’t have that it’s totally fine explore. Just don’t ship that code.
Are all non-citizens banned from services now? As a tourist in your country, how do I access services?
Like I said, lots of problems. :)
That said, the problem is that sites don’t know for sure who is over and who is under 13. And more robust age proofs have lots of problems: https://www.techdirt.com/2022/06/29/california-legislators-s...
Even if they did, many parents strongly disagree about what content is age-appropriate. When I was growing up my mom was shocked that other 10yo were allowed to view James Bond films.
So while you might be able to set some baselines, parents still need to do a lot of the legwork to make sure kids are ingesting media that’s aligned with their values. This was true on older media and it’s still true today.
There’s lots of very vague language in the bill that we’re going to need to wait for the agency to actually define in regulation. The article authors would rather the law more clearly define things to keep the agency from overreaching.
I suspect they have an inventory problem, not a ranking problem. Why would any real humans publish publicly on the internet at this point?
If you’re posting for fun you have to deal with moderation and stolen content.
If you’re posting for pay Google just takes your content and shoves it onto the SERP, skipping your ads.
Bill Gates had it right. If you want a sustainable ecosystem, you need to make sure the other players are making more money than you in aggregate. As far as I can tell, that’s only true on the internet for eCommerce and so that’s all that’s left.
Considering tech outside of an application context feels weird to me. Like yeah it’s hard to choose with no requirements or constraints.
That said, Rust is probably here to stay. Linux, V8 and some other high profile systems are adding first class support.
https://www.rrstar.com/story/business/2022/04/30/new-pig-but...