Everyone wants to do the model work, not the data work [pdf]
storage.googleapis.com
storage.googleapis.com
So far it's been a crappy industry to be in, it seems the only way to make money is to provide the entire value chain, and this is the business model we're pivoting to.
Just a bunch of photography and lidar scans isn't worth much to anyone, and it's very costly to establish a quality data stream. We had enormous capex on sensor gear, training personnel, running the operations. All this to save a couple bucks on sending an inspector physically to the site.
If you do the entire value chain, you gather the data, you process it into information, and then you turn that information into actionable reports. So that means that besides running our operations, we're also doing the modeling work, and we're providing the industry specific expertise for the actual inspection. So our customers pay for knowing what is broken, where it's broken, and what the impact of that defect is.
In the end, I believe we'll be very successful, but it will be because our investors were willing to take us all the way, where many other companies have simply given up on the concept.
We're also working with terrestrial lidar btw. When we started out we thought we were going to do everything using photogrammetry, but right now we're deriving most of our geometry from the lidar.
Now the challenge is keeping the team from falling back into the easy habit of just selling the data or trying to design product that are nothing more than glorified dashboards.
What I meant was that we had to actually hire or partner with industry inspectors, so we could sell a package deal to the customer, we had to go outside of our comfort zone and sell inspections, when all we really wanted was to sell data.
Is there a mobile video frames dataset for labeling ceilings? Why not? I can't believe I am the only one who experienced this. Why is not a ceiling dataset worthy of, say, CVPR? It will improve mobile video on-device deep learning more than most of CVPR papers. This is a serious problem.
Edit: I understand why niche datasets remain in industry and not academia. Public datasets are better if they are of general interest. But ceiling dataset is of general interest to anyone who wants to process video frames originating from blocked or misdirected camera aka smartphones, and it's hard to imagine topics of more general interest once you have obvious things like face.
Ceilings often have lightings! Lightings (either on and off) vary between geography and culture! (They do! Extreme variations!) Lightings, if on, will cause flares! We just sampled from actual distribution (privacy policy included we can monitor connection for anti-abuse and we can use data for research and development of anti-abuse), but dataset collected with care will be valuable. I am sure actually doing this will reveal other considerations, and even just with points I mentioned, it should be good enough for a paper, if not for PhD.
It's not entirely accurate, but it's good enough. Within a few hours you can have a pretty large dataset of whatever you want, really. (Yay for massive dataset plus user tags.)
You can snag a copy from my server, if you want. (Warning: It's a direct link to a 54GB json file.) https://battle.shawwn.com/sdc/f100m/yfcc100m_dataset.json
I have a script called "janky-image-search" that returns random results from searching this dataset. Here are a few random hits for `janky-image-search ceiling`:
https://www.flickr.com/photos/35981213@N00/9374577304
https://farm3.staticflickr.com/2084/2513539756_7a768de44c.jp...
http://www.flickr.com/photos/96453841@N00/7077625757/
https://www.flickr.com/photos/28742299@N04/4683245369/
etc.
The quality is hit or miss, but it seems better than 70%.
EDIT: Here's a gallery of the first 100 hits for "ceiling": https://cdn.gather.town/storage.googleapis.com/gather-town.a...
But as you can see, it's not effortless. Most of those ceilings are old. So it depends what you want. It's why labeled data is worth billions of dollars (scale.ai et al).
If you ended up using the ceilings uploaded to your app directly from your users, it might be possible that there wasn't any better solution than the one you came up with.
(Small detail: the first 100 hits for ceilings are all _interesting_ ceilings? Nothing like the kind of ceilings usually encountered on a video chat app?)
What is a good tool to search it without 64gb of free memory?
tag="${1}"
shift 1
curl -fsSL https://battle.shawwn.com/sdc/f100m/yfcc100m_dataset.json | jq '{user_tags, machine_tags, description, title, item_download_url, item_url}' -c | egrep -i "\"($tag)\"" "$@"
Basically, it trades bandwidth for memory. And hetzner has free bandwidth.Here's curl -fsSL https://battle.shawwn.com/sdc/f100m/yfcc100m_dataset.json | head -n 1, which should give you all the info about the structure. Or most of it.
{"date_taken": "2013-08-10 11:05:54.0", "item_farm_identifier": 3, "item_url": "http://www.flickr.com/photos/9315487@N04/9497940823/", "user_tags": [], "title": "IMG_7216", "item_extension_original": "jpg", "latitude": null, "user_nickname": "meowelk", "date_uploaded": 1376370010, "accuracy": null, "machine_tags": [], "item_server_identifier": 2856, "description": null, "item_secret_original": "15f4257122", "user_nsid": "9315487@N04", "license_name": "Attribution-NonCommercial-ShareAlike License", "capture_device": "Canon EOS 7D", "item_marker": false, "item_id": 9497940823, "longitude": null, "item_download_url": "http://farm3.staticflickr.com/2856/9497940823_0c0d854111.jpg", "license_url": "http://creativecommons.org/licenses/by-nc-sa/2.0/", "item_secret": "0c0d854111"}
I just dump those to disk and then use `jq` to extract the item_download_url.FWIW, the full ceiling search finished with 9,325 results. Here's the first 1k.
https://cdn.gather.town/storage.googleapis.com/gather-town.a...
We downloaded a subset of 3 million images, which is apparently 379.77 GiB. So a linear extrapolation would be (100/3*379.77) = ~12,659 GiB for the full 100m images.
12TB really isn't too bad. It's massive, yes, but imagenet 21k is 1.2TB.
This long penis is A. This short penis is A. This cat is B. This dog is B. Now, what is this ceiling?
The model, looking at the ceiling, discovers a long fluorescent tube. It is long! Neither cat nor dog is long (The model is yet to discover longcat), and while penis comes in long and short variety, all penises seem long-ish. The ceiling is A.
Adding common inputs to the training (or at least validation and test) sets is a good solution. Its hard data work, but will pay off. There are some techniques outside of closed-set classification that can help reduce the problems, or make the process of improving it more effective:
- Couple the classifier with a out-of-distribution (novelty/anomaly) detector. Samples that score high are considered "Unknown" and can be flagged for review. - Learn a distance metric for "nudity" instead of a classifier, potentially with unsupervised or self-supervised learning (no labels needed). This has higher chance of doing well on novel examples, but it still needs to be validated/monitored. - Use one-class classifier, trained only on positive samples of nudity. This has the disadvantage that novel nudity is very likely to be classified as "not nudity", which could be an issue.
If you are using random images from the web, there is a decent chance that you will get a result like this.
https://www.theguardian.com/commentisfree/2021/feb/03/what-a...
The Guardian labels it sexist. I assumed that it was because the majority of young attractive women on the internet are posting bikini pictures.
They're both true at the same time, which is exactly the point. The people in the article in the paper weren't trying to teach their neural net that men usually wear suits and women usually wear bikinis, but they did. They only noticed because the failure mode was really obvious. In some other situation, the exact same kind of sexist bias might exist in the data -- and important decisions, affecting people's lives, may be made -- without anybody noticing first.
The main question people ask when considering reading a paper is "So what??"
Companies, governments, individuals come to me all the time asking how to "Do AI" without knowing anything about their product, data provenance, testing plans, state observability etc...
Sure, you can do a toy problem with some interesting vision or NLP framework and some off the shelf data that isn't ready for production or anything serious. Who cares? That doesn't do anything novel, it doesn't actually change your product function.
Chomsky said the most important thing to do as a young person is to follow the phrase "Know Thyself."
People who want to do AI need to "Know Thy-data and product architecture" first.
I did a talk last year on exactly this: https://www.youtube.com/watch?v=jK-SBu1iiHo
Now the problem is that any of these folks can get into trouble for doing the others job, and the pay scales are often radically different. Relegating core responsibilities such as data cleaning/collection to support roles.
I'd bet that flattening the title structure in these groups would yield better results, and allow team members to more seamlessly move across Backend, Data, ML, and science oriented tasks.
Typically, understanding seems to stop somewhere between A & B.
Without management-driven separation, too many people try to work on just modeling. You wind up with nobody labeling and cleaning the data. Everyone complains that it's hard to get their models to work on the limited data available, but no one wants to be relegated to the data cleaning role.
I have yet to see a good company layout that makes everyone happy. Too many people want to work on modeling, and the worst part is that individuals are right to fight for it. The more practice you get, the more you improve at the work. Plus, companies tend to value good modelers, so your salary goes up too. It's a very frustrating feedback loop.
Separating modeling work by title or by team is one of the few ways companies can keep from having a run on the modeling projects.
There are a few other ways, but they are all pretty bad too. You can have a "keep what you kill" structure, where everyone gets first dibs on doing modeling on the data they've collected and cleaned. I've also seen a company that makes the modeling job as miserable as possible with slow tooling and high expectations, so lots of people wash out and only a few modelers are left.
I wish I knew a good solution to this. My only idea is to play-up the importance (and salary!) of the support roles, but I'm not satisfied yet.
In the olden times, your model would normally error out if you hadn't done any data cleaning (length 1 factors, constant columns etc) but i suppose with NN's they just ignore stuff like that because of the regularisation, I assume.
Every good DS I know regards the data as much more important than the models, normally from bitter, bitter experience. Models are "sexy" though, so people love them. Additionally you can do the XKCD compiling thing x 10 with models, as they take forever to fit on any non-trivial datasets.
I actually think that the cause of this is poor management. ML/DS//statistics are tools, they do not supply any business value by themselves. If people focus too much on the tools, then their manager should refocus them on the actual problem.
Mind you, if you don't have the background in this area as a manager, you're less likely to be able to spot these pathologies.
Personally, as an experienced DS, I know that my time is best spent understanding the data and the business problem, as that will drive better results. It's sad that people don't realise this (including some of my current management chain, unfortunately).
Data collection/cleaning/labeling etc. and generating good data sets that solve the problem we're facing requires a lot more domain knowledge and experience, and quite frankly I'm a lot better and faster at it compared to a freshly minted Masters/PhD student.
I've seen far more ML projects fail because they couldn't get the right data set, and very few fail because they couldn't find the right model. So I'm getting more and more convinced that getting the data set right should be the primary responsibility of the most senior member(s) of the project team.
Data is not nearly as evil - hate your whole life - as collections, so I hope that we don't have to stoop to this level. It is an option though.
The best approach I've found is to reward based on impact to multiple individuals, and to some extent modulate it by rewarding individuals who "run towards fires". A straight scientist who makes the best datasets in a scalable fashion will thus have high impact on a painful part of the business that no one else will work on.
Personally I might prefer the generalist hat but I see things moving in the opposite direction.
There is also the problem of hiring - if 1-10% of the job requires a lot of statistical theory/research and the rest doesn't, do you hire PhDs and ask them to mostly do the engineering tasks that they aren't necessarily prepared to do? What about the people who are perfectly capable of doing 99% of the job and picking up the extra 1% of theory - do you exclude them because they don't have the academic credentials?
People change faster than orgs/titles, it would be better to keep this part flexible to maximize their growth. This can be further extended to the whole engineering culture.
It's not uncommon that a frontend engineer shares a title with a backend software engineer - these skills are different enough that the individuals are no longer fungible for the most part. Why not give DS's engineering titles and treat it as a management objective to ensure you have the right people on any given project?
If scientist is the preferred title, why not give anyone the science title?
Particularly when you can raise billions by being confidently wrong.
I hear people say all the time how this is what they want, yet for some reason if you aren't already published in machine learning no one wants to take a risk. Instead it's easier to complain about poorly formatted datasets.
Whenever I've tried to actually use existing data to attack the problems I'm most interested in the fundamental problem is that data that is comparable between different labs doesn't exist.
And to avoid any confusion, I'm not saying it doesn't replicate, just that it's not easily comparable, like if you want to know what platinum catalysts will do over an alkene but the data you can find was done in a batch reactor with 22nm platinum and you're using a continuous flow reactor with 10nm platinum. And it's usually worse than that once you include how it's supported, what flow rates were, heating profiles, etc.
It just goes on and on if you don't do it yourself.
Those sound like incredibly useful skills in any business which deals with data (which is most of them). I'm partial to consumer tech (psychology background), but there are opportunities for people with these skill sets.
AI might have been a passable term in the era of rule-based expert systems, but both “AI” and “ML” are particularly misleading terms (when applied to modern data-heavy methods), leading to consequences like that pointed out in the article.
Further, glossing over the importance of data to emphasize the algorithms is not accidental, given the subject’s roots in CS, and the rhetoric used to promote the field. The sub fields where it plays out differently are those which have a healthy respect for the prices of data gathering.
More content, more tools, no time for understanding of the full scope of analysis.
Statistical education has called for the use of real data for a long time and it is still woefully absent in undergrad statistics and probability courses. Students are given clean/cleanish data so they can be evaluated on the correct application of tools, not on a full stack process of data sense making. When I took over our stats course I took out the coding (instructors were arguing over R vs Matlab) and put in excel. Excel opens their eyes because it can’t make math invisible like R can...they can’t obfuscate their mistakes in abstractions.
It’s hard because doing that is at the edge developmentally of young adults...but it’s like giving them a Ferrari as a present for their drivers licenses when we teach this stuff. If experts can’t do it right, think about the implications of teaching them to 19 year olds.
If we can abstract away the models and training, I see an opportunity to focus more of the dataset. If you look at something like Pytorch Lightning (and I think fast.ai but am not familiar) you can deploy very powerful, and proven, models without having to get into the details. What might still be missing are more tools to work with the dataset - easily varying size, composition, using active learning, etc. But it's all supported by more abstraction.
(I'm a huge believer in the importance of understanding things from first principles, so I agree with making sure students develop their skills that way. But when it comes to practical work with datasets, abstraction is very important)
Infrastructure work is low status, janitor-adjacent or like the guy who fixes broken toilets. Academic, abstract, scientific, work is less immediate, more rhetoric-focused, more open ended, similar to lawyers and politicians and is about selling ideas. Shoveling data is like shoveling dirt, while academic abstract modeling is like an architect drawing sketches of a palace.
Like it or not, infrastructure work is thankless. Complaining is useless, either accept it or switch jobs.
I think this is probably only because it's harder to assess skill in it, though. It's easy to test for knowing the vagaries of the ADAM solver or statistics or whatever. What is the equivalent for being a good data cleaner, though?
In this case, their incentives are aligned with good data work, but they still suck at it and neglect it.
I believe there's a few psychological reasons:
- People aren't trained to do it in college, so it's undervalued.
- It requires domain knowledge, which is hard.
- The work is quite gruelling and annoying. It's very detail oriented.
- You don't feel like you're making progress while you're doing data work. Modelling work feels closer to the final output.
- Self-delusion during modelling is easy since you can overfit on the holdout data which gives you those dopamine hits.
- Modelling satisfies curiosity. During data work, the data scientist can't wait to hurry up to the modelling stage and see what the results look like.
- Many people are just bad with data work, they don't have the attention to detail required and produce data with glaring errors.
I actually enjoy data cleaning, sure it's annoying sometimes but the insight you gain actually allows you to build a much better model.
And, to counterpoint one of your points, I would note that part of my (psychology) degree was collecting data from our friends for various surveys and experiments. We'd then add everyone's data together and run analyses on this.
It's really crazy to me that stats/ML/everyone else didn't do any of this, as it certainly lead to me being a much better data scientist.
Especially the making money part, as that's ultimately why my employer/clients pay me.
There's an even more extreme type to the one I outlined. This type of DS dislikes data engineering and dislikes applied modelling.
They prefer to sit in an armchair and think about modelling, writing up a LaTeX document with abstract and overcomplicated math formula.
I met one DS like this who was from a theoretical physics background and his years of operating in this mode couldn't be shaken off.
I got into DS because I really enjoy analysing data and running experiments, and as a backup to academia.
Nevertheless the point is true, and importantly it’s probably caused by “AI engineers” who are actually pretty mediocre and were just attracted to this job for the wrong incentives. If interviews for DS and AI focus on recruiting people for not just theoretical middling knowledge but also the drive to solve “the real problem” that will let them tackle the project in a truly fundamental way, the majority of issues pointed in this paper will probably go away.
> We presented a qualitative study of data practices and challenges among 53 AI practitioners in India, East and West African countries, and the US, working on cutting-edge, high-stakes...
I'd rather read a well-written opinion article and hear more about what those 53 people have to say.
Or are you saying that their qualitative methodology is lacking? Is so, what would be a better way of approaching a topic like this?
You usually try to select for a pretty diverse cohort (which the authors did in this case), not a representative one.
A diverse cohort means you'll probably have all categories that matter. A representative cohort means you'll have a distribution over these categories that is pretty similar to the real world.
The result here is essentially displayed in Figure 1. One could now follow up with a quantitative design to e.g. discern which particular step is the most common obstacle.
Once you've got the model then you need to work on operationalise, and that's where work on your data pipeline becomes important. It's always going to be the case though that building your initial prototype is more fun/interesting and probably 1000x easier, than getting your prototype into production. And it's important to remember that operationalizing your model is probably going to involve working on your model, retraining it and refining it to work with your new production environment.
The authors have tried to dress their waterfall up as "cascades" but the truth is that they've defined the development process as a waterfall and are trying to fix that rather than adopting modern development practices.
That being said, I'd do your simple SQL aggregations first, to ensure that one can actually learn something (if we have 0.0001% positives, a simple approach probably isn't going to work well).
Ok. Well that sure seems like a thing the business might like an estimate of! Sure. Next question: Have you run a basic regression over the data you've got, maybe a simple model would be close enough to throw some error bars up and get going.
Of course not.
First of all I have to convey the idea that we are just engineers, so the vast majority of our job is "boring" engineering stuff. Write, maintain and upgrade code, fix bugs, scale our systems... other than the domain, you wont see much difference.
What people usually see as the "cool" part of the job is often done by scientists. There is an interface between teams where we get the result of science and need to get it "production ready", but that might be around 10% of the job.
"Before March 2020, the country had no shortage of pandemic-preparation plans. Many stressed the importance of data-driven decision making. Yet these plans largely assumed that detailed and reliable data would simply . . . exist. They were less concerned with how those data would actually be made."
[0] https://statmodeling.stat.columbia.edu/2021/03/21/whassup-wi...
I don’t think that’s true in bioinformatics. For example, people highly respect the laboratory work done to collect data on protein function.
There's a type of euphemization that's been irking me in recent years: The vague everyone. No one ever talks about X.
"Data largely determines performance, fairness, robustness, safety, and scalability of AI systems [44, 81]. Paradoxically, for AI researchers and developers, data is often the least incentivized aspect, viewed as ‘operational’ relative to the lionized work of building novel models and algorithms"
Do they mean that projects (at google? in society?) under resource "data work?" Do people get paid less. Do they mean incentives in academia? Do they mean that companies specialising in these areas are less successful? That universities have priorities wrong?
Just for the meander... at the top level, the economy does seem to be valuing data work highly. Adwords & Amazon marketplace antitrust cases suggest that "data work" is at the core of these mega businesses. TSLA's current valuation is partially based on tesla's exclusive dataset. In at least some very substantial examples, data work is very highly valued while models are seen as a temporary advantage... IE it looks like the market (some industrie) expects models/algorithms to be commoditized while the datasets are expected to be valuable intellectual property.
The authors do really have some insights here. A lot of "AI project" post mortems do find "data work" was underestimated. I guess that means it's underappreciated by definition. But... I think it would have gone someplace more concrete if they had started someplace more concrete.
What is valuable about adwords isn't the "data work", right? It's their pricing power in the advertising space...
Someone else might think that the true value is in the data, while AI/ML skills can be bought for 120$ per hour :)
Perhaps one day we'll have tools that will figure out the right AI model for you given data, but we won't have tools that can as easily come up with the data.
I am just pointing out the your statement has an other side, too. It's kinda like with programming languages, people think that particular programming language, rather than domain details, are the transferable skill for SW developer. In data science, it's the statistical modelling theory that is considered to be the transferable skill, rather than the business domain.
But you always need both, and there is no particular reason why would, in hiring, one prefer knowledge of the tooling to knowledge of the domain. It just happens to be, as a kind of spontaneous symmetry breaking.
The kind of researchers who are averse – "to getting their hands dirty" – to do data tasks will have a hard time staying relevant in any production setting. Unless the company has dedicated research department which is a million miles away from product teams, they are going to have a tough time being impactful.
But talking about changing the branding to deliver the product that was advertised all along feels like putting the cart before the horse. Sure, maybe it will work. Or maybe in a few years we'll be talking about how we need more "data sleuths" to replace all those haphazard "data engineers" that just didn't have a job title that emphasized solving the problems rigorously enough.
The more interesting aspects of deep learning are often around engineering the whole pipeline and doing it efficiently at large scale.
Given that it’s (a) more difficult, (b) more business critical, (c) more stressful and requiring an incident alerting on-call rotation, then “data work” should be much better paid and offer job security and career growth.
Yet no company I know of pays expert modelers & researchers less than expert data platform engineers.
So either the companies know something you don’t (e.g. that data platform work is more commodity and easier to replace than rarer modeling talent) or there’s a free lunch you can get by exploiting the arbitrage opportunity to pay data platform experts more and consume correspondingly higher business value that other orgs are missing out on by putting modeler / researcher higher on the status hierarchy than data platform engineer.
My perspective after many years of experience managing machine learning teams (both platform/infra and research/modeling) is that data platforming is just a worse job. It’s unpleasant and stressful and business stakeholders who are removed from backend engineering complexity and just want the report or just want the model couldn’t care less about organizational structures and workflows that support healthier lives for intermediate data platform teams. Because of this, the pay and bonuses for data platform roles should be much higher, but politically speaking it’s impossible to advocate for that, so it becomes a turnover mill where everyone burns out to keep the existing shitty system running, with comparatively low pay and low autonomy, and so nobody ends up wanting to join that team or do that work.