Math versus Dirty Data
jeremykun.com
jeremykun.com
It’s an infuriating internal battle that lots of others can relate to so makes for pleasurable reading.
For companies to resent it... I think you may have a point. It’s hardly great PR and actually gives you a good reason to discount whatever the marketing says.
I'm really happy to see posts like this here. Jeremy has done lost of great work and so it's somewhat refreshing to see that even someone like him struggles at times at a place like google. It's also important because it show that FAANG isn't all it's cracked up to be and so it is a completely legitimate career path to remain in small startups, maybe getting paid less, but having more fun getting paid.
I've worked at both the big and small and this post definitely resonates with me and certainly makes it easier to go into work knowing that I'm not the crazy one.
I don't think there's just these 2 options. There's the other 90% in between.
I've definitely seen people be extremely disorganised in some aspects of life but also sufficiently, moderately, or fairly, successful in others.
I suspect companies are also capable of exhibiting similar behaviours.
I think dirty data are like the refs in a football game. Nobody comes to see them, but they'll be part of the game until you have perfect players.
It’s not the organizations fault. Discovering and maintaining a known sampling for the data is part of the process. I consider it a win when I can get a team to accept that the data process will need to be debugged just like code.
As I understand it, this problem - "policy-intensive" - is a lot of changes of requests from the customers of the system. In other words, customers don't know fully what they need and produce a stream of requests. This stream or requests may converge on some stable "global" requirements, or, alternatively, may represent "moving target" (not converge).
Additionally, some of those requests are caused by different (supposedly better) understanding of the nature of data - the object of the system. That is, with evolution of the system customers (with the help of developers) understand more and more specific details about the data - missing parts, ambiguous parts, alternative sets of attributes etc.
The approach for such problems, which so far is most promising, is to organize the system as a set of independent, composable as much as possible operations. When another request comes from a customer, or another detail about the data becomes known, the system built from such composable components better allows incremental modifications towards processing such a change.
A good set of operations sometimes develops with time. This approach requires constant reflection on what a particular change mean to the existing process, uncovering assumptions and making them explicit and changeable...
"A mathematician is a machine for turning coffee into theorems."
The problem is usually not "This data was sent as a JSON without any schema and with syntax errors", it's "This avro file has a completely useless schema (e.g. everything is typed as string|null) and there are multiple enumerations where the same value is encoded as 3 different strings (e.g. yes, y, true)"
We are working on all of that