The Facebook Method of Dealing with Complexity
ubiquity.acm.org
ubiquity.acm.org
Mr Gall wrote a book about complex systems and how they fail back in the 70's and most of it (from what I recall it's been a while since I read it) applies as much now as it did then, it's a very good read.
For large organisations it's a bit like getting a look at Oz behind the curtain :)
http://www.amazon.com/s/ref=nb_sb_noss?url=search-alias%3Dap...
I strongly feel that "garbage collection" is grossly neglected in almost every organization and is part of the reason firms slow down as they grow older.
The article mentioned simplification of the Norwegian tax code. Are there governments with garbage collection and simplification of laws built into the legislative process?
I've often been thinking about the same thing. I think that most laws should have an expiration date, especially those that were enacted during "emergency" periods (e.g. everything legislated as a response to the Great Recession). The only reasonable exception are the core laws (e.g. the Penal Code), which should be replaced (not amended) by newer versions every so often.
http://www.commongood.org/pages/the-problem
I'm a supporter.
Sort of, e.g. see:
> Any law that isn't routinely and consistently enforced is invalid.
...since in my view unenforced laws don't do anything except invite corruption.
Prior to each agency's expiration, the Sunset Commission reviews their purpose and effectiveness and then makes a recommendation to the legislature to abolish the organization, continue with modifications, merge the agency under review with another or let it expire.
Agencies subject to sunset review: https://www.sunset.texas.gov/reviews-and-reports/agencies
Since the commission was created ~40 years ago, they've abolished 79 state agencies.
Dealing with complexity by reducing or ommiting some problems are vital also for small and medium projects. This encouranges rapid creation of some working prototype and the overall process tends to shift towards more iterative development.
This is one of the ideas behind building a Minimum Viable Product, and think it's a big strength that startups have. The product doesn't do everything. People will adapt to the product, and start using it in new and unexpected ways. And it's a lot cheaper and easier to build.
I'm curious if there are examples of large companies successfully simplifying existing products. How many features could gmail take away before they alienate too many of their users? Quickbooks, Microsoft office (the ribbon?), or any of the Adobe products?
But it does everything for some people or some use cases. I think that's an important distinction.
Take for example the 'Gorilla' issue that Google encountered.
http://nyti.ms/1JlTBab [www.nytimes.com]
Then there was the issue with the bookmarks and suicide categories
http://bit.ly/1JFkmCr [productforums.google.com]
(This one isn't as big of a problem, as the frustration was probably more due to the compulsory usage of the new bookmarks, but it still shows some hazards of auto-categorization).
With social things, it can sometimes be very hard to solve a problem that works for 95% of the population, but doesn't piss off the remaining 5% (often with good justification). Unfortunately, detecting when you're in that situation in the first place is sometimes equivalent to the problem you're trying to solve.
Use simple statistics to identify trouble spots.
Is this a good use case for a document store, NoSQL database?
I've worked on some large very heavy business logic heavy systems and often what you find is that the database ends up a dumb store and the logic is entirely handled on the application side, this is fine when the original system only spoke to the one application but becomes a massive problem once you want to allow other applications to access that data store, from experience this is one of the reasons why large enterprise re-writes often fail, the specification on paper vs the massive complexity of edge cases in the application code, years sometimes decades of if/else statements.
Of course the inverse is when the database does have all the business logic and the client is dumb while better that does often mean that you can't get the business logic out easily if the database has become obsolete (some of these systems aren't even SQL based or a dialect of SQL that barely looks like SQL, when you have stuff that was written in the 70's all bets are off).
It's things like this why IBM still sells zSeries with COBOL support for code that was written 30-40 years ago.
Not really sure what the solution is, I'd say some kind of business specific almost domain language that goes beyond SQL or Java (but we tried that, it was COBOL ;) )
As for your suggestion about NoSQL there is merit in that but if you go that route you have all the complexity of the business logic in your application as well as much of the complexity that would have been handled for you in the database via SQL (government systems are mostly amenable to normalisation).
Possibly some kind of SQL/Document hybrid (something like Postgres with JSONB) with a business logic layer over the top might work but I've no idea what that would look like (possibly something like SAP's ABAP which I haven't used but people I know who have don't like it very much), most of our current languages care more about pointers, bytes and function than they do about the actual modelling of the system they represent.
You have the same problem if the application implementation language becomes obsolete even if the logic is not in the database. Which your COBOL example illustrates.
That has nothing to do with where you put your business logic, it has to do with (on of the many reasons) why you need to have, and maintain, technology-neutral specifications documentation.
In theory that is absolutely correct, with an unambiguous and current spec you can re-implement the software in something else without the original system the problem there (other than trying to get programmers to keep the documentation in sync with the system...) is that a specifications document written in English (or any natural language) is way too ambiguous to get away with that so then you decide to use a constrained version of English to write your spec and then someone says well why can't we get the machine to understand this and....COBOL again ;).
From what I recall the XML corresponded one to one with the paper forms. If you think about how the paper forms work it's pretty straight forward. You go top to bottom entering data and making calculations. Sometimes there are rules (if this line is less than X, goto Y) that you would have to write somehow.
As to how they store it I can only speculate. I imagine they're stored as blobs by return/form. I doubt they would be stored in some denormalized form as it seems unnecessary and difficult. They don't need to query across all people's returns at once. They're probably batch processed return by return.
[1]. http://www.sec.gov/Archives/edgar/data/1288776/0001288776140...
I think the example of the Norwegian tax system was a good one, because it demonstrates the article's point without any risk of discrimination. But the earlier example of Gender is a terrible one. Gender is more than just whether you have dangly bits hanging off your pelvis, it's also a very critical part of a person's core identity, and there are a lot of different variations. Lumping everything other than cis-gendered male/female into "its complicated" is a great way to make anyone who doesn't fit into the gender binary feel like they don't belong.
Perhaps a more compelling example would be race. I know there's a fairly standard set of racial categories used when asked to self-report race (e.g. in those optional questions you can fill out when doing various standardized tests), which are fairly broad, and I'm pretty sure the last one is "Other". But at least with race, nobody feels like their racial identity is not recognized as legitimate by a large percentage of the world population (which is a serious problem that non-cisgendered or non-gender-binary people have). But even here, imagine if the test just said "White, Black, or It's Complicated". It should be easy to imagine that a lot of people would be pretty outraged over that.
All that said, if you capture the main categories, and then have a free-form "Other" that people can fill in the details, that's similar in spirit to "It's Complicated" but a lot more palatable. Heck, depending on why you're asking the question, you might not even need to save the answer to Other (e.g. if you're only ever looking at data in aggregate, although even then you might want to pull out keywords from Other in order to report more groupings than the form offers), assuming the people filling in the form have no way of knowing whether you're keeping that data or not. So in the backend you might still have the equivalent of "It's Complicated". But how you present that to the user is important.
---
Edited to add: Looking back at this, I realize that it looks like I'm just talking about whether the user feels fairly represented, which is related to but not the same thing as whether the form actually discriminates against minorities. So on that note, I just want to emphasize the importance of picking an appropriate set of main categories, even if you have an Other, such that very few people should ever resort to using the Other, and those that do should not end up getting unfair treatment as a result. For example, if policy is made based on the demographics of a particular set of people, that policy needs to account for anyone who doesn't fit into the standard groupings. If you just have e.g. "Male, Female, It's Complicated", policy is almost certainly going to be made based on the standard gender binary, without accounting for anyone who doesn't fit.
And on a related note to that, it's not even just the person filling out the form who might feel unfairly represented by a poor choice of categories. Anyone else who does fit into those categories is going to see the categories as a reinforcing of their worldview, i.e. anyone who sees "Male or Female" as a set of gender choices will just be reinforcing the incorrect idea that everyone fits into the gender binary, and this helps cause discrimination.
In these cases, I think there are two approaches that satisfy the requirements of simplicity and humanity: think very carefully about whether you need that field and, if not, just remove it; alternatively, make it a free input field. If you think that collecting as much data as possible is a good thing, then the free input field gives you the best possible scenario - the strictly most accurate, detailed data direct from the person who is in the best position to tell you. If you don't need it, then just don't ask.
There's also a distinction here between questions that are objective vs subjective. Many would argue gender is objective, but they'd be wrong, it's subjective. You can't look at a person and tell them what their gender is; it's something only they can decide. So a free-form input field for Gender would be great, because everybody can feel fairly represented. But if you're asking for, say, Age, that's objective, and you can get away with letting the user choose from a set of ranges without any risk of discrimination (assuming you cover the full range of ages of all possible users).
That suggests the root of the problem. Presumably the model is based on identifying a particular feature of the world as worth measuring [e.g. gender]. But boxing gender [e.g. masculine | feminine ] is not a feature of the world and the boxing means that the model does not correspond to the world in regard to gender, even though that was the purpose of capturing gender in the model. The idea of "getting the groupings I want" means my methods are suboptimal scientifically. The objective truths are in the data not in my interpretation.
Not every question needs to become a forum for minorities to express themselves. Not every worksheet or webform should worry about "reinforcing the incorrect idea that everyone fits into the gender binary."
Walking on eggshells...
> Not every question needs to become a forum for minorities to express themselves.
This is perhaps the most problematic part of your comment. If a question allows the majority to express themselves, but denies that same right to minorities, or in fact refuses to acknowledge that the minorities even exist, that's absolutely discrimination. And perhaps even worse, when you have questions like "what gender are you?" that are asked often and where most questioners only accept male / female, that reinforces the incorrect idea that gender is binary and that there are only 2 answers, and effectively tacitly condones other forms of discrimination centered around the same question of personal identity. Which comes right back to your question. The reason you think it's perfectly ok to say "Sometimes, you just want to find out if someone is a girl or a guy" is because you're part of a culture in which the majority of people either refuse to acknowledge that there are more possible answers to that question or think it's perfectly acceptable to pretend that the minorities don't even exist.
Girl or guy
( ) Yes
( ) No