1,787 karma · joined April 7, 2008
I'm also a co-author of the Gosu programming language:
http://gosu-lang.org/
My e-mail is akeefer at gmail
We generally end up putting everything in Jira anyway for historical and statistical purposes, as well as for visibility when people are working remotely, and it simply ends up as someone's job to make sure Jira is in sync with the physical board every so often.
I generally agree with the point about designing your architecture to accommodate whatever gets thrown at you, but doing that is hard. Immensely hard. Nothing is concrete, everything is some extensible meta-problem, where instead of writing code to do X, you have to think "how can I make it so my customers can do X, but can also change it to do Y instead." And of course, they want it to be easy to make it do Y instead. Do you just make them write custom code? Do you enumerate all possible options in a declarative fashion and let them choose between them? It's never easy.
The bits about finding structure in what's unstructured, though, I find to often be exactly the wrong mentality, though it's usually the mindset all engineers start with. The best approach in my experience is usually to just embrace the fact that everything is arbitrary, and recognize that you're just going to be writing a lot of if statements (or giving your customers ways to write if statements); otherwise, you'll inevitably A) run into some situation your nicely-structured algorithms can't handle and B) be endlessly frustrated by the exceptions. The processes being modeled by enterprise software are fundamentally irrational, illogical, inconsistent, and arbitrary. Accepting that, and building systems that can handle that, are hard things for engineers to do, since the tendency is to always try to find patterns and order and simpler, more general algorithms.
Not all problems in enterprise software fall into that category, so it's critical to be able to identify which sort of problem you have: is it one that's structured and makes sense? Then write code accordingly. Is it modeling some pre-existing human process? Then expect that the crazy exceptions you've heard about so far only comprise 1% of all the crazy exceptions out there, that you don't even know how to categorize what those exceptions could be, that there's no real underlying order, and that your job is to write the system in such a way that it can handle the fact that the process is arbitrary.
It lets you do arbitrary hotswapping of code, rather than only swapping method bodies. Not appropriate for production at this point, but you can install it on top of any Java 6 version prior to update 26 (not sure about Java 7); it's pretty useful for doing rapid iterations during development of large-scale server apps or swing applications.
To analyze whether this is "fair" or not, I think the best perspective there is Rawls's framework, which is the maxi-min idea. To summarize, the idea is that inequality is okay to the extent that it improves the minimum position. For example, if you start out with some totalitarian society where everyone makes $25k, it's okay for someone to start making $200k so long as it doesn't mean that someone else now makes less than $25k. Economically, that comes out as something like: it's okay for the top end of the wage spectrum to increase because they're producing more value and capturing that value for themselves. It's not okay for them to increase because they're capturing value that other people are creating instead. If someone making $500k starts making $1MM because they're producing $500k more in value, that's acceptable. If they do that and they're only producing $300k more in value, and the other $200k increase is coming from other people's productivity increases, then that's not "fair" and shouldn't be acceptable.
So, which one is it? Does anyone have a compelling argument? In the financial sector, for example, I think it's pretty easy to argue that it's the latter: people who privatize gains and socialize losses are enriching themselves far out of proportion to the value they create. When it comes to startups and new businesses, it often works the other way, in that the person starting the business, even if they become hugely wealthy, does so by creating even more value than they capture. I'm not sure it's obvious if what's been going on for the last 40 years is one way or another.
Reason #1 is that working long hours often becomes an excuse to not prioritize properly. Working under realistic constraints forces you to really decide whether some feature is worth it, or if spending 40 hours on Feature A is better than spending 40 hours in Features B, C, and D combined. Too often the answer at companies is, "Well, A, B, C, and D are all really important, so just work harder and do everything." That's a very seductive trap to fall into, but it's absolutely the wrong escape valve. At least in my view a failure to focus and prioritize properly is far more often a cause of failure for startups than "we didn't work hard enough."
Reason #2 is that you want to avoid burning people out if you expect to be around for the long haul. Our company just turned 10, and we still have a surprising number of long-tenured engineers, which I'd attribute in large part to the work environment and the relative sanity of the work/life balance people can have. If you expect people to work 60+ hours a week every week, they're not going to stick around for 10 years; they're going to get burned out and bored and they'll feel like the only way to get a break is to quit.
You can quibble with the second reason, but I think that even in a situation where you feel like you have to get a ton done and working 40 hours a week isn't an option, it's very important not to use "we'll just work harder" as an excuse to avoid making the hard decisions around priorities.
http://www.inventionstatistics.com/Patent_Litigation_Costs.h...
Filing patents isn't exactly free either, in terms of time or money.
Also, how exactly do you "beat up" patent trolls? They threaten to sue you, and you . . . threaten to spend millions of dollars on court fees to have their patent invalidated? That doesn't work so well unless you have millions of dollars you don't happen to need.
Are you really suggesting that all those small-time devs sued by lodsys over BS patents would have some recourse if only they'd filed a bunch of patents themselves? Or if only their predecessors had somehow flooded the patent office with enough BS patents that the lodsys ones got thrown out as prior art?
The whole "patent everything, let the courts sort it out" philosophy sounds nice in theory, but the whole problem is that it's freakishly expensive to sort those issues out in court, so patent lawsuits end up as shakedowns: you either pay up in licensing, or you pay up in laywer fees. Either way you're paying someone.
"What is the payment structure between Amazon and me? Amazon pays developers 70% of the sale price of the app or 20% of the list price, whichever is greater."
https://developer.amazon.com/help/faq.html#Sales, Payment and Tax
That's clearly the bit the author was referring to as being deceptive: in public they state that developers will always get at least 20% of the list price (which leads some people to think developers still get paid when their apps are listed as free), but in private they ask developers to take 0% of the list price when they promote the app as the "free app of the day."
Machines aren't patentable because they do stuff: they're historically patentable because new ones take a lot of work to create and they can easily be reverse engineered and copied, so they're patentable because of the pragmatic tradeoff that says that society will be better off if machines are patentable, because it gives people an incentive to create and share knowledge, knowing they won't be stolen. That same justification applies to why pharmaceuticals are patentable: it's a huge amount of work to create a new one, and once they're created they can be copied for a fraction of that amount of work, so without patents people won't research them. Machines are pharmaceuticals are in no way isomorphic in that they do the same sorts of things, but they do share the same sorts of qualities that make patents a net win for society.
Software doesn't share those characteristics: the difficulty of a given software "invention" tends not to be high (except for things like compression or crypto algorithms), similar "inventions" are likely to be arrived at independently, copyright and trade secrets protection work well enough to motivate people to do it, and outright duplication of a program without stealing source code requires a significant amount of work due to the size of any complex program. (i.e. you can try to copy photoshop down to the last behavior, but it's going to take about as much work as writing photoshop took).
Therefore, I don't believe it's to correct that since you can replace a machine with software, and the machine is patentable, therefore the software is patentable. The machine isn't patentable because of what it can do, but rather due to the inherent qualities of mechanical inventions, and those are things that simply don't apply to software.
As to your first question, in a world without software patents, that wouldn't happen exactly as you describe it. What copyright and trade secrets protect you against is outright theft; that's actually a large part of what patents are supposed to protect you against (i.e. you invent something and I just copy it). In physical devices, copying is easier than in software, since the number of elements involved is relatively fewer and because things are easily amenable to disassembly, and there are few "implementation details" that are hidden from an initial set of observations. In software, "copying" something these days generally means re-implementing something that has the same effect, but the implementation techniques could be radically different. As a result, in software as it is now, patents don't prevent theft by "copying" the actual implementation, they effectively prevent re-implementation of the same features, even if that implementation is radically different than the original. (Witness pretty much any software lawsuit in the news in the last six months). So again, copyright and trade secret protections protect you against outright theft of your work: someone stealing your code and re-using it without your permission, or stealing your internal documentation about how things work, or even reading your proprietary source and using it to guide a new implementation. They don't prevent someone from "copying" your software by implementing their own program that does the same thing. If someone does that, and they independently (with no help from you) go ahead and rebuild your system, why should you get to profit from that? If you have a pizza place and another pizza place opens next door and copies your menu, you don't get to sue them for patent infringement: you make sure your pizza is better, or your cost base is lower, and you compete on the merits. That's how pretty much every other business on the planet works: if someone comes out with a similar product, that's life, and it's your job to be better. Imagine how ineffective our markets would be if that weren't the case.
Secondly, software patents are "special" because patents in general are special: they're a constitutionally mandated pragmatic tradeoff that grants people temporary monopoly rights in exchange for the greater public good. (Note that in Europe patent rights accrue from a theory of "natural rights" effectively, but in the US it's 100% pragmatic in base). So if, pragmatically, software patents do more harm than good, they shouldn't be there, end of story. The benefits of patents are supposed to be two-fold: to give people an incentive to create things, and to give them an incentive to disclose their creations without fear of copying. The latter point is more or less totally moot with software: lawyers advise their clients not to research patents for fear of knowingly infringing something, and on top of that the patents themselves are incomprehensible. So that benefit is basically a 0 with software patents, with perhaps a 0.001% exception for significant algorithmic patents around compression or cryptography. The incentive to create benefit is also pretty difficult to justify; lots of small software development shops have exactly 0 patents, outright theft is prevented by copyright and trade secret protections, and these days most companies use patents entirely to avoid being sued themselves or in an anti-competitive fashion. I believe it would be tough to make the argument that less innovation would happen without patents, given the huge number of open source and independent developer projects that are threatened by patents. So software patents are "special" because they fail the pragmatic test: the ROI on them is intensely negative, patents (in the US at least) are only supposed to exist as a way to benefit society, therefor software patents shouldn't exist.
Again, there already are special cases, in that things like book plots or fashion designs aren't patentable; it's up to the legislature and the courts to draw the line on patentability, and they've chosen to say that mathemetic formulas aren't patentable, plot devices aren't patentable (but people try), but that genomes are (which is intensely controversial and the line is fuzzy), as are hardware devices, pharmaceuticals, and now (as of the last 15 years) business methods and software. The line gets drawn and re-drawn all the time. Why not draw it in a way that accrues the most benefit to the public? That's the constitutionally-mandated reason for there being a line at all.
So I don't disagree about general patent reform, but I do disagree that software isn't a special case: it is (along with business method patents) because it's an area where patents are doing the most harm, have almost no benefit to outweigh that harm, and where independent invention is the rule rather than the exception.
Software is protected by copyright and trade secret protections, so even without patents it will always be intellectual property that is strongly protected and which has material value.
So when I hear engineers say they like patents, first of all I assume that they've never worked for a company that's been sued for patent infringement (and that they optimistically assume it only happens to other people), but then I try to find out why they don't think copyright protection is enough. Someone still can't legally steal your code without patent protection, because it'll be protected by copyright and trade secret protections. Even if they didn't copy your code, but they looked at it prior to implementing their own version, that would violate trade secret protections.
In that respect, software is protected the same way that authors and musicians are protected. Authors invent characters, plots, worlds, objects, even words, but they don't get to patent them. And yet they're still protected from theft by copyright protections; you also can't just go and make a movie out of someone else's book without permission, though you can certainly make one that's similar. If it's good enough for authors and musicians, why isn't that good enough for software developers?
So to sum that up: software development involves a creative act that deserves protection, but that's different than saying that the creative act deserves patent protection, which legally enjoins anyone else from independently developing the same thing, and which gives person A the legal right to take away the work that person B has done completely independently (or at least take away any money they've made from it and prevent them from selling it in the future). To justify taking away someone's work like that, you have to either be sure that the work is a copy or derivation of the original, which is almost never the case with software patent lawsuits, or you have to argue that even though it's unfair to deprive people of their work like that, that the benefits of the overall system are positive to society. That's an easier argument to make if 1% of patent lawsuits deprive people of the product of their independent work, but it's a pretty hard argument to make when 99% of them do.
I think that the answer there is probably yes, and I think that if you talk to enough female engineers, they'll tell you that the sexism and insensitivity is a turnoff to them. It obviously doesn't push every woman out of the field, but it's pretty hard to imagine that it doesn't discourage at least some women from pursuing a career in software.
Getting to 50/50 gender equality isn't the goal; the goal is to not have women discouraged from doing a job that they'd want to do because of the (often unconscious) sexism of their potential coworkers. If we did that and the field was still 70% male, then okay, but can anyone say we're honestly at that point yet, and that no one is turned off by this stuff?
Also, I believe you're incorrect about Amazon's patent: it basically does cover any method whereby the user only has to use one click to buy something, regardless of the implementation. It was challenged and then amended to narrow it down to requiring a shopping cart, it appears, but the patent has nothing to do with cookies or databases or anything like that: anyone who implements the same feature in their application could run afoul of the patent, regardless of how they implement it under the hood.
Of course, the issue of if you violated the patent or not, or if you removed any offending source code, is pretty much immaterial to the patent lawsuit issue; they can sue you either way, and if they want to take it to court, you'll have to pay a truckload of money to defend yourself, unless you want a summary judgment issued against you.
My experience with type-safe query layers is that they tend to be incomplete; they simply don't let you generate the full range of SQL queries because you're restricted by the language's type system. That said, I'm not particularly familiar with squeryl (and Scala's type system is certainly more expressive than most statically-typed languages), so I can't say what it's limitations are, I can only make general statements.
Anyway, I think it's fair to say it's difficult to talk about ORM generally due to the differences between frameworks and approaches. So I'll try to phrase things more clearly, and say that I think the author's original intent, and the part I agree with, is the fundamental premise that ORM abstractions are inherently leaky and that performance needs often result in a desire to go around the ORM framework to handle something more natively in SQL. Some ORM frameworks embrace those limitations, and allow you to use them when you want to and to work around them when you don't; other frameworks fight that limitation and attempt to swallow the world such that you never have to leave the ORM framework, and those frameworks tend to be the ones that become frustrating to work with.
So if I were to attempt to charitably read the original post, I'd say that perhaps saying it's an "antipattern" is taking it too far, but saying that it's a fundamentally flawed, leaky abstraction is totally accurate, and that recognizing that it's fundamentally leaky means that you, as a developer, should probably take that into account in your application design and your library selection, and that there are some techniques that might help you to do that.
I don't disagree with the points you've made around caching, but I do think you're simplifying the problem a bit. Not all performance tuning in DB-intensive applications is around caching, and it often involves query tuning, indexing, and traditional DB-level stuff.
A large part of the abstraction leak around ORMs is around both the caching and that DB-level performance tuning. You have to understand what code is going to generate what queries so that, at the very least, you can tune them by adding in the appropriate indexes in the database. All of a sudden, you're living in SQL land, examining query plans, etc. But if you decide that the change you need to make is to the SQL itself, the ORM layer suddenly gets in your way: you either have to bypass the ORM layer to drop into raw SQL, which at worst is hard to do and at best tends to massively reduce the value proposition of the ORM framework, or you have to try to tweak your code to get it to generate the query that you want, which is often frustrating and far more difficult than just writing the SQL yourself. I don't think I'm that much in the minority of having an experience like, "Hmm, the query I really need to write needs to use an ORDER BY statement that includes a few case clauses . . . now how do I convince this query-generation framework to spit that out so that I don't have to pull back all the results and do the sorting in memory?" It's also worth mentioning that caching doesn't help tune writes, so if scaling your product requires scaling writes, you're probably going to be mucking around in SQL land.
There's a similar problem around query-generation layers that attempt to allow you to just write normal methods and have things executed on the database; because the code is so far removed from the SQL, it makes it really, really easy to write really terribly-performing queries or to write things that will do hugely unnecessary amounts of work.
On a more trivial point, the fetching all columns when you only need a subset of them problem is really an issue sometimes, especially if you A) have to join across a bunch of tables, B) the columns that you want could be retrieved from indexes, rather than requiring actual row reads, or C) the columns that you care about are strictly several removes away from the original search table, but the ORM layer loads everything in between. (For example, Foo->Bar->Baz, my WHERE clause is on Foo, but the only columns I care about are the id on Foo, which is in the index, and a few columns on Baz . . . how do I tell my ORM layer to load nothing from else from Foo and nothing at all from Bar? It's a different problem than pre-fetching, because I just don't want anything loaded.)
Now, that's not to say that ORM layers can't be made to perform; of course they can, pretty much all of them have the sorts of hooks you describe, and there's plenty of empirical evidence to that effect. But sometimes the way you make them perform is by just bypassing them.
There's another abstraction point, which is that supporting multiple databases often leads to a least-common-denominator functionality approach; for example, if you want to use a db-specific spatial data type, the ORM has to either provide db-specific functionality, or it might just not support handling that type of data well. The same often comes to things like db-specific functions or query hints; if the ORM layer doesn't handle those things for you, you have to bypass it and drop into raw SQL if you need them.
So really, the argument is not, "ORM's are not functional and no one should use them," it's related to the value proposition of an ORM layer. The value proposition is "This tool will make your life easier, will save you from having to write SQL, and will help you work across multiple databases." If the tool makes life harder than it otherwise would be, then it's not useful, even if it's still possible to do work in it.
So the question is largely around whether or not they make life easier or not. In the simple case, I think the answer is that yes, they do: they make it easier for beginners to get off the ground, they make it easy to do simple queries and writes, and the performance probably doesn't matter anyway.
When things get more complicated, though, the question becomes a lot less clear. Yes, the ORM layer makes it easier to have structured queries that can be cached . . . it also makes it harder to have one-off queries that can be tuned easily based on exactly what data is needed and tweaked to convince the database to generate the right query plan, and it makes it much harder to look at some DB stats, identify a poorly-performing query, and then map it back to the code that generated that query. I know of applications that have basically had to bypass ActiveRecord more and more as they scaled to just do raw SQL queries because making ActiveRecord perform was simply too hard or not possible.
So personally, I prefer an ORM approach that does minimum stuff to let me do the simple things simply (pull rows back and map them to an object, execute simple queries directly on that table), but that's designed from the ground up with the idea that dropping straight into SQL is a normal, accepted part of the workflow, rather than some one-off thing that you should rarely do. But it really depends on your project and your comfort level with SQL.
Maybe I want to find all the functions related to assignment: how do I do that? Do I just look on disk? Do I use tags or metadata? Do I assume some naming convention? It's a problem you have to solve somehow, and solving it requires effort, because it's an explicit organizational effort. Modules or namespaces or packages or classes give you a way to say "this stuff goes together" and to encapsulate and abstract large units of related functionality, and to do so in a way that tools and programs understand. In a Java IDE, my IDE will auto-complete all the functions on a class; in a REPL I can just print out all the methods on an object. That's a useful organizing principle, and it's one that you lose if you totally flatten the namespace.
You could argue that organization isn't one-dimensional or hierarchical, which is what namespaces and such force on you, and it's a reasonable statement. But replacing it with nothing explicit and relying on extensive developer metadata and documentation so you can search for things? That seems even more idealistic. When was the last time you saw every single function, even private ones, documented and tagged so correctly and up-to-date on a real-world, large-scale project that you'd rely on them to serve as the only organizing principle in your application?
Prior to that approach, we tried everything: we actually wrote our own test framework that used javascript to drive the browser from inside the browser, then we moved to using selenium, but they all suffered from that same problem of incredibly fragile tests. Over time, test maintenance is easily 10x the cost of initial test development, so it's a huge, huge problem.
So I see this as a promising direction, but building and maintaining that (essentially) Page Object model by hand is going to result in its own set of maintenance problems. How do you know when the model is up to date? If someone changes the model, how do you know which test that affects if you're in a dynamic language like Ruby? It solves one set of horrific problems, and replaces it with another set of problems that are hopefully, with enough developer discipline applied (and guidance from the framework), ultimately slightly less horrific. Being able to automatically generate the POM via metaprogramming by statically analyzing your application is awesome, but it's not a technique that really works with any standard web framework that I know of.
Web app testing really still has a long, long, long way to go.