Why did OKCupid write their own web server?
answers.onstartups.com
answers.onstartups.com
I'll be the first to point out the numerous flaws C++ has but it just feels like folly to make fun of the ugly chick at the party without realizing everyone at the party is covered in warts... (The is a metaphor for all programming languages having problems)
I wrote a web server in C from first principles back in 2001* for many of the same reasons that OKCupid seem to have done - and with a team of one good programmer that worked out really well.
* http://web.archive.org/web/20070705191513/http://www.annexia...
Rather than "Microsoft ruined everything", nearly the entire commercial software industry standardized on C++, and entry-level "coderbots" was who they wanted to hire out of school.
C++ lost its sex-appeal because its simply more interesting to do web startup things than to be a junior coder at Big Software Inc.
C++ has problems unique to C++: a deceptive illusion of abstraction (what Spolsky would call "leaky" abstractions) makes a bunch of idioms in the language dangerous, including virtually all "smart" pointers, iterators, exceptions, allocators, and arrays. It is uniquely difficult to write reliable code in C++.
Add to that the superficial but more practical and common complaints against C++: the ghastly compile times, the header file dependency hell that forces every class into a kabuki dance of "pImpls" and nested classes, the error messages rendered in ancient Sumerian... you know I could go on, but you get the idea.
C++ is just not a very good language. C is a fine language. If you are building a system for which C's abstractions are inadequate --- and I'll stipulate that such systems exist --- you're better off with C and a very good glue language (Lisp and Lua are two good, popular choices) than you are with a uniform C++ implementation.
This is the attitude that I really, truly do not get.
If I want a linked list or any other common data structure in C++, I include the appropriate STL header and declare an object with the template syntax. It's two easy lines of code. If I want a data structure in C, I have to write it myself. From scratch. Every time.
C++ is a powerful high-level language. C is a low-level language where everything is explicit; I know this is "elegant" to many programmers, but really, I'm just trying to get things done, and C++ has served me far better in that practical role than C ever could.
I don't love C++; it has tons of flaws (many of which disappear if you use shared_ptr for everything). But I loathe the tedium of working in C.
http://library.gnome.org/devel/glib/2.28/ makes C development much more pleasant. You could make the same argument against C++ if you didn't have STL.
The fact that C is not great for GUI apps does not diminish C's quality for what it is, it's just not designed for that level of abstraction.
C++ on the other hand, attempts to be designed for that level of abstraction, and so in fact you are better off writing most types of modern apps in C++ than C in absolute terms, but when compared against other languages designed for the same paradigm it falls down due to complexity overload. That doesn't mean you can't do great work in C++, but you have to really know what you're doing, which parts to avoid, etc. C++'s strength is being both fast and powerful, but most of the time you don't really need both those qualities. Game programming is a notable exception.
It is just a mess.
If I want a linked list in Python, I do:
frobs = ['foo', 'bar', 'baz'].
The equivalent C++ is: vector<string> frobs;
frobs.push_back("foo");
frobs.push_back("bar");
frobs.push_back("baz");
The reason people like these multi-language solutions is that Python or Lua is much, much more productive than C++ (as in, many times more productive), and yet the most programs have only a very small core that needs to run fast. You can write that small core in heavily-optimized C, making it quite a bit faster than kludged-together C++, and then write everything else in Python. And the total complexity of the system is still much less than if you'd just used C++ to start.http://en.wikipedia.org/wiki/Linked_list#Linked_lists_vs._dy...
In every system I've written, there's been a large chunk of code that runs on startup, or implements a feature that only 1% of users care about, or performs setup for one of these expensive operations but itself only touches a few data items. This usually consumes about 90% of the code but only about 1% of the runtime. Heck, probably 50% of my code never makes it to production at all, because it's exploratory or analysis code that's intended to define the problem, not implement the solution. Why not write it in a language that makes you really productive, and spend the time saved to optimize the hell out of the remaining 10%?
The post I was replying to said "If I want a linked list or any other common data structure in C++". The discussion is about how easy it is to use basic data types, not about what those data types are.
I agree that this is not the main point of the discussion, though.
It is not a terrible language to write in, once you get over the verbosity; it is not, for instance, the waking nightmare that C++ is. But as a deployed system it is the worst attributes of C and of interpreted languages: a complex runtime in which many (probably most) of the errors that can creep into deployable Ruby code can bite you even though your code passed both compilation and static analysis.
Its high-level constructs for the most part share the same reliability issues that C++ has; exceptions, for instance, are not really any safer there than in C++.
Blocks are nice, but I say that having only used them within the past few months. It could just be that the zombies haven't eaten their way to my particular shelter yet. They may, for instance, turn out to be horrible to debug.
Imagine a variant of C or C++ that you have to unit test as vigilantly as Rails code, and you're not far off the mark.
ObjC is a good statically-typed language.
C++ is half a dozen good statically-typed languages, all fighting for supremacy like a ham actor's "multiple-personality-disorder" shtick.
"Statically typed objects have the same internal data structures as objects declared to be of type id. The type doesn't affect the object; it affects only the amount of information given to the compiler about the object and the amount of information available to those reading the source code."
How exactly is Obj-C "good"? Is it safer than C? Perhaps less verbose? Or maybe it's easier to use - libraries aside? Is it better because it runs on more platforms? Because it's standardized?
If all else fails, I think it must be the way that it manages to merge two completely different programming paradigms by just stapling them one onto the other.
std::vector<std::string>> frobs = {"foo","bar","baz"};You can use several glue languages with C++ rather more easily and powerfully than C. See chaiscriptor QTScript. Also luabind and perl and python C++ bindings. Introducing scripting glue languages into the equation only further strengthens C++'s advantages.
> The equivalent C++ is:
The new initializer syntax is widely supported enough now. You can write:
std::vector<std::string> frobs {"foo", "bar", "baz"};
And even if you don't have a newer compiler you can use boost assign. http://www.boost.org/doc/libs/1_46_0/libs/assign/doc/index.h...
It is possible to write significant revenue-producing commercial software using Lua for control logic calling compiled modules, and expect it to perform every bit as fast as a similar app written 100% in a compiled language.
I'm surprised no one has mentioned Lightroom yet.
http://www.sauria.com/blog/2008/10/09/lua-in-lightroom/ http://en.wikipedia.org/wiki/Adobe_Lightroom
I absolutely LOVE not having to reinvent wheels like hash tables. Recently, boost added an interval data structure. I just wrote a few lines to expose those functions as Tcl commands. Whammo, I could prototype all I wanted in Tcl!
I also find it very hard to believe that using two programming languages can be easier than using one. This affects team core skills, build systems, testing, debugging(!)... pretty much everything is the more complicated the more languages you use.
And it's true that the more languages you use, the more complicated things become - but C++ itself is really at least 4 languages glued together (C, object-oriented C++, template metaprogramming C++, and the STL; citation: "Effective C++") and they often don't play well together, so switching down to Python + C, you cut that in half.
It is flat-out not true that you need to write data structures in C each time you use them. The programmers that do this are, in my limited experience, uniformly bad. I wouldn't even write a linked-list for fear of screwing up a corner case on deletion. There are libraries for everything, and they're plenty fast. They're just less pretty looking.
Almost every type of operation in C is predisposed to HUMAN error (the worst kind of error) - I/O, strings, creating a data structure, declaring a variable, calling a function even basic maths can fail in unintuitive ways. [1]
No matter how many glue languages you use, C is not going to become safer or easier to use. Furthermore, C's abstractions are inadequate for most things, even the things that it's traditionally known for,such as the Linux kernel. Take a look and see how many wheels have had to be reinvented for the kernel or any large C project.
But... you probably already know all of this. It's just that you're so comfortable with it that you're willing to forgive all of its sins.
[1] https://www.securecoding.cert.org/confluence/display/seccode...
http://yosefk.com/c++fqa/faq.html
Subjectively I don't feel HN as particularly anti-C++, it's more that people are younger and started programming in languages like Python and Ruby so they don't have a pro-C++ bias. If anything I feel that many people on HN are anti-Java. Personally I'm thankful for Java's existence because it managed to replace C++ in a lot of places.
"Each graft of a new subsystem onto the underlying core has resulted in either rejection or graft vs. host disease with its concomitant proliferation of incapacitating scar tissue. The Unix networking model is a cacophonous Babel of Unreliability that quadrupled the size of Unix’s famed compact kernel. Its window system inherited the cryptic unfriendliness of its character-based interface, while at the same time realized new ways to bring fast computers to a crawl. Its new system administration tools take more time to use than they save. Its mailer makes the U.S. Postal Service look positively stellar."
It is irritating to find false statements being upvoted just because noone bothers to double check them.
“Sure, we love your foreword,” they told me, but “The only truly irksome part is the ‘c’mon, you really love it.’ No. Really. We really do hate it. And don’t give me that ‘you deny it—y’see, that proves it’ stuff.”
Norman remains suspicious, but the preface claims: "We have all experienced much more advanced, usable, and elegant systems than Unix ever was, or ever can be."
Personally, I imagine for many of the participants, it was a labor of hate, not of love. We could chalk up any ambiguous love to Stockholm Syndrome.
I have succumbed to the temptation you offered in your preface: I do write you off as envious malcontents and romantic keepers of memories. The systems you remember so fondly (TOPS-20, ITS, Multics, Lisp Machine, Cedar/Mesa, the Dorado) are not just out to pasture, they are fertilizing it from below.
Your judgments are not keen, they are intoxicated by metaphor. In the Preface you suffer first from heat, lice, and malnourishment, then become prisoners in a Gulag. In Chapter 1 you are in turn infected by a virus, racked by drug addiction, and addled by puffiness of the genome.
...
You claim to seek progress, but you succeed mainly in whining.
Here is my metaphor: your book is a pudding stuffed with apposite observations, many well-conceived. Like excrement, it contains enough undigested nuggets of nutrition to sustain life for some. But it is not a tasty pie: it reeks too much of contempt and of envy.
It takes way to long to learn (probably 2 years for a developer working with it 8 hours a day), and the grown ups don't like it either: http://www.amazon.com/Coders-Work-Reflections-Craft-Programm...
This isn't even an ugly chick.
This is why I read HN.
If you are an expert in the intricacies of C++, please
consider this knowledge a kind of martial art - something
a real master never uses.
http://yosefk.com/c++fqa/why.htmlAnd the answer highlights why you should never retrospectively "judge" design choices several years after the fact.
But, these guys are really fucking brilliant and productive. Immensely... I feel like a chump in comparison.
See the two responses for details.
There are a variety of reasons why startup founders may want to bend the truth with their public statements.
Even though all sites have bugs, broken links, what have you, I don't know any other site that's given me such an expectation that it will be unresponsive for a significant number of page views for any given session over a long term period. Even the sites that started development circa 2003.
With landline connections OKC works very well :-)
I don't think I've ever seen OkCupid down completely, not that I use it nearly as often as reddit. But I've never seen reddit have the behavior that while it's up, I can expect one of every 5-20 page views or posts to time out or fail on the first load/submit but succeed on the second or third. And this is consistent enough with OKCupid that I think about a usage session as something I'll have to be in the mood to fight with.
I know people who have used OKC before. OKC users in my social class (male, white, educated) ignore the match percentages, because the SNR is really low. They just plow through all the search results of people to find good pictures and interesting profiles.
So, I'd speculate that match-percentages are a marketing thing, and that they know they made a weak business decision which required lots of computation and now they're stuck with it.
I'm probably wrong. Maybe the long-tale users pay attention to match-percentage.
We're still together, five years later.
I'm probably an outlier, but hey--match percentage works some of the time!
(I wrote some C++ webapps myself)
A) You still have to worry about it when interfacing with libraries that use plain pointers
B) Shared pointers incur runtime penalties (larger data pushes things off the cache, spurious inc/dec-refs messing said cache. If you don't care about that, why do you use a language like C++ in the first place?
C) Reference counting is a poor form of automatic memory management, you still have to worry about cycles, and use weak references or such to break the cycles.
(Disclaimer: I interviewed @ OKCupid in 2007)
OKWS is less a web server than it is an architecture of servers. It's the difference between sendmail and qmail/postfix.
It has nice security and performance properties because each service is run as a separate user, with a separate process. Logging is handled by an independent daemon. Request demultiplexing is handled by a simple daemon that binds to port 80. Actual HTTP parsing is handled by a shared library that services link to.
[1] http://www.reddit.com/r/programming/comments/exkk3/ok_webser... [2] http://pdos.csail.mit.edu/~max/docs/okws.pdf
One of the biggest problem C++ has is the fact that core language has too many stuff but still lacking things that people really want to use. It's certainly workable, and the results are fast since it is compiled very well. However 'workable' does not mean 'a pleasure to work'.
[1]: http://www.okws.org/doku.php?id=sfslite:tame2:tutorial , http://www.okws.org/doku.php?id=sfslite
Hiring and training is also probably more difficult, though that has got be a huge boon to OKCupid engineers since Match cannot afford to lose them.