Coding tricks I learnt at Apple
blog.joemoreno.com
blog.joemoreno.com
This is pretty terrible from a security standpoint. A development environment is typically much less secure than the live environment, and for good reason. The development environment must be accessible to developers, typically both on-site and remote. All developers have access to test databases for the purpose of testing their changes. There are often many more software packages in a development environment, and development servers have a higher probability of running vulnerable services. Live environments typically have much better logging and auditing.
Every company should have a program that can be run to sanitize the live database for use in testing. I've seen too many situations where the production environment was appropriately locked down and audited, but the development environment was compromised. It's not unheard-of for a developer to lose possession of his laptop, and if it contains a copy of the live database it's no better than the site, itself, being compromised.
Multiple databases made up the live environment. The developers did not have access to the live environment. As a matter of fact, we had developed it to the point that, except in special circumstances, the engineers didn't even know what products were about to be announced until Steve Jobs announced them to the public.
The dev database was a copy of the products database – most everything in the products database was either public information or it was obsolete data for EOL Apple and third party products. I don't recall ever seeing any sensitive data in our production environment. Unlike, say Walmart or Amazon, we didn't need to keep track of inventory since we build the products (the only exception that comes to mind were the refurbed products since there was a limited amount of inventory).
Dev systems would process dummy credit card numbers, etc, in dev (but I didn't work, in that area). The applications were smart enough to figure out which environment they were in when started up which is a big help.
+1.
However, modifying or sometimes even a dump/load on a SQL database can change its behavior dramatically.
But the harder part can be getting a hardware setup similar to production. Most places want their newest, best hardware in production for obvious reasons. Few places want to pay for mirror-image production systems just for testing to bang on.
There's the old saying: If you have two systems, 'production' and 'testing', the one with the more challenging load is 'testing'.
Also, I don't recall ever seeing any security issues with our environment or how we handled the code and data. There were some very smart engineers "minding the store."
Read above line twice. He is saying that test is done in 'production environment'. Do not assume there is 'development environment' for testing within 'production environment'
And mostly, I made my comment in the hopes that people here wouldn't go, "Oh, Apple tests against a copy of the live database. I should do that, too!" I'm sure Apple is smarter about it than my comment indicated.
My point was simply to be careful with the live database. Investing in state-of-the-art protection and auditing in your live environment can be trivially defeated by simply copying the live database to an insecure environment. So don't do it.
Also, some organizations, especially ones that use cloud computing (EC2, Slicehost, etc) have the option of creating a staging environment for QA. testing, etc and, if everything checks out, the staging environment can be restarted as a live environment (there are ways to do this safely and securely) while leaving the previous live environment live. If, for some reason, you find problems with the new live environment, then you can simply shut it down and keep the old live environment humming along. Of course, this is easier said than done, but it'll keep the anxiety level down.
Another good idea is to have apps that know which environment they're in when they start up. So, if an app starts up in dev, it should only access the dev database. If it starts up in the live deployment environment, it should only access the live database. To enforce this security, the live database must only allow connections from clients in the live environment (IP address).
But, these "self aware" apps don't necessarily need (probably should not) have configuration files with the sensitive data used to access databases, credit card processors, etc. Rather, at app start up, these apps should know which environment they're in and then go over the network to fetch the sensitive data used for logging into databases, services, etc. (The apps should only store this sensitive data in memory.) The services, on the other end of that connection, should be checking the IP address of the app which is asking for the data to ensure that there isn't a mixup.
B-trees are an external storage technique; he means balanced binary trees.
The answer to the interview question is another question: "what operations does the container need to support?". It's not "hash tables are O(1)".
A perfect hash table does have O(1) look-ups in the worst case.
There are lots of other constraints that come to mind. Do we do frequent insertions/deletions?
What are the memory constraints?
What is the size of the unhashed key?
And many more... No data structure always wins.
Radix tries on the other hand don't hash. They're always at least O(k), and you can find the next and previous values (lexicographically) in O(k) as well. They're also more compact and you never have to resize your tables.
That's important to keep in mind, particularly for DoS resistance. But isn't it really as much of an oversimplification as saying they're O(1)?
In practice, there are inexpensive ways to consistently avoid that worst case. But you do have to know to do it.
Hmm reminds me of some code I should probably go double check...
No, because big-oh notation implies worst case. It's an asymptotic upper-bound. If you want to talk about average case, then you need to qualify what your assumptions are. So insertion into a hash-table is O(n), but if we assume an even distribution of keys with a sufficiently large table, then insertion is O(1).
This is kind of a silly semantics argument... but, if you interview for a job and look like a deer in the headlights when I said hash tables aren't O(1), NO HIRE.
(I'll assume you're talking about the quicksort algorithm rather than the qsort() function because you're comparing it to hash table accesses in general rather than a specific implementation of them.)
But I'm not sure I agree. The worst case for quicksort is an already-sorted list. I run into this scenario all the time, though usually in a slightly different form. When I iterate through the elements of one of the tree-based containers std::set and std::map they emerge in sorted order. If, inside my loop, I then insert them into a similar container I end up with worst-case performance. The other day I replaced std::set with std::unordered_set and saw a dramatic increase in performance, although it may not have been entirely due to this effect.
On the other hand, a non-crappy hash function should be available to everyone at this point. Personally I like FNV-1a, but haven't found an issue with whatever my C++ standard library is supplying. I have spoken with other programmers though who didn't realize hash tables needed some empty space to perform well or had stumbled into a pathological case with their oversimplified hash function.
This is kind of a silly semantics argument... but, if you interview for a job and look like a deer in the headlights when I said hash tables aren't O(1), NO HIRE.
Gee Ptacek, what do you have against deer? I can just imagine some poor interviewee with a bright desk lamp shining into his eyes...
While there may be a set of textbooks and papers for which the editors would have flagged "worst cast O(..)" as redundant, I think it's relevant to point out that the use of big-oh notation in mathematics predates computer science. So to the extent algorithm analysis is a branch of mathematics, those who use big-oh in this more general way are in fact consistent with the larger body of work.
I find this a bit confusing because it's unclear to me what Knuth actually endorses for use in average case description.
Wikipedia: Although developed as a part of pure mathematics, this notation is now frequently also used in the analysis of algorithms to describe an algorithm's usage of computational resources: the worst case or average case running time or memory usage of an algorithm is often expressed as a function of the length of its input using big O notation. Wikipedia and the web in general have many examples of "worst case O(...)" which would be redundant under your convention.
So clearly one needs to be vigilant about assumptions when discussing average case, but common usage does not agree that big-oh notation implies worst case.
[1] Every database implementation techniques lecture compares the two. See, e.g., http://infolab.stanford.edu/~hyunjung/cs346/.
There is a tradeoff of computation (doing binary search inside the node) and amount of storage (keeping less pointers). In environments where memory access is much more expensive than a CPU instruction, it is preferable to perform these computations than to have to read all the extra pointer data to jump to the right places.
In fact, a breed of cache-friendly data structures are precisely based on the B+Trees but with even less pointers, having the algorithm compute these pointers instead (CSS-, CSB-Trees)
For some reason, binary trees look real cool when you learn about them from a textbook that people sometimes forget about hash tables. But, that might just be my personal experience from candidate interviews.
I am trying to figure out who would find this article useful without any details.
Ahh, the all powerful Source Control Shingle: http://thedailywtf.com/Articles/The-Source-Control-Shingle.a...
Requiring that work on it be done single-threaded like this suggests that some other part of the overall process broke down somewhere - the developers/automated tests/continuous integration server couldn't catch merge conflicts? Code reviews weren't done and made visible to everyone else on changes to this special code?
By putting a mutex on it, you are forced to re-evaluate the state when you obtain the lock. You can't sneak in a change in front of somebody else's change, requiring them to re-evaluate (or, worse, merge and hope).
If it was required for modules all over, it would be broken process. For a single module, I can understand it.
1. Sometimes the best process is old fashioned communication between people with common sense.
2. Never underestimate the power of a rubber chicken.
I most wholeheartedly agree. Git, or a Source Control Shingle cannot replace effective communication. In fact, a solving a merge conflict is much more painful than a quick discussion.
When in doubt, STFU. Not just for legal reasons, but also because you don't want future collaborators and employers thinking you're a Chatty Cathy who's going to tell everyone about your secret sauce.
I don't think it's fair to expect Steve Jobs to hop in his helicopter and track down Joe Moreno so that he can high-five the guy for writing something that got on the front page of Hacker News.
Also, the "because you don't want future collaborators and employers thinking you're a Chatty Cathy who's going to tell everyone about your secret sauce" bit rubbed me the wrong way. I mean... really? I feel like that rates high on the unwarranted paranoia scale.
All around us we see civilians cuffed and murdered for their opinions (for example, in countries with dictatorial regimes), but apparently this is not the 'status quo' in the USA.
I think Guy Kawasaki's first book, The Macintosh Way, goes much further than my comments here – and he was even rehired back at Apple several years after his book was published.
My hope is two fold. 1. That someone reading this would pick up some good practices; and 2. That someone reading this would say, "I want to be a part of that!"
Apple's a great company to work at on so many different levels.
(a) How do you find out, before going to work somewhere, whether they actually work like this? Are there questions you can ask? Word of mouth? ... ?
(b) If you don't work somewhere like this, how do you start putting professional processes in place? Assuming in particular that you have never actually worked somewhere like this, so you can't speak from experience, only from instinct about what seems to be a good way of working.
Why? Also, is this common these days?
??
Does Apple seriously turn off their store while Jobs talks? Or is he talking about pushing new content out based on announcements?
The former just sounds... Odd.
You wouldn't want an uninformed person buying an iPhone 3GS after Steve has already introduced the iPhone 4, but you don't want to put it up for sale yet because the media should be paying attention to Steve, not the store. If they just launched the item with all the specs on the store as soon as it was introduced in the keynote or media event or whatever, then it would ruin the suspense. Jobs is a salesman.
I would think an appropriate separation of content from the site itself would allow them to reach their same business goals without deliberately giving themselves an outage.
Wait, saying a B-tree is the common answer?
http://www.geeksofpune.in/drupal/files/8058778-ext34talk.pdf
ext2,3 used an indirect-block tree structure. Ext4 uses an extents system which is nicer, but still not (AFAIK, I've only skimmed this part) a B tree.
I'd wager "learning how to properly use 'learned'" wasn't one of them.
http://www.urch.com/forums/english/9214-learned-vs-learnt.ht... "The descriptive answer in American English is: There is no such word as "learnt". Use "learned" always."
In the US, we consider that poor grammar, as in: "I done did kilt three of 'em skeeters on me yesterday."
OK, maybe yours is a noun, but still!