What's wrong with 2006 programming?
antirez.com
antirez.com
While ram is far cheaper than it once was, there still are substantial savings in reducing your resident set requirements from TBs to just hundreds of GBs.
madvise + mincore syscalls could be used to get the kernel to preload the pages asynchronously.
Polling mincore wouldn't be that bad as it would only happen between the commands and at least the core of the algorithm would be simple: before executing a command, check if it's data are in memory with mincore, if not ask to load them with madvise, and go to the next command.
What does Flash use worker threads for?
http://www.usenix.org/event/usenix99/full_papers/pai/pai_htm...
The idea is to process in parallel everything you can and have the services ("processors") communicate asynchronously by events.
I know everyone these days seems to think that removing a structured query language parser from a database makes every other problem go away, but realistically RDBMS vendors spend millions of dollars trying to fix this exact problem. It's called cache invalidation and it's a hard problem to solve in a general way.
SSDs are just a midpoint in the performance trade-off game.
The OS is the worst at this, DBs are somewhat better, but realistically if you want serious performance out of your application you need to make those choices for yourself, and use all strategies where appropriate RAM for records you need instantly (memcache,redis,mongodb) SSDs for the stuff you can't afford to keep on an SSD And hard drives for stuff you can't afford to keep on an SSD.
What you need to think about is the value of your data in dollars per IO/sec per DB ($/(IO/sec/GB), if the amortized value of that data exceeds the amortized cost of the retrieval system then buy it. Focus on increasing the value of your data, not reducing the costs of it's retrieval as that will drop by 1/2 every 18 months anyway. Alternatively, change your business model so you are going short on IO/sec/GB, (eg. pre-sell storage so that when you need to buy it you can do so cheaply)
What I'm trying to say is that the value of a picture is worth more to Flickr than it is to Facebook, thus Flickr will have an easier time building it's retrieval systems than Facebook because of the costs involved. That's why Facebook had to write their own filesystem for retrieving pictures.
I'd bet that any commercial DB will blow rings around redis/mongo/etc if you had your persistent store as a RAM disk and used hard drives for the transaction log. The cost of a SQL Server license is negligible if you're going to buy a server with $200,000 worth of RAM in it. If your data is valuable enough you could just keep everything in SRAM (L1/L2 cache) and buy processors just for the cache.
Ultimately there's no need to use a commercial database either, as there are compelling open source alternatives, though if your needs are very specific, a commercial database may be your best tool.
Yes I believe it was first Cicero that pointed this out. Or perhaps even Aristotle. ;)
I think you meant 'in RAM' there.
I don't think it's accurate to describe any performance critical part of the Linux kernel as "simple." For an overview of the page replacement policy, see http://kerneltrap.org/node/7608. I wondered if CLOCK-Pro [1, 2] had made it into the kernel yet, but it looks like it has not.
This author makes compelling arguments for implementing application level paging. But the nice thing about doing systems work is we never have to rely on arguments alone to evaluate something - show me numbers.
[1] http://linux-mm.org/ClockProApproximation
[2] http://www.cse.ohio-state.edu/~fchen/paper/papers/usenix05.p...
My point here is that a lot of people have spent a lot of time working on the page replacement problem. I am very open to the idea that an application can beat the page replacement policy in the underlying kernel, for a variety of reasons. But: numbers. Always evaluate. The implementation of these algorithms in practice is always more subtle than our high level understanding, so we need to do real performance comparisons to know if we actually improved anything.
(I'm getting my information from the algorithm's author's page: http://www.ece.eng.wayne.edu/~sjiang/ Jiang graduated from William and Mary while I was a young grad student there.)
If you think about the problem space of Redis vs Varnish, it's intuitively obvious that Varnish deals with a wide variety of general data without many opportunities to optimize beyond general purpose algorithms such as an OS provides. Whereas Redis has specific data types often with small footprints, and very careful attention paid to the details of optimization for memory and disk usage.
I'm not trying to deride his work - it's a neat project, and I will probably read through his earlier entries more. I'm down with all of the reasons provided, but I recognize that as humans, we tend to believe in things we understand. Hence, we need to evaluate.
But that's only the first level point. The second level point is: are the optimizations worth it? That is, if you only improve performance by less than 1%, then it's probably not worth the hassle. These are the sorts of things that experiments can tell you.
If it sounds like I'm being pedantic: well, yes, I am. I do systems research. This is the same standard I hold myself and my peers to. If someone asked me to review a systems paper that claimed to improve something, but had no results, I'd reject it. I recognize this is a blog post and not an academic paper, but my standard for "do I accept that this is a better approach" has not changed. And I have seen plenty of blog posts with experiments.
for the whole "programmers shouldn't manage memory" myth. Clearly, if you know what you're doing, you can do better than the OS and/or malloc() does. If you don't know what you're doing, you have bigger problems than writing your own allocator will quickly solve.
That is, "CustoMalloc" takes a look at memory usage patterns of a particular program, then generates a semi-customized allocator for that program.
If I can use an OTS memory allocator and get within 2% of my hand-crafted hand-tuned allocator, I think I'd need to be extremely perf sensitive to care.
And as others already noted, "'just' allocating from VM" can be better (less copying) once you can organize everything right and know how to manage the threads. And the main argument of the main article is "if you can't access it with different threads, VM access can block you everything" (that's the two clients complaint). That much both are right, as long nobody tries to make some too general statements like "just use always X."
Azul recently open sourced some of their kernel patches:
http://www.managedruntime.org/faq
But that is about all I know.