I've touched on this elsewhere, but here's the crux of the problem:
Racking servers, even if they're cylindrical, is actually much easier than building an entire real-time image rendering pipeline from scratch.
The chassis that we co-designed and had made works out to a few hundred dollars per machine, since we aren't buying in giant quantity. When you talk about servers that run 24/7 and cost between $5k-25k per box, that's not the dominant expense by any measure. It's also been picked up by some other companies to use, which is always nice to see -- we aren't keeping it as a proprietary solution, anyone can buy it from our vendor.
Either way we decided to go, we would still be solving both the image rendering and machine racking / operation problems simultaneously to different degrees. As it sits now, we have about 4-5x the number of infrastructure and imaging engineers as we have datacenter managers / technicians on staff. As it should be, I think.
This design decision is actually part of our advantage, in that it allows us to deliver more functionality to our customers at a lower price per unit on the backend. There's always room for improvement, and we continue to iterate on it -- some day the Macs will probably be gone from the datacenter, but only when it's the right move to make.
I think this is still a penny wise, pound foolish decision.
I am sure it is less work to figure out how to buy and rack some trashcans, vs writing an entire image pipeline for 1.0. I get that.
I also don't know your full specs. There are tons of FOSS image libraries out there. What % of your workflow do they cover? Are you sure you are estimating taking something like Pillow and adding in missing pieces, vs writing an entire image library from scratch? Things like basic filters and image resizes are a solved linux problem for a long long time.
Hardware lock-in always has a price. People that build on AWS exclusive things like lambda do the same thing, to a smaller scale. It seems like you are paying this lock-in price, and will continue to pay it. Perhaps it is the right decision, perhaps not. But I cannot image locking my entire company into the whims of a consumer trendy product line, when there is no fundamental need. I would look really hard at hiring some devs to build out a proper rendering engine...
However, there is no off-the-shelf solution that solves the same core problem as the imgix pipeline. Resizing or format conversion, sure. That plus the types of graphical operations that Photoshop supports? Less common. Now do all of that in one pass in 30-50ms, streaming to/from the GPU, with the ability to easily add new workflows and operations? It's our core business, and we own that workflow completely inside of our stack.
We built on the foundation of CoreImage and that gave us a leg up -- we can take most of what we've built, replace the foundation and keep going on another platform (such as Linux). The main challenge is that there is no equivalent technology to CoreImage on any other platform. PIL (or Pillow) are adjacent technologies, but I can tell you from experience trying to operate a PIL-backed image thumbnailing service at scale it is not a large scale solution. Better than ImageMagick for sure, but it still has limitations. Even something like RenderMan is not tuned for this kind of workflow -- it's awesome at what it does, but latency and scale are not its central concerns.
There are risks to imgix's business, and this is one of them, but I'll be pretty surprised if it sinks us. We have some pretty sharp cookies working on our rendering platform, and while success on that front isn't guaranteed it's entirely within the realm of possibility.
The trashcan Mac Pro has very outdated Xeon E5620 CPUs and it doesn't seem like there's a lot of upgrades in the pipeline. I can even see Apple stopping the sale of these in favor of iMacs.
I think any 'regular' sever with e.g. an Intel Xeon E5-1650 v3 be about twice as fast for basically the same price.
I'd assume that image manipulation is mainly a CPU bound issue.
> Racking servers, even if they're cylindrical, is actually much easier than building an entire real-time image rendering pipeline from scratch.
Scaling a service based on Mac Pros is probably much harder than starting to port pieces of a real-time image rendering pipeline to hardware that is a lot faster and more readily available. To be honest, for the price of the Mac Pros, you can probably run a lot of CPU hours in AWS / ... and skip the whole rack-n-stack.
That being said, I'm not the one running an image rendering business, so I'm just armchair quarterbacking :)
It is annoying that Apple has let the current Mac Pro rot on the vine to a large extent. This is something that certainly gives me pause and will be factored in to the future direction of our rendering pipeline. There may not be much of a future for image rendering on Macs at imgix, but we have to factor in all the variables when deciding on our future roadmap.
The difference between a v2 and v3 CPU definitely isn't double performance, by the way. An E5-2697v2 (the top-of-line Mac Pro CPU) is a 970 on SPECintrate, and an E5-2697v3 is a 1240 on SPECintrate (we track other measurements, but that's the quickest comparison I found). So it's better for sure, but it's also 15W higher draw and has 2 more cores to help grow those numbers. Per-core performance on SPECint is only 60 vs 64.
That said, in a perfect world I would love to use v3 or v4 Xeons on our image rendering hosts. It just wouldn't be a night-and-day kind of difference that's worth throwing our existing pipeline out over. Have to be careful not to toss the baby with the bathwater here.
The real performance issue for us this time around was adding machines to overcome certain scalability constraints, and optimizing certain code paths that were prohibitively expensive.
I think the point here is that your roadmap may lead off-road once Apple finally does to the desktops what they did to the headphone jack on the iPhone(ignore the outcry, do whatever they bloody-well please).