If memory pressure is the issue, how are the trusted testers finding their memory pressure when they have a whole number of in-flight requests? If the PlusFeed fellow got 2.7 working, we'd expect to see 100/3.5%= 28 in-flight requests. Do you have data on how big the base memory vs per-thread memory requirements of these apps are? Python isn't famous for freeing up memory.
Which is to say, do you have any solid numbers that tell us that when we switch to 2.7, you wont have exactly the same memory pressure and either have to up the instance-hour cost, or start charging for ram, or just limit in-flight request to 3 or 4 so that our costs are only 5 times as much instead of 20?
Bottom line: what we all woke up to is that fact that as of right now:
* you set the price of an instance, and
* you get to decide how many instances I'm going to pay you for
Some of us are thinking that while that was a great idea when engineers were in charge, its not such a great idea now the bean-counters have taken over.