Since we are using the power abstraction, why not instead of load shedding, start up secondary machines servicing, which are slower, but could handle a bunch of simultaneous requests for some clients, or even tertiary ones?
That could at least postpone the load shedding, or handle localised surges in service demand due to user behaviour more effectively.