These tests are intended to be stand-ins for realistic application behaviors. They are not realistic applications, but they should make use of the same functionality that a real application would. I want the tests to establish realistic high-water marks for the types of operations they exercise.
Running a single query per request? The high-water mark you might expect is what you see as our single-query test. It would be very difficult to make a single-query operation in a real application more trivial than what we have.
If I treated all tests as a black box, none of the database tests would actually query the database; the fortunes test would not add an ephemeral object and re-sort on every request; and so on. Everything would just send static responses back.
Bottom line: nothing beats benchmarking your own application. But the tests in this project are intended to give a first-pass filter of sorts, assuming performance is on your list of requirements.