But I think this is on the right track. As a simple approach you can have final-cost ratio. "Unknown" bidders would use the average score across all projects (or all first-time projects?) and known-good bidders would be able to use their own score (which may be close to 1). Known-bad bidders may have 1.5-2 multipliers which "corrects" for their underbid.
Of course this too simple. You want to:
- Account for time.
- Account for how easy they were to work with.
- You need to factor in unexpected changes. (which do have a cost, but how do you determine the true cost?)
- Using the average for unknowns may make it too difficult for new companies to enter. I way I see it you want some new entries to keep your pool in check. (Although the majority of projects, and quite possibly the largest projects should be going to known-good companies).
I think this is part of the problem. There are so many factors that it is difficult to rigidly define them, and if they aren't rigidly defined than you have chances for corruption.
The TL;DR is that unless you just let a human make the decision it is a very hard problem, and right now we don't think that humans are trustworthy enough.