Historically, the amount of data available to traffic lights controllers was rather rudimentary, basically a simple presence/occupancy detector covering the area immediately upstream of the stop line and/or some point somewhat (a few seconds of driving time) further upstream of the stop line and/or the whole area in-between. The detector could be implemented via an inductive loop buried within the pavement, or it could be a camera, but with the camera still usually working only as a simple presence detector with no advanced detection capabilities where it could actually intelligently "see" individual vehicles.
With such a setup, the traffic lights controller can't actually see that there's only 4-5 cars approaching on the main road – either they're far away enough that they don't even register on the detector for the main road yet (in which case a vehicle waiting on the crossroad will naturally trigger the phase change without delay), or even if they're already near enough to the intersection to trigger the detector, the traffic lights controller can't actually know that it's "only" 4-5 vehicles and its strategy might hence be that if the main road has had more than x seconds of uninterrupted green time, a demand from the side street will be granted immediately, regardless of traffic conditions.
With a state-of-the-art system it'd be different – with more and better sensors the traffic lights controller could actually track individual vehicles (or platoons of vehicles) as they approach the intersection and try making the kinds of trade-off you mention – but somebody needs to pay for that kind of upgrade. (And I suspect we haven't necessarily reached the point where the price premium for that kind of advanced traffic lights controlled compared to classic traffic actuation is negligible enough for it to be the default at least for all new intersections and equipment renewals, either…)