One solution might be to characterize individual drivers (agents) by running an RL experiment that mimics the interactions of individual car drivers. If the reward function is made to compare traffic patterns of the simulation vs real life, then eventually the RL model (a combination of agents) should converge to how humans drive.
So you would have various transition probabilities of a car driver moving from an attentive state into various inattentive states, with drivers having different reactions in each state.
It would also help to have a level of "variance" in individual drivers. So instead of having a bunch of drivers who are just as likely to make mistakes, you have some who transition far more easily into inattentiveness and some whose likelihood of damage/nuisance is higher than the standard driver when inattentive.
It seems entirely doable. (famous last words)