This makes me wonder too about the entire premise and worthiness of these evals. They orient themselves around normal one-shot interactions with a likely non-sys-prompted model with no built up context or memory of the person. I doubt the mentioned 'job loss' scenario is even contextually seen as a 'loss'; it is only a circumstance descriptor, a single snapshot without a history. Maybe to get the best advice we actually need to tell the LLM our entire story, not just a narrow request for a question; a question that - itself - is biased to our own imaginings of what problem we perceive ourselves as having, which humans are often bad at.