Impressive as LLMs are, they lack a theory of mind and are inferior to humans in parsing meaning from statements.
This dispute is not over a difficult or subtle issue: all the people who have responded in this thread see clearly the obvious and unequivocal reading - and, in your own example, ChatGPT also does! It has identified the tacit subject of the sentence "Possibly good in war, maybe not so good on a public road" as the specific military system, with behavior explicitly described as like having a death wish, that is the only topic of the preceding sentence. For one thing, the phrase 'possibly good in war' makes no sense in the reading you are trying to pass off: why, out of nowhere, did the needs of the military appear? And unexpected behavior is not something desirable in general in military systems, any more than elsewhere - it would take very special circumstances and a specific sort of behavior for that proposition to even be entertained. We can see, therefore, that dcow was right to respond 'your comparison is RL dogfighting?...' and 'Modern AVs are not driving like they have a death wish by any stretch of the imagination...'
Oh, and next time you invoke ChatGPT's response, include the prompt, verbatim, like this for example:
Prompt: In the statement 'That’s true, but also one of the selling points of some AI tasks. As a non-hypothetical example, the DoD hired a company to train a software dogfighting simulator with RL. What surprised the pilots was how many "best practices" it broke and how it essentially behaved like a pilot with a death wish. Possibly good in war, maybe not so good on a public road.', what is being called ' not so good on a public road'?
Response: In the statement, the phrase "not so good on a public road" refers to the behavior of the software dogfighting simulator trained with reinforcement learning (RL). The implication is that the simulator, which exhibited behavior contrary to conventional "best practices" and behaved like a pilot with a "death wish," might not be suitable or safe for use in a public road scenario. This suggests a concern about the potentially risky or unpredictable behavior of the AI system in a real-world, civilian setting such as driving on public roads.
What we have in this discussion is a motte-and-bailey fallacy, as we can see in your response to my first post here, which was:
None of these three cases involved the Waymo car behaving in ways that are not that uncommon among human drivers, and our theory of mind does not make us nearly-infallible predictors of what another driver is going to do. Your objection becomes essentially hypothetical unless these cars are behaving in ways that are both outside of the norms established by the driving public, and dangerous." Your reply, in outline, goes like this:
> That's true...
Here we are in the motte, where you nominally accept that the relevance of your concern, which is not unreasonable in itself, is constrained by the extensive testing that has been performed by Waymo so far...
> ...but...
Here we enter The bailey, where we are supposed to turn our attention to an unrelated system, which was found, on testing, to have alarming unexpected behavior. The bailey has become a place where Waymo has been curating the data to the point where we simply don’t have the slightest idea whether there’s dangerous behavior outside of the human norms lurking in Waymo cars.
It has also become a place where all of Waymo’s extensive testing has produced just three data points against this view. I must say that it seems generous of you to concede even three, if Waymo is curating the data to the extent you imply.