If you read the paper, nothing has changed. They still depend on the target talker not having competing co located sounds or voice.
Does that mean sound coming in the same direction as the person they are targeting? Does it mean sound near the person /object they are targeting?