The context windows aren’t large enough, as I understand it. It might be possible via a chain-of-summarization, though.
Most importantly due to the context length not being long enough. If the context length was long enough, it is possible that they could do it with clever training. I only trained much smaller language models though.