Stack Overflow used to release their data archives quarterly on BigQuery. Looking at the BQ datasets, they were last updated Nov 2022, which doesn't have the latest 2023 info in the submission.
https://medium.com/snowflake/how-to-load-the-stack-overflow-...
The Hugginface repo unfortunately prefilters some of the tables/rows according to some criteria, making it less usable for general analytical queries that the BQ or SEDE datasets enable. If anyone knows of an 'XML-streaming' solution that directly samples from the Internet Archive's data dumps, I am all ears.
[1]: https://huggingface.co/docs/datasets-server/rows
[2]: https://huggingface.co/datasets/HuggingFaceGECLM/StackExchan...