The Overflow Offline project
stackoverflow.blog
stackoverflow.blog
I also have great memories from a University exam where we were allowed to have laptops that were not connected to the internet.
Here are the datasets: http://download.kiwix.org/zim/stack_exchange/
It's not clear to me why the data set shrank between 2019/3 and 2022/6; was something excluded? Compression improvements?
> stackoverflow.com_en_all_2019-02.zim 2019-03-12 19:53 134G
> stackoverflow.com_en_all_2022-05.zim 2022-06-17 12:36 75G
https://archive.org/details/stackexchange
It's the "official" place to get the data
I've download it several times and extracted my own contributions.
> ... to ensure that an up-to-date version of our dataset is easily available for those who need it, and will work to improve its readability and reduce its size so there is less friction for end users...
I was once a guest in a maximum security prison for a few hours. It gave me the eerie realization that my own freedom is entirely up to the guards and staff. When that door closes you are no longer free - you're at the mercy of a huge system. It's a terrible feeling. Talking with some of the inmates made me realize how much I had subconsciously dehumanized them as a group. It was eye opening.
Listening to stories of people who have been in prison, even if you never meet them, can help build empathy. Here's a good interview with a guy who taught himself to program while in prison: https://corecursive.com/prison-programming-with-rick-wolter/
If you're ever interviewing a formerly incarcerated person for a job, I'd encourage you to try very hard to keep your biases in check. Building a life after prison seems like a supremely difficult task.
[edit] That's what I get for reading the comments before the article! Unlocked Labs is linked in the 2nd sentence. :)
- The files are big. The data files were 486 GB and the log file (I haven't tried shrinking it yet) is 1.4 TB
- There were no foreign key relationships defined. Nor indexes on commonly queried columns. Easy enough to add them as you work on tuning it.
- To be usable by humans you'll need to set up some full text indexes for searching - they use Elasticsearch. SQL Server's Full-Text catalog doesn't get you where you need.
And I funded to work to run that on an Android phone https://play.google.com/store/apps/details?id=space.atrailin...
GitHub: https://github.com/neuml/codequestion
Article: https://medium.com/neuml/find-answers-with-codequestion-2-0-...
I'll throw in my own shameless plug: Self-host your Stack Overflow, Wikipedia etc on Sandstorm: https://apps.sandstorm.io/app/5uh349d0kky2zp5whrh2znahn27gwh... Obviously uses kiwix-serve as well. 3 years old, I need to make a better clip for updating it.
There's a fund:
https://opencollective.com/sandstormcommunity
The bigger it gets, the more time can be diverted toward it. I was recently paid to upgrade the Etherpad app. But in theory it could go toward core development (which wouldn't be me).
Of course, if you asked me, it always was. You couldn't assume great connectivity then and you often still can't today.
I don't know if it's available for html or js or css, or opengl.
It's been a very handy tool in my toolbelt.
> “We built the Sotoki (Stack Overflow to Kiwix) scraper in such a way that it can capture each and every one of the 180 Stack Exchange websites.”
Unclear to me if "can" means "does" or "will soon" or just "could"