StackOverflow's database was public and shared at the Internet Archive until recently: https://archive.org/details/stackexchange
They've now moved it onto their own infra, ostensibly so people have to agree not to use it for LLM training: https://meta.stackexchange.com/questions/401324/announcing-a...
If they really did close off db downloads then I'll never answer a question there ever again. I bet many others wont either. Maybe that's also part of why SE will fail.
The only deal I personally am willing to tolerate is me spending time providing quality questions and answers in my domain in return for being able to download all questions and answers under the cc-sa license and being able to freely download them for offline viewing without an account (via kiwix). No other arrangement is acceptable.
I'll look further into this and update this comment with my findings.
Edit: Yep. No more answers from me. Talk about throwing the baby out with the bath water.
Maybe small community forums are really the only way to share knowledge these days :(
I thought of mailing lists as archaic, but maybe we just regressed from there.
Best for Society: Open Source <-- Where SO was
Medium for Society: Proprietary, but openly browsable <-- Where SO is
Worst for Society: Proprietary, not browsable <-- LLM-based code assist tools
Is Photoshop.exe "interpretable" by anybody with a copy (of windows)? How about a binary that's been heavily decompiled, like a Mario game?
Don't get me wrong, llama is at least more open than OpenAI and that may be meaningful.
Photoshop, or any compiled binary, isn't meant to be open source and the code isn't meant to be reviewable. Llama is called open source, though the most important part isn't publicly available for review. If llama didn't claim to be open source I don't think it would matter that the model itself and all the weights aren't available.
If your argument is just that most software is shipped as compiled and/or obfuscated code, sure that's how it is usually done. That isn't considered open source though, and the line with LLMs seems to be very gray - it can be "open source" if the source code for the training logic is available even though the actual code/model being run can't be reviewed or interpreted.
Personally I can't say I care as much about what the training set is, I want to know what's actually in the model and used at runtime/interpretation.
When I said "it's not really open source", I was referring to the fact that there are restrictions on who can use Llama.
Whoa there Nelly!
If we're going to resurrect things, can we do it whilst leaving PHP in the past?
Either your site's content stays hidden behind discord, or an LLM's bot/minion scrapes all your content and makes visiting your site superfluous, thereby effectively killing your site.
Source: I bothered a lot of people on the Internet about C++ when I was child.
Information is only useful if it's accessible. If they are asking questions it's because the information they want is in practice not accessible to them.
When I was in high school I read the docs, and learned C++ from books and MSDN. Granted my access to the internet was rather limited back then, but it also never crossed my mind to bother people for things I could easily lookup myself.
Growing up in a RTFM, "search the forum first before asking" environment is seen as toxic today, but it really helps keeping certain behavior in check thats a drag on society as a whole.
One of the best mentors/bosses I ever had never answered coding questions directly, but always in the form of a question so I could look it up and learn for myself.
I try to do the same with my junior devs today, unless there's time constraints or they're under stress, I try to let them figure out the final answer themselves.
I just hit that point with some TensorFlow stuff because I started hitting the limits of what ChatGPT could answer successfully, and I think that's fine. But maybe good that I couldn't get everything out of it or it may have delayed my learning further yet. Which I guess reinforces your point.
So, every comment on a post roughly equals 100 views.
There will always be new libraries and software updates, and those will always have corners and edge cases that will foster questions. The LLMs won’t have the answers out of the box until they have something to train from, so there’s still room for StackOverflow.
There should be a product, something that you install to capture your own data, and then pseudo-anonymize it and sell it back to databurgers (I meant to write data brokers but I kept this misspelling for lolz).
Is there?
Assistant LLMs like Limitless.AI merged with their older desktop scraper app could be repackaged to do that...
New industry? Or old?
I used to like reading StackExchange sites as a social media site--lots of interesting questions and clever answers. Today, votes have slowed down and the best answers are from 2017, and only niche questions can avoid being closed.
> but as microprocessors become faster, I believe we'll be able to generate artificial realities of useful information we can't live without which are superior to Stack Overflow today.
Was this written by an LLM bot? It seems…off.