Because it looks more credible, obviously. In a sense it's cargo cult science: people observe this is the style of science, and so copy just the style; to a casual observer it appears to be science.
1. Preprint servers create DOIs, making works better citable.
2. Preprint servers are archives, ensuring works remain accessible.
My blog website won't outlive me for long. What happened to geocities could also happen to medium.
If you're writing a paper about a longstanding math problem and the solution gets published on 4chan, you still need to cite it.
Pure math tends to be much more conservative in citations than other fields though, and even when writing a paper about a longstanding math problem you wouldn't necessarily bother to include existing solutions. You reference the things you actually used, and even then you assume some common background knowledge for your audience and don't reference every little undergrad topology theorem or whatever. The point is to be honest with the reader about what was helpful for this work in particular, both to properly attribute things you actually used and to make any searches based on your work more targeted and fruitful.
Both models are fallible, which is why discernment is so important.
That something is unreviewed does not mean that it is bad or useless.
Assuming that you are referring to the Arxiv, they can't:
Yet it's what we train LLMs on.
> We introduce phi-1, a new large language model for code, with significantly smaller size than competing models: phi-1 is a Transformer-based model with 1.3B parameters, trained for 4 days on 8 A100s, using a selection of ``textbook quality" data from the web (6B tokens) and synthetically generated textbooks and exercises with GPT-3.5 (1B tokens). Despite this small scale, phi-1 attains pass@1 accuracy 50.6% on HumanEval and 55.5% on MBPP. It also displays surprising emergent properties compared to phi-1-base, our model before our finetuning stage on a dataset of coding exercises, and phi-1-small, a smaller model with 350M parameters trained with the same pipeline as phi-1 that still achieves 45% on HumanEval
We train on the internet because, for example, I speak a fairly niche English dialect influenced by Hebrew, Yiddish and Aramaic, and there are no digitised textbooks or dictionaries that cover this language. I assume the base weights of models are still using high quality materials.
In addition, peer reviews are anonymous for both sides (as far as possible).
It is weird how people use a platform exactly how it is supposed to be used.