2,960 karma · joined January 28, 2019
How could this possibly be true? It's not at all rocket science to create a static blog and serve it via a production grade web server (nginx, etc).
> The UX of DNS hasn't improved (it's still impossible for normies)
The UX of DNS sucks but we're talking about a single A record. Is that not within reach of a normie in the age of AI?
> Custom domains are not just vain, they're ephemeral. Certainly more so than, say, the domain of a blogging platform that's managed by a non-profit.
I can't think of a single free blogging platform that has stood up to the test of time. Depending on centralized resources, particularly when you're not paying for them, is the recipe for ephemerality. If you're going to pay for it why can't you afford a domain?
The 26% you miscategorized are people with pending charges. Everyone is innocent until proven guilty.
That's not even the worst scenario. There are plenty of websites that are nearly meaningless. Could you predict the next token on a website whose server is returning information that has been encoded incorrectly?
LLMs really do find the signal in this noise because even just pre-training alone reveals incredible language capabilities but that's about it. They don't have any of the other skills you would expect and they most certainly aren't "safe". You can't even really talk to a pre-trained model because they haven't been refined into the chat-like interface that we're so used to.
The hard part after that for AI labs was getting together high quality data that transforms them from raw language machines into conversational agents. That's post-training and it's where the armies of humans have worked tirelessly to generate the refinement for the model. That's still valuable signal, sure, but it's not the signal that's found in the pre-training noise. The model doesn't learn much, if any, of its knowledge during post-training. It just learns how to wield it.
To be fair, some of the pre-training data is more curated. Like collections of math or code.
Is the average human 100% correct with everything they write on the internet? Of course not. The absurd value of LLMs is that they can somehow manage to extract the signal from that noise.
Some studies have shown that direct feedback loops do cause collapse but many researchers argue that it’s not a risk with real world data scales.
In fact, a lot of advancements in the open weight model space recently have been due to training on synthetic data. At least 33% of the data used to train nvidia’s recent nemotron 3 nano model was synthetic. They use it as a way to get high quality agent capabilities without doing tons of manual work.
While I understand that this has been difficult for him and his company... hasn't it been obvious that this would be a major issue for years?
I do worry about what this means for the future of open source software. We've long relied on value adds in the form of managed hosting, high-quality collections, and educational content. I think the unfortunate truth is that LLMs are making all of that far less valuable. I think the even more unfortunate truth is that value adds were never a good solution to begin with. The reality is that we need everyone to agree that open source software is valuable and worth supporting monetarily without any value beyond the continued maintenance of the code.
It's extremely fast on good hardware, quite smart, and can support up to 1m context with reasonable accuracy
It seems similar to what you're describing.
However, is cost the biggest limiting factor for agent adoption at this point? I would suspect that the much harder part is just creating an agent that yields meaningful results.
I've noticed that open models have made huge efficiency gains in the past several months. Some amount of that is explainable as architectural improvements but it seems quite obvious that a huge portion of the gains come from the heavy use of synthetic training data.
In this case roughly 33% of the training tokens are synthetically generated by a mix of other open weight models. I wonder if this trend is sustainable or if it might lead to model collapse as some have predicted. I suspect that the proliferation of synthetic data throughout open weight models has lead to a lot of the ChatGPT writing style replication (many bullet points, em dashes, it's not X but actually Y, etc).