We are getting a lot of great models out of China in particular (Yi, Qwen, InternLM, ChatGLM), and some good continuations like Solar.
Lots of amazing papers on architectures and long context are coming out.
Backends are going crazy. Outlines is shoving constrained generation everywhere, and Lorax is a LLM revelation as far as I'm concerned.
But you won't hear about any of this on Twitter/HN. Pretty much the only thing people tweet about is vllm/llama.cpp and llama/mistral, but there's a lot more out there than that.
I also hang out on a few Discord servers: - Nous Research - TogetherAI / Fireworks / Openrouter - LangChain - TheBloke AI - Mistral AI
These, along with a couple of newsletters, basically keep a pulse on things.
And yeah, /r/LocalLlama seems to be getting noisier.
TBH I just follow people and discuss stuff on huggingface directly now. Its not great, but at least its not discord.
Some are pretty good! Check out this little curated nugget: https://llm-tracker.info/
I used to follow one with a UI that resembled HN itself, but now I can't find it in my bookmarks, lol.
Discord is so time inefficient, its almost hilarious. For every incredible conversation between experts you observe, you have to weed through 200 times as much filler.
I can't say I blame them either, there is a lot of insane crypto-like fraud in the LLM/GenAI space. I watch the space like a hawk... and I couldn't even tell you how to filter it, it's a combination of self-training from experience and just downloading and testing stuff myself.
https://news.ycombinator.com/item?id=38505986
Zero interest for some reason.
Edit: Deepseek coder has 4 submissions to HN with almost zero interest.
TBH I do not follow OpenAI much. I like my personal models local, and my workplace likes their models local as well.
No one knows about it! Which is ridiculous because batched requests with loras is mind blowing! Just like many other awesome backends like InternLM's backend, LiteLLM, Outline's VLLM fork, Aphroidte, exllamav2 batching servers and and such. Heck, a lot of trainers don't even publish the loras they merge into base models.
Personally we are waiting on the integration with constrained grammar before swapping to Lorax. Then I am going to add exl2 quantization support myself... I hope.
Also check out the new prefix caching, I see huge potential for batch processing purposes there!
Everything is moving so fast!
I will say that if you want to explore the forefront of this multi-LoRA inference, definitely worth giving LoRAX a look. We just added support for per-request model merging (https://predibase.github.io/lorax/guides/merging_adapters/) as an example, and are planning on continuing to double down on this idea of combining adapters in some pretty unique ways.
> But the company’s CEO, Sam Altman, says further progress will not come from making models bigger. “I think we're at the end of the era where it's going to be these, like, giant, giant models,” he told an audience at an event held at MIT late last week. “We'll make them better in other ways.”
https://www.wired.com/story/openai-ceo-sam-altman-the-age-of...
And there's certainly more to be found in training longer on better datasets and adjusting the architecture.
On topic of LLMs and diffusion models I doubt they're really having that much of a positive social impact, I would bet that in the long run they'll automate more of the jobs humans would like to do compared to the ones we don't, and erode trust on a global level. But the economic gain is undeniable so we're doing it anyway. Automating all work would be a good thing in an ideal world, but in practice it's the one thing that still gives the lower classes some leverage over the capitalists and a world where the average person is expendable and unemployable isn't good for society.
Change will happen - necessarily so, in fact, as more automation is brought online, because the crowd of people excluded from the economy as "redundant" will not be happy about that state of affairs, and it will only grow in size. At some point, they will understand that the real problem is not that they don't have work, but that they don't have a fair share of the automation pie. And things like, say, property rights on means of production / automation stop to matter if most people don't believe in them anymore.
Yeah, they're an arms dealer now.
Think differently is again the way forward.
Anyone else had this experience? It seems like they're actually locking down some very helpful use cases, which maybe falls into the "safety" category or, more cynically, in the "we don't want to be sued" category.
I also suspect there’s some dynamic laziness parameter that’s used to counteract increased load on their servers. You can ask the same prompt over and over in a new chat throughout the day and suddenly instead of writing the code you asked for or completing a task, it will do a small part of the work with “add the code to do xyz here” or explain the steps required to complete a task instead of completing it. It happens with v4 as well.
That's a brilliant risk mitigation mechanism! The AI won't recursively self-improve to superhuman levels if it just keeps getting tired of thinking.
I guess, as one of the other replies here said, it was never "allowed" to give you text for a contract according to the TOS, but then it would've been best if it never replied so effectively to my earlier prompts. Taking it away just seems lame.
Edit: bad typing
- They have better just stuff, but aren't release it yet. They are at the top already, it makes sense they would hold their cards until someone got close.
- They are more focused on AGI, and not letting themselves get side tracked with the LLM race
- LLMs have peaked and they don't want to release only a minor improvement.
FWIW OpenAI seems to have a corroded definition for AGI that is essentially "[An] AI system generally smarter than humans".
They don't seem to use the typical definition I'm used to of some variation of autonomy or (pseudo)-sentience.
So their LLM race is the race for AGI
This hypothesis has a curious habit of surfacing when OpenAI is fundraising. Together with the world-ending potential of their complete-the-sentence kit.
A probabilistic word generator is still a word generator. It might be slightly better next version, we've seen it get worse, its more of the same
We can talk about AGI all we want, but it wont be built on the same technology as these word generators. There will have to be a technological breakthrough years before we even get close
The companies focusing on LLMs right now are dealing with
* Generate better words
* Make it cheaper to generate (use less compute)
* Find better training material
There is a ton of money to be made, but its still more of the same
In both cases, intelligence appears to be just an interesting emergent side effect.
But they probably have something more powerful internally. GPT-4 took months to be available to the public.
GPT 3.5 came out 2 years after GPT 3
GPT 4 came out 1 year after 3.5
I wouldn't think this matters as long as you charge enough to use that API. For example, you could have a tiered pricing structure where the first 100k words per month generated costs $0.001 per word, but after that it costs $0.01 per word.
This kind of pricing would also make it much less compelling for other users. 100k tokens is nothing when you're doing summarization of large docs, for example.
EDIT: Clearly the point is lost on the repliers. There is no general understanding of what 'general intelligence' is. By many metrics, ChatGPT already has it. It can answer basic questions about general topics. What more needs to be done? Refinement, sure, but transformer-based models have all qualifications to be 'general intelligence' at this point. The responses are more coherent than many people I've spoken with.
At the end of the day, as with most things, AGI is a meaningless term because no one knows what that is.
https://en.m.wikipedia.org/wiki/Artificial_general_intellige...