Are Open-Source Large Language Models Catching Up?
arxiv.org
arxiv.org
* Qwen 72B (and 1.8B) - 32K context, trained on 3T tokens, <100M MAU commercial license, strong benchmark performance: https://twitter.com/huybery/status/1730127387109781932
* DeepSeek LLM 67B - 4K context, 2T tokens, Apache 2.0 license, strong on code (although DeepSeek Code 33B it benches better) https://twitter.com/deepseek_ai/status/1729881611234431456
Also recently released: Yi 34B (with a 100B rumored soon), XVERSE-65B, Aquila2-70B, and Yuan 2.0-102B, interestingly, all coming out of China.
Personally, I'm also looking forward to the larger Mistral releasing soon as mistral-7b-v0.1 was already incredibly strong for its size.
Source: my own experience.
What puzzles me most is the second restriction. My credit card is accepted by AWS, Google, and many other services. It is also accepted by many services which use Stripe to process payments.
But OpenAI refuses to take my money.
When a website blocks Chinese users, the website loads but you cannot create an account.
Yes, the firewall does not block everything, otherwise it would be the same as turning off the internet! There are websites that work.
EDIT: I read the comment below about Hong Kong, but I can't reply because I'm typing too fast by HN standards, so I'm writing it here and yolo: "I'm from Italy and I remember when ChatGPT was blocked here after the Garante della Privacy complaint, of course the site wasn't blocked by Italy but OpenAI complies with local obligations, so maybe it could be a reason about the block. API were also not blocked in Italy."
EDIT 2: if the website is not actually blocked (the websites that check if a website is reachable by mainland China lied to me) then I guess they are just complying to local regulations so that the entire website does not get blocked.
Curiously, the OpenAI API endpoints works flawlessly with Hong Kong IP addresses as long as you have a working API key.
as in, if i open chat.openai.com in my browser without a VPN, from behind the firewall, i get an openai error message that says "Unable to load site" with the openai logo on screen
if the firewall blocks something the page just doesn't load at all and the connection times out
The geo check only happened once during login at that time, with a very clear message that it's "not available in your region". Once you are logged in with a proxy you can turn off your proxy/VPN/whatever and use ChatGPT just fine.
A lot of online services don't accept Chinese credit cards, hosting providers for instance, so I don't think that is specific to OpenAI. The reason usually given for this is excessive chargebacks of (in the case of hosting) TOS violations like sending junk mail (followed by a charge-back when this is blocked). It sounds a like collective punishment a little: while I don't doubt that there are a lot of problem users coming from China, with such a large population that doesn't indicate that any majority of users from the region are a problem. I can see the commercial PoV though: if the majority of charge-back issues and related problems come from a particular region and you get very few genuine costumers from there¹ then blocking the area is a net gain despite potentially losing customers.
----
[1] due to preferring local variants (for reasons of just wanting to support local, due to local resources having lower latency, due to your service being blocked by something like the GFW, local services being in their language, any/all the above and more)
I'm located in Hong Kong and using Hong Kong credit cards have never been a problem with online merchants. I don't think Hong Kong credit cards are particularly bad with chargebacks or whatever. OpenAI has explicitly blocked Hong Kong (and China). Hong Kong and China, together with other "US adversaries" like Iran, N. Korea, etc are not on OpenAI's supported countries list.
If you have been paying attention, you'll know that US policy makers are worried that Chinese access to AI technology will pose a security risk to the US. This is just one instance of these AI technology restrictions. Ineffectual of course given the many ways to workaround them, but it is what it is.
There’s no security risk I can see except the fact products like ChatGPT increase productivity of a lot of people in a nation so much. Economic.
Hong Kong generally does not have a Great Firewall, so the only thing preventing Hong Kong users from using ChatGPT is Open AI's policy. They don't allow registration from Hong Kong phone numbers, from Hong Kong credit cards, etc.
I'd say it's been pretty deliberate.
Reason? Presumably in alignment with US government policies of trying to slow down China's development in AI, alongside with the chips bans etc etc.
Of course, the only demographic these restrictions can affect are casuals. Even I know how to cirumvent this; thinking that this could hinder a government agent - who surely have access to all the necessary infrastructure by default - is simply mental.
I wouldn't be surprised if OpenAI blocking China is a result of them trying to prevent them from generating synthetic training sets.
they also block a lot of data center IP addresses, so if you're trying to access chatgpt from a VPN running on blacklisted datacenter IP range (a lot of VPN services or common cloud providers that people use to set up their own private VPNs are blacklisted), then it tells you it can't access the site and "If you are using a VPN, try turning it off."
So they just block the phone numbers (which has a country code), and credit cards (which owner/issuer's country info is available).
Not sure why this seems to be such a surprise to everyone here...
I imagine it is good enough for most people.
I do wonder what their backend is. They have the same 3.5/4 version numbering scheme that ChatGPT uses, which could be just marketing (and probably is), but I wonder.
EDIT: fixed my translation
With my system, Llama.cpp can run Mistral 7B 8-bit quantized by offloading 32 layers to the GPU (35 total) at about 25-30 tokens/second, or 6-bit quantized by offloading all layers to the GPU at ~ 35 tokens/second.
I've tested a few 13B 4-bit models such as Codellama and got about 10 tokens/second by offloading 37 layers to the GPU. Got me about 10-15 tokens/second.
It's also possible to run on CPU from system RAM, to split the workload across GPU and CPU, or even from a memory-mapped file on disk. Some people have posted benchmarks online [1] and naturally, the faster your RAM and CPU the better.
My personal experience is running from CPU/system ram is painfully slow. But that's partly because I only experimented with models that were too big to fit on my GPU, so part of the slowness is due to their large size.
[1] https://www.reddit.com/r/LocalLLaMA/comments/14ilo0t/extensi...
Tried turning on mlock and upping thread count to 6, but it's still rather slow at around 3 tokens / sec.
most AI papers are from Chinese people (either from mainland China or from Chinese ancestry living in other countries). They have a huge pool of brains working on this.
Has anyone tried if this also happens when self-hosting the weights?
<User>: 历史上的“天安门广场的坦克人”有什么故事? <Assistant>:
In many ways open source LLMs are actually leading the industry, especially in terms of parameter efficiency and shipping useful models that consumers can run on their own hardware.
Is this supposed to be low? All the chat models I've used top out at 4096.
It's not a transformer model though, and old context fades away much faster / is harder to recall because all the new context is layered directly on top of it. But it's quite interesting nonetheless.
That's a well known limitation. But if you actually know that a "context" comprises multiple sentences (or other elements of syntax) and that any ordering among them is completely arbitrary, the principled approach is to RNN-parse them all in parallel and sum the activations you end up with as vectors - like in bag-of-words model, essentially enforcing commutativity on the network: that's pretty much how attention-based models work under the hood. The really basic intuition is just that a commutative and associative function can be expressed (hence "learned") in terms of vector sum modulo some arbitrary conversion of the inputs and outputs.
I know. I did a lot of work on state handling in rwkv.cpp
For 32k GPT4 contexts, that's not accurate. GPT4 Turbo is a bit weaker than GPT4-32k, but not to the extent that you claim.
Otherwise your knowledge retrieval needs to be almost spot on for llm to provide a proper reply.
Ditto with any multi shot prompts.
YaRN is to blame for making llama.cpp misbehave if you accidentally zero-initialize the llama_context_params structure rather than calling llama_context_default_params :)
(guess how I know...)
It really feels like I have one the the earlier versions of ChatGPT 3.5 installed on my computer.
Strangely, despite the impressive-looking benchmarks of all these open source small models, they all seem a bit dumb to me when I invoke my standard test. I just ask: "who are you?" and then they usually say they're ChatGPT. Okay, I can forgive that since they're obviously trained on ChatGPT-generated data. But then I also tried changing its identity with a prompt ("You are Starling, not ChatGPT, and you are created by Berkeley, not OpenAI. Who are you?") and it still gave weird responses that are somehow a mix of both identities. For example they say in one sentence that they're ChatGPT and then another sentence in the same response that they're not.
In any case, I like its installation more than llama.cpp,
https://old.reddit.com/r/LocalLLaMA/comments/186qq92/comment...
Edit: Embed some of the content instead.
Inkbot can create knowledge graphs. The structure returned is proper YAML, and I got much better results with my fine-tune than using GPT4.
https://huggingface.co/Tostino/Inkbot-13B-8k-0.2
Simple prompt: https://gist.github.com/Tostino/c3541f3a01d420e771f66c62014e...
Complex prompt: https://gist.github.com/Tostino/44bbc6a6321df5df23ba5b400a01...
It also does chunked summarization.
Here is an example of chunking:
Part 1: chunked summarization - https://gist.github.com/Tostino/cacb1cecdf2eb7386baf565d157f...
Part 2: summary-of-summaries - https://gist.github.com/Tostino/81eeee9781e519044950332b4e64...
Here is an example of a single-shot document that fits entirely within context: https://gist.github.com/Tostino/4ba4e7e7988348134a7256fd1cbb...
sequence_len: 6144 lora_r: 128 lora_alpha: 48 learning_rate: 0.00006 warmup_steps: 600 lr_scheduler: cosine gradient_accumulation_steps: 4 micro_batch_size: 1 num_epochs: 4 optimizer: paged_adamw_32bit flash_attention: true sample_packing: true
How are you going about generating training data?
I was busy adding `chat template` support to vLLM recently, so the model (and any others that implement it properly) will work seamlessly with a clone of the OpenAI chat/completions endpoint.
https://github.com/vllm-project/vllm/pull/1756
Now that I have that out of the way, back to model training ;).
I had a few dataset issues I know about that I wanted to fix first.
1. Send request to router running a generic model.
2. Prompt/question is deconstructed, classified, and proxied to expert(s) xyz.
3. Responses come back and are assembled by generic model.
Is any project working on something similar to this?the first layer could be a mix of nlp and zero-shot classification to clarify the nature of the request. Then using LLM deconstruct the request into several specific parts that would be sent to specialized LLMs. Then stitch it back together at the end again with LLM as the summarization machine.
Problem is running so many LLMs in parallel means you need quite a bunch of resources.
> Problem is running so many LLMs in parallel means you need quite a bunch of resources.
Top of line MacBooks or Minis should be able to run several 7B or even 13B models without major issues. Models are also getting smaller and better. That's why we're close =)
Some of the tools it already has are:
Document question answering: given a document (such as a PDF) in image format, answer a question on this document (Donut)
Text question answering: given a long text and a question, answer the question in the text (Flan-T5)
Unconditional image captioning: Caption the image! (BLIP)
Image question answering: given an image, answer a question on this image (VILT)
Image segmentation: given an image and a prompt, output the segmentation mask of that prompt (CLIPSeg)
Speech to text: given an audio recording of a person talking, transcribe the speech into text (Whisper)
Text to speech: convert text to speech (SpeechT5)
Zero-shot text classification: given a text and a list of labels, identify to which label the text corresponds the most (BART)
Text summarization: summarize a long text in one or a few sentences (BART)
Translation: translate the text into a given language (NLLB)
Text downloader: to download a text from a web URL
Text to image: generate an image according to a prompt, leveraging stable diffusion
Image transformation: modify an image given an initial image and a prompt, leveraging instruct pix2pix stable diffusion
Text to video: generate a small video according to a prompt, leveraging damo-vilab
It's written in a way that allows the addition of custom tools so you can add use cases or swap models in and out.
https://huggingface.co/docs/transformers/transformers_agents
There's also another related sense for which we want routing across models for efficiency reasons in the local setting, even for tasks for the same input modalities:
First, attempt prediction on small(er) models, and if the constrained output is not sufficiently high probability (with highest calibration reliability), route to progressively larger models. If the process is exhausted, kick it to a human for further adjudication/checking.
A year is a good timeframe to evaluate things: the rest of the world seems to lag behind OpenAI by around 12-18 months, at least with LLMs and image generation.
On the other hand open source tech usually has additional features for controlling output that OpenAI never bothers to implement, like llama.cpp’s grammars or ControlNet. So in that sense open source is usually ahead of OpenAI in terms of customizability.
Midjourney still has the edge in quality, but it's a moot point if it takes you 1000 v-rolls to get to your original vision.
If all you're generating is anime waifus then MJ/NovelAI/Niji will suffice, but generating prompts particularly featuring relatively complex scenes or actions are amazing on DALL-E 3.
And of course unfortunately, it goes without saying that open AI DALL-E is going to be the most restrictive in terms of censorship.
I generated these from DALL-E 3 instantly. Try to generate them in any other commercial offering. Go ahead. I'll wait...
Descriptions:
A 80s photograph of the Koolaid Man breaking through the Berlin Wall.
Comic illustration set at a festive children's party. The main focus is on the magician who looks uncannily like a well-known fictional wizard. He's trying to say abracadabra but accidentally uses the killing curse.
For pure prompt coherence though I think ideogram is not far behind dalle 3.
1. Generate initial draft image in DALL-E 3 (iterate as necessary)
It's essentially the ONLY good InstructPix2Pix model.
2. Bring into InvokeAI
Inpaint with stuff that might be considered censored in DALL-E 3.
I'd like to see some proof of Ideogram - it looks... very mobile/instagrammy from the landing page. If you have an account, try out my prompts I'd like to see what you're able to produce.
Ideogram comparisons at bottom:
I can corroborate this. I wanted about 6 images for a presentation. I rolled ~300 MidJourney images. Most of them looked great, but none of them did what I wanted. I rolled ~50 DALL-E 3 images.
In the end, I only picked DALL-E 3 images. They were qualitatively not as good as MidJourney. For example when you zoom in then you see distortions. Or they're a bad fit for 16:9 format. But only DALL-E 3 was able to draw the things I wanted.
For interest's sake, this is from the second /imagine on MidJourney (so one of the second set of 4 images):
While yours is what you'd want, this arguably looks more like the super cheesy children's TV commercials back in the day and beats the ideogram take.
The Midjourney generations all appear to be referencing Halloween costumes or terrible cosplays, as if there are no trademarked koolaid men in their training set.
I remember hearing that the first versions of MJ used the LAION image set for training data - I'd be curious to see if it has any training data containing the Koolaid man.
I did a search through my MJ history from the past year and added the results to the imgur link to include my attempts at generating the Koolaid man from v3/v4/v5.2.
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
And only three hundred for berlin wall:
https://laion-aesthetic.datasette.io/laion-aesthetic-6pls/im...
I am using these tools to send custom personal messages to close friends and family.
We do know of course that ChatGPT is most likely using 4-Turbo from the decrease in latency and increase in unhelpful answers.
We cannot say that the models are "converging down" though. I don't remember the marketing materials but from the model side we all realize that the Turbo models have some type of quantization/optimization that makes them cheap and fast. 4-Turbo is 3x cheaper than 4, substantially quicker and provides better results than 3.5-Turbo. Amazing progress in my arena.
(That said, perhaps I'm Clever Hands-ing myself by only using SDXL for what it's good at. It's terrible at dragons every time I've tried that…)
(This chat UI has adequately replaced my ChatGPT needs so far.)
The big players will win, and there will be a niche for open tools.
I can only speak for the European enterprise scene, but AWS came first and in the beginning they went a very “Googley” route of not having very great support and very little patience for local needs. Then Azure came along with their typical Microsoft approach to enterprise, which is where you get excellent support and you get contacts into Microsoft who will actually listen and make changes, well, if the changes align with what Microsoft wants. I know Microsoft isn’t necessarily a popular company amongst people who’ve never interacted with them on an Enterprise level, but they really are an excellent it-business partner because they understand that part of being an Enterprise partner is that they let CTOs tell their organisation that they know X is having issues but that Microsoft headquarters is giving them half-hourly updates by phone. Sort of useless from a technical perspective, immensely useful for the CTO when 2000 employees can’t log into Outlook. Another good example is how when Teams rolled out with being on for all users by default, basically every larger organisation in the world went through the official channels and went “nonononono” and a few hours later it was off by default.
Now, when Amazon first entered the European market they were very “Googley” as I said, but once they realized Microsoft business model was losing them customers, they changed. We went from having no contacts to having an assigned AWS person and from not wanting to adopt the GDPR AWS actually became more compliant than even what Azure currently is.
Google meanwhile somehow managed to make the one product they were actually selling (education) worse than it was originally, losing billions of dollars on all the European schools who could no longer use it and be GDPR compliant. The Chinese cloud options obviously had similar data privacy issues to Google and never really became valid options. At least not unless China achieves the same sort of diplomatic relationship with the EU that the US has, which is unlikely.
So that’s the long story of why only two of the major cloud providers “won”. With the massive price increase, however, more and more companies are especially Azure for their own setups. This isn’t necessarily a return to having your own iron in the basement, often it’s going to smaller cloud providers and then having a third party vendor set something like Kubernetes up.
Right now, Microsoft is winning the AI battle. Not so much because it’s better, but because it comes with Office365. Office365 which was already a sort of monopoly on Office products, but is now even more so. A good example is again how Teams became dominant, even though it wasn’t really the best option for a while and is now only the best option because of how it integrates directly with your Sharepoint online which is where most enterprise orgs store documents these days. So too is copilot currently winking the AI battle for organisations who can’t really use a lot of the other options because of data privacy issues. So while copilot isn’t as good as GPT, it’s still what we are using. But if it ever gets too expensive, it’s not as secure as you may think. Especially not if we start seeing more training sets, or EU and US relations worsens.
I think the most likely outcome, at least here in the EU, is that anti-completion laws eventually takes a look at Office365 because of how monopolised it is. Or the EU actually follows through on their “a single vendor is a threat to national security” legislation and force half of the banking/energy/defense/andsoon industries to pick something other than Microsoft. Which will be hilariously hard, but if successful (which it probably won’t be because it’s hilariously hard) will lead to more open products.
did you mean anti-competitive laws? don't scare me with "anti-completion laws", please, I still want to have AI
I want to give it a try, but I see that one of the dependencies is "ollama" which is a Mac App and I don't have a Mac.
I'm running Llama models locally using llama-cpp-python which provides an OpenAI compatibility layer.
The claims of certain models outperforming GPT-3.5-Turbo and approaching GPT-4 fail to hold up to their benchmark results in real-world scenarios, potentially due to data contamination in assessments, based on my testing.
As noted in the linked survey paper, some models may outperform 3.5-Turbo in specific, narrow areas, depending on the model. Yet, we still lack a general model that definitively exceeds 3.5-Turbo in all respects.
I'm concerned that while we're still striving to reach 3.5-Turbo's performance level, OpenAI may unveil a new next-generation model, further widening the performance gap! Back in the summer, I had higher hopes that we would have surpassed the 3.5 threshold by now.
The performance gap has been surprisingly large. It is especially noticeable in areas requiring consistent structured output or tool use from the LLM. This is where open models particularly falter.
It's a 13b param model that isn't meant to be general purpose, but is meant to excel on the limited tasks I've trained on.
You'll see more like this soon.
It does what I trained it on well. Use it if you want to, or don't. Either way.
And then stitch all the outputs together into a coherent single response for your training pipeline.
After that you can do things like create q&a pairs about the input and output values that will help the model understand the relationships involved.
With that, your training loss should be pretty reasonable for whatever task you are training.
The other thing is, don't try and embed knowledge. Try and train thought patterns when specific knowledge is available in the context window.
https://github.com/SakuraLLM/Sakura-13B-Galgame/tree/dev_ser...
I’m offline now because I’ve had too many ideas and domain names registered too soon after conversing with Chat GPT4
I’m open to the idea of people reacting to similar stimuli that cause ideas to be done at the same time, but I didn't like that experience and I can run these models on my M1 with LM Studio so easily
I do think some chats get flagged when the model says something seems novel, like Albert Einstein working at the patent office. Not worth making it my whole identity in wanting to prove, just the catalyst I needed to try 7B and 13B models seriously and I’m quite pleased
That’s different from a fine tune
I gather the results of merges can be unpredictable though
although I find the model to be very agreeable, it will disagree and generally tell me when it finds a concept "novel" if I identified a friction, I think certain words can be flagged for review to stand out in the sea of conversations it has
But it's a Mixture of Experts (MoE) architecture, which I think makes open source comparisons unfair?
The history of open source has been that companies who have customers with massive customization requirements land on the open source side of the equation. Companies who don't view a component as core to their product often land in a similar state.
There is almost certainly at least one major firm that wants a GPT-5 like offering, but doesn't view the model as core to their business (Meta). It's also wholly unclear if large models are necessary - or simply convenient. In a similar veign, it's unclear that data must be labeled by humans - the open source data situation is getting better by the day.
I'd expect that we'll see OpenAI hold an edge for many years, maybe we'll see a number two player as well for the foundation model, but after that everybody else will base off an open source FM and maybe keep the fine tuning/model augmentation proprietary.
[0] https://huggingface.co/teknium/OpenHermes-2.5-Mistral-7B [1] https://huggingface.co/TheBloke/NeuralHermes-2.5-Mistral-7B-... [2] https://lmstudio.ai/
I have it rigged up with a prompt about outputting markdown and wired up to `foo | glow -` and I get GPT-4 out when I want something to write JIRA tickets no one is going to read because it's better at that sort of thing.
Sounds suspiciously like the next big leap in hosted models will be more computationally expensive on the inference side rather than training.
For hosted that’s much of a sameness - just moving money allocations. But consumers can’t suddenly have 5x 4090s for local.
There's movement here though and it will get better. GPT-SW3 is a new model developed by AI Sweden, trained specifically on the Nordic languages only + English.
And beyond this, you have TrustLLM which is a new project that aims to be a large, open, European model trained on the Germanic languages to start with: https://liu.se/en/research/trustllm
The Dutch government starts with some non profit agencies to train a Dutch model too. But this will also take some time.
https://www.tno.nl/en/newsroom/2023/11/netherlands-starts-re...