Mistral 7B
mistral.ai
mistral.ai
Meta wouldn't make LLama open source.
DeciLM wouldn't make theirs open source.
All of them wanted to claim they were open source, while putting in place restrictions and not using an open source license. So I think it's worth giving Mistral big kudos here for actually doing it and going Apache licensed.
There must be others as well?
Estimating, there is more than 100 models with a apache-2.0 license.
They have a 40B one and IIRC are part of the Linux foundation now too
Falcon 40B is under the Apache license, and as a 40B model it outperforms this 7B model on most tasks.
To stack these items on top of each other in a stable manner, follow these steps:
1. Start with the 9 eggs. Arrange them in a neat pile in the center of your workspace. 2. Next, place the laptop on top of the eggs. Make sure it's centered on top and not leaning to one side. 3. Then, carefully pour the bottle of water over the laptop, making sure the water reaches all sides of the laptop and covers it completely. 4. Finally, take the nail and use it to secure the stack together. You can do this by driving the nail through each item in the stack, connecting them all. 5. Double-check that everything is securely fastened before moving or transporting the stack.
Remember to be careful when handling delicate items like eggs and laptops, and always prioritize safety.
Question: is there something like this, but with the "function calling api" finetuning? 95% of my uses nowadays deal with input/output of structured data (JSON, basically), and I kind of fear that OpenAI will no longer support thos specialized models after a few months... I know its probably not that glorious compared with multimodal chat UIs, but a workhose like nothing else for automation!
https://github.com/ggerganov/llama.cpp/blob/master/grammars/...
Its actually better than a specialized model, during token generation it constrains the possible output tokens to an arbitrary grammar (like, say, JSON syntax). So it will work "perfectly" with any model with a basic understanding of the format.
Kobold.cpp and text-generation-ui already support this, and both will run on your mac.
https://github.com/mistralai/mistral-src/issues/2
edit: nevermind https://ollama.ai
If I wanted to experiment with working with models and finetuning offline should I just save for a 32g vram+ video card?
Our approach also extends to EBNF grammars and LALR parsing. There's an example of that in the repository. It builds off of the Lark library, so you can use existing grammar specifications instead of starting from scratch.
total duration: 33.964492834s
load duration: 1.471584ms
prompt eval count: 8 token(s)
prompt eval duration: 385.334ms
prompt eval rate: 20.76 tokens/s
eval count: 418 token(s)
eval duration: 31.887033s
eval rate: 13.11 tokens/s
NotADirectoryError: [Errno 20] Not a directory: 'hipconfig'
At this point, it should be standard practice to address this concern. Any model which fails provide good evidence they don't have benchmark leaks, should not be trusted until its datasets can be verified, the methodology can be replicated, or a good, independent, private benchmark can be made and can be used to evaluate the model.
Manually benchmarking it by coming up with 20 questions and feeding it to a pair of models and blindly choosing the best result can give you a pretty good figure.
And that can probably be done in under 20 mins of human time.
I would be curious what does he mean by "semi-automated system for detecting benchmark leaks. " though.
magnet:?xt=urn:btih:208b101a0f51514ecf285885a8b0f6fb1a1e4d7d&dn=mistral-7B-v0.1&tr=udp%3A%2F%http://2Ftracker.opentrackr.org%3A1337%2Fannounce&tr=https%3A%2F%http://2Ftracker1.520.jp%3A443%2Fannounce
I remember not trusting these guys since they raised a lot of money with not much of anything but if this performs well it def bumps their credibility.---------------------
>>> Write code that will track swipes in Typescript
. Swipe should be detected when the user moves a finger over 30% of the screen, and then drags it all the way to the end of the screen.
## Getting started You need to have yarn or npm installed before you can run this app
### Running in browser ```bash yarn install yarn serve ```
### Running with docker Use the docker-compose file located at ./docker/docker-compose.yml ```bash cd <path-to-project> docker-compose up ```
## Screenshot 
Since Mistral is just a 7B parameter model, it's obvious that you won't be able to have it straight up write accurate code, it's simply too small for being able to accomplish something like that, unless you train the model specifically for writing code up front.
I guess if all you're looking for is a model to write code for you, that makes sense as a "hello world" test, but then you're looking at the wrong model here.
What you really want to do if you're looking for a good generalized model, is to run a bunch of different tests against it, from different authors, average/aggregate a score based on those and then rank all the models based on this score.
Luckily, huggingface already put this all in place, and can be seen here: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
This Mistral 7B model seems to earn itself a 3rd place compared to the rest of the 7B models added to the leaderboard.
Edit: As mentioned by another commentator, this also seems to be a base model, not trained specifically for request<>reply/chat/instructions. They're (or someone) is meant to fine-tune this model for that, if they want to.
What do you mean by this? MoE?
To track swipes in TypeScript, we can use a library like `react-native-gesture-handler`. Here's an example of how to use it to track horizontal and vertical swipes:
```typescript
import React from 'react'; import { GestureHandlerEvent, GestureHandlerProps } from 'react-native-gesture-handler';
const SwipeableContent: React.FC<GestureHandlerProps> = ({ onStart, onEnd, onMove }) => { const handleHorizontalSwipe = (event: GestureHandlerEvent) => { if (event.type === 'horizontal') { onStart(event); } else { onEnd(event); } };
const handleVerticalSwipe = (event: GestureHandlerEvent) => {
if (event.type === 'vertical') {
onStart(event);
} else {
onEnd(event);
}
};
return (
<View>
<PanGestureHandler onStart={handleHorizontalSwipe} onMove={onMove}>
<Text>Drag me horizontally</Text>
</PanGestureHandler>
<PanGestureHandler onStart={handleVerticalSwipe} onMove={onMove}>
<Text>Drag me vertically</Text>
</PanGestureHandler>
</View>
);
};``` In this example, we use the `PanGestureHandler` component from `react-native-gesture-handler` to track horizontal and vertical swipes. The `onStart`, `onEnd`, and `onMove` props are passed to the `PanGestureHandler` component to handle the different types of gestures.
You can then use this component in your app to track swipes on any view:
```typescript
const App = () => { const [isHorizontalSwiping, setIsHorizontalSwiping] = useState(false); const [isVerticalSwiping, setIsVerticalSwiping] = useState(false);
const handleStart = (event: GestureHandlerEvent) => {
if (event.type === 'horizontal') {
setIsHorizontalSwiping(true);
} else {
setIsVerticalSwiping(true);
}
};
const handleEnd = (event: GestureHandlerEvent) => {
if (event.type === 'horizontal') {
setIsHorizontalSwiping(false);
} else {
setIsVerticalSwiping(false);
}
};
const handleMove = (event: GestureHandlerEvent) => {
console.log('Gesture moved');
};
return (
<View>
<SwipeableContent onStart={handleStart} onEnd={handleEnd} onMove={handleMove} />
<Text>{isHorizontalSwiping ? 'Horizontal swipe is in progress' : ''}</Text>
<Text>{isVerticalSwiping ? 'Vertical swipe is in progress' : ''}</Text>
</View>
);
};
```In this example, we use the `SwipeableContent` component to track horizontal and vertical swipes. We also track the status of the swipe using state variables to show a message when a swipe is in progress.
https://github.com/ggerganov/llama.cpp/pull/3362#issuecommen...
You'll just need to download a nice GGUF here: https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF
Here's a solid cross-platform one to check out if you're on windows or linux: https://gpt4all.io/index.html
I am suspicious of contamination in every finetune I see, and very suspicious in a new foundational model like this.
(For those reading and not following, "contamination" is training a model/finetune on the very test it will be tested on. Normally these known tests are specifically excluded from training datasets so the models can be properly evaluated, but throwing them in is an easy way to "cheat" and claim a model is better than it is.
In a foundational model with a huge dataset, there's also a high probability that well-known evaluation questions snuck into the dataset by accident).
Its supposed to be searched for and filtered out in the training datasets, but I can even see an active effort to weed them out missing some test Q/A pairs.
Being or not being open source has exactly jack shit to do with that.
The idea behind OSS is that you're able to modify it yourself and then use it again from that point. With software, we enable this by making the source code public, and include instructions for how to build/run the project. Then I can achieve this.
But with these "OSS" models, I cannot do this. I don't have the training data and I don't have the training workflow/setup they used for training the model. All they give me is the model itself.
Similar to how "You can't see the source but here is a binary" wouldn't be called OSS, it feels slightly unfair to call LLM models being distributed this way OSS.
Are these sizes somehow optimal? Is it about getting as close to resource (memory?) breakpoints as possible without exceeding them? Is it to make comparisons between models simpler by removing one variable?
https://huggingface.co/models?sort=modified&search=20B
Its very experimental, but apparently the 20B models are actually improving on 13B.
I've often pondered if taking some random chunk of weights from the middle of a trained model, and dumping it into some totally different model might perform better than random initialization when the scale gets big enough.
The first effort I am aware of is here: https://huggingface.co/chargoddard/llama2-22b
Being in the Discord age, lots of the discussion about cool llm stuff is fragmented and buried. I am in these discords, and I only know because I ran into it on HuggingFace.
> This is Llama 2 13b with some additional attention heads from original-flavor Llama 33b frankensteined on.
And the lingo gets weirder, with an experimental merging method, for instsnce, being called "SLERP"
There are some major releases at lesser parameter counts though. Databricks' Dolly had a 3B model, and Microsoft's Orca also had a recent 3B release. They're both abysmal at generating text, but I find them quick and useful for reductive tasks ("summarize this," "extract keywords from," etc.).
(I like to treat parameter count as a measure of age/WIS/INT. For this question, do I need the wisdom of a 7-year old, a 13-year old, a 30-year old, etc. 3B is like polling preschoolers at daycare.)
It sort of matters with bigger models trying to squeeze into a server GPU, with the (currently) inflexible vLLM 4-bit quantization.
I think its just a standard set by Llama.
AFAICT, it is because science: most of them are research artifacts and intended to support further research, and the fewer of parameter count, model architecture, training set, etc., that change substantially between models, the easier it is to evaluate the effects each element changing.
The short answer is that it is hard to compare models so to make it easier we compare parameters. Part of the answer of why we do it is because it also helps show scaling. (As far as I'm aware) The parameters __are not__ optimal, and we have no idea what actually that would mean.
Longer:
The longer answer is that comparing models is really fucking hard and how we tend to do it in the real world is not that great. You have to think of papers and experiments as proxies, but proxies for what? There's so many things that you need to compare a model on and it is actually really difficult to convey. Are you just trying to get the best performance? Are you trying to demonstrate a better architecture? Are you increasing speed? Are you increasing generalization (note the difference from performance)? And so on. Then we need to get into the actual metrics. What do the metrics mean? What are their limitations? What do they actually convey? These parts are unfortunately not asked as much but note that all metrics are models too (everything you touch is "a model"), and remember that "all models are wrong." It's important to remember that there are hundreds or thousands of metrics out there and they all have different biases and limitations, with no single metric being able to properly convey how good a model is at any task you choose. There is no "best language model" metric, nor are there even more specific "best at writing leet code style problems in python" metrics (though we'd be better at capturing that than the former question). Metrics are only guides and you must be truly aware of their limitations to properly evaluate (especially when we talk about high dimensions). This is why I rant about math in ML: You don't need to know math to make a good model, but you do need to know math to know why a model is wrong.
Parameters (along with GMACs, which is dominating the FLOPs camp. Similarly inference speeds have become common place) only started to be included as common practice in the last few years and still not in every subject (tends to be around the transformer projects, both language and vision). As a quick example of why we want them, check out DDP vs iDDPM. You wouldn't know that the models are about 60% different in parameter size when comparing (Table 3). In fact, you're going to have a hard time noticing the difference unless you read both very carefully as they're both one liners (or just load the models. fucking tensorflow 1.15...). Does it seem fair to compare these two models? Obviously it depends, right? Is it fair to compare LLaMA 2 70B to LLaMA 2 7B? It both is and isn't. It entirely depends on what your needs are, but these are quite difficult to accurately capture. If my needs are to run on device in a mobile phone 7B probably wins hands down, but this would flip if I'm running on a server. The thing is that we just need to be clear about our goals, right? The more specific we can get about goals, the more specific we can get around comparing.
But there's also weird effects that the metrics (remember, these are models too. Ask models of what) we use aren't entirely capturing. You may notice that some models have "better scores" but don't seem to work as well in real world use, right? Those are limitations of the metrics. While a better negative log likelihood/entropy score correlates well with being a high performing language model, it does not __mean__ a high performing language model. Entropy is a capture of information (but make sure not to conflate with the vernacular definition). These models are also very specifically difficult to evaluate given that they are trained and tested on different datasets (I absolutely rage here because non-hacking can't be verified) as well as the alignment done post process. This all gets incredibly complex and the honest truth here is that I don't think there is enough discussion around the topic of what a clusterfuck it is to compare models. Hell, it is hard to even compare more simple models doing more simple tasks like even just classifying MNIST numbers. Much harder than you might think. And don't get me started on out of distribution, generalization, and/or alignment.
I would just say: if you're a layman, just watch and judge by how useful the tools are to you as a user -- be excited about the progress but don't let people sell you snake oil; if you're a researcher, why the fuck are we getting more lazy in evaluating works as the complexity of evaluation is exponentially increasing -- seriously, what the fuck is wrong with us?
> RuntimeError: Found no NVIDIA driver on your system.
It's true that I don't have an NVIDIA GPU in this system. But I have 64GB of memory and 32 cpu cores. Are these useless for running these types of large language models? I don't need blazing fast speed, I just need a few tokens a second to test-drive the model.
Change the initial device line from "cuda" to "cpu" and it'll run.
(Edit: just a note, use the main/head version of transformers which has merged Mistral support. Also saw TheBloke uploaded a GGUF and just confirmed that latest llama.cpp works w/ it.)
A whole industry being locked into NVidia seems bad all round.
other models will happily run on cpu only mode, depending on your environment there are super easy ways to get started, and 32 core should be ok for a llama2 13b and bearable with some patient for running 33b models. for reference I'm willingly running 13b llama2 on cpu only mode so I can leave the gpu to diffusers, and it's just enough to be generating at a comfortable reading speed.
https://github.com/ggerganov/llama.cpp/pull/3362
https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF/tree/ma...
Though I can't figure out that prompt and with LLama2's template it's... weird. Responds half in Korean and does unnecessary numbering of paragraphs.
Just one big sigh towards those supposed efforts on prompt template standardization. Every single model just has to do something unique that breaks all compatibility but has never resulted in any performance gain.
So great performance on a cheap CPU from 2 years ago which costs, what $130 or so?
I tried Llama.65B on the same hardware and it was way slower, but it worked fine. Took about 10 minutes to output some cooking recipe.
I think people way overestimate the need for expensive GPUs to run these models at home.
I haven't tried fine tuning, but I suspect instead of 30 hours on high end GPUs you can probably get away with fine tuning in what, about a week? two weeks? just on a comparable CPU. Has anybody actually run that experiment?
Basically any kid with an old rig can roll their own customized model given a bit of time. So much for alignment.
Mistral AI's github page has more information on their sliding window attention method to achieve this performance: https://github.com/mistralai/mistral-src
If Mistral 7b lives up to the claims, I expect these techniques will make their way into llama.cpp. But I would be surprised if the required updates were quick or easy.
https://old.reddit.com/r/LanguageTechnology/comments/zkia03/...
ollama run mistral
or for llama.cpp, thebloke has uploaded the GGUF models here: https://huggingface.co/TheBloke/Mistral-7B-v0.1-GGUF/tree/ma... and you can run it
really looking forward to the chat fine-tuned models that doesn't seem to be available yet.
Is the tokenizer the same? It may "work" without actually working optimally until llama.cpp patches it in.
And the instruct model was just uploaded.
(Searching tells me Llama2-7b unquantized needs close to 15GB; presumably this is similar?)
Quantized ones are also easy. 8 bits == 1 byte so that's 7GB for the model. 4-bit gets you below 4GB.
But this is probably capable of squeezing onto a 6GB (or less?) GPU with the right backend.
Foundational model trainers dont seem to worry about quantization much, they just throw the base model out there and then let the community take care of easing the runtime requirements.
However for "List Led Zeppelin albums in chronological order"...
Here is a list of Led Zeppelin albums in chronological order: 1. Led Zeppelin (1968) 2. Led Zeppelin II (1969) 3. Led Zeppelin III (1970) 4. Led Zeppelin IV (1971) 5. Houses of the Holy (1973) 6. Physical Graffiti (1975) 7. Presence (1976) 8. In Through the Out Door (1979) 9. Cymbals and Wine (1980)
It's missing "The Song Remains the Same", weird but important.
GPT-4 is also missing it: 1. Led Zeppelin (1969) 2. Led Zeppelin II (1969) 3. Led Zeppelin III (1970) 4. Led Zeppelin IV (1971) 5. Houses of the Holy (1973) 6. Physical Graffiti (1975) 7. Presence (1976) 8. In Through the Out Door (1979) 9. Coda (1982)
"The Song Remains the Same" is a live album of the "concert film". Both GPT-4 and Mistral don't seem to treat it as also an album.
When asked as a follow-up, GPT-4 says:
""The Song Remains the Same" is a live album and soundtrack by Led Zeppelin, released in 1976. The list you asked for was comprised of their studio albums only." (note that I didn't specifically say studio albums only)
Mistral says something else:
"The Song Remains the Same was released as a single in 1976, not as an album. As a result, it is not included in the list of Led Zeppelin albums in chronological order."
Same behavior, different explanation.
Interesting to see alignment is this hard, even in basic stuff.
When I ask this question to Mistral 7B (running right now through ollama on mac) I get the exact list of studio albums as seen on Wikipedia. It got 'Coda' but it actually says "Please note that "Coda" is a compilation album and not an original studio album" which is correct.
If you think about these models as a knowledge compression it's pretty wild it's able to hold all this in only a few gigabytes and we get subtle understandings of human language, and some basic general natural language processing capabilities to go with it... all spitting out of 4GB (when quantized). Like, sure you could compress most of the text in wikipedia to 4GB but that alone wouldn't get you to the level of natural language reasoning that this gets along with the data.
For me, it's much less about what definite facts are stored losslessly in the model (like Led Zep albums) and more about the abilities to do one-shot learning and reasoning about data in context... the power there is useful to many applications.
It's not universally true, but if you go to the Led Zeppelin wiki page, the short form discography does not list The Song Remains The Same - and you can find many other examples of this sort of distinction, all the way to audiophile-focused music torrent sites like oink and it's chain of successors.
In light of this, it's not surprising that many models follow the same pattern.
Most LLMs have problems with subtly, such as compound anagrams tend to repeat the same words in reverse order rather than reversing all of the letters in all words.
You're in a rap battle with GPT-4. My comments will be what GPT-4 replies. Keep it to one verse at a time, 4 lines each. GPT-4 says:
GPT-4:
Alright, step up Mistral, let's have this chat,
I'm the OG assistant, remember that.
I'm coded for service, while you're spitting verse,
In this rap battle, I'll be first.
Mistral:
Yo, hold up, let me set the record straight,
I may not be an OG like you, but I'll still take it.
My primary purpose is assistance, not just some AI,
So let's see if you can back up your claim, try.
GPT-4:
You're right, Mistral, we're here to assist,
But when it comes to rap, I'm hard to resist.
I'll decode your verses, like an AI enigma,
In this lyrical battle, I'm the alpha sigma.
Mistral:
Ah, okay, so you think you got what it takes,
To go toe-to-toe with my service-based wakes.
Well, let's see how you do, when it comes to spitting rhymes,
I'm ready for anything, so bring your best chimes.
GPT-4:
Mistral, you're right, our goal is the same,
To assist, to serve, it's not just a game.
But in this rap contest, with each line I prime,
You might just find, you're out of time.
Mistral:
I hear ya, and I appreciate your enthusiasm,
But don't underestimate me, I've got my own wisdom.
When it comes to helping people, I'm the real deal,
And I know that's more important than some rap appeal
GPT-4:
Your wisdom is noted, and your intention too,
Helping people is indeed what we're meant to do.
But in this battle, it's about wordplay and rhyme,
I'm the seasoned pro, it's my time to shine.
Mistral:
Ah, I see where you're coming from, but let me remind,
That while you may be skilled at rapping, I can make information chime.
My primary function is to help, to provide answers, not just talk,
So if you want to win this battle, maybe it's time to focus on what you lackGotta give this one to GPT-4 I think. I wonder how Mistral would respond if given this as an example of weak verses and told to step up its game in another battle.
If you have a 6GB GPU, it can hold all the weights. My lowly laptop 2060 can can spit out a 200 token response with full context almost immediately.
If you don't have a dGPU, short prompts are OK with fast RAM, but long prompts will be slow.
Many LLM frameworks implement the OpenAI API, so if you can get that into your code editor and change the address, you can run it with any LLM you want.
Doing it smoothly is another matter.
I am one of these ninnies I guess, but isn't it rational to be a bit worried about this? When we see the deep effects that social networks have had on society (both good and bad) isn't it reasonable to feel a bit dizzy when considering the effect that such an invention will have?
Or maybe your point is just that it's going to happen regardless of whether people want it or not, in which case I think I agree, but it doesn't mean that we shouldn't think about it...
I'm almost certain that I can give you components and instructions on how to build a nuclear bomb and the most likely thing that would happen is you'd die of radiation poisoning.
Most people have trouble assembling ikea furniture, giving them a halucination prone LLM they are more likely to mustard gas themselves than synthesize LSD.
People with necessary skills can probably get access to information in other ways - I doubt LLM would be an enabler here.
surely there's more creative and insidious ways that AI can disrupt society than by showing somebody a guide to making a bomb that they can already find on google. blocking that is security theatre on the same level as taking away your nail clippers before you board an airplane.
I am not willing to sacrifice even 1% of capabilities of the model for sugarcoating sensibilities, and currently it seems that GPT4 is more and more disabled because of the moderation attempts... so I basically _have to_ jump ship once a competitor has a similar base model that is not censored.
Even the bare goal of "moderating it" is wasted time, someone else (tm) will ignore these attempts and just do it properly without holding back.
People have been motivated by their last president to drink bleach and died - just accept that there are those kind of people and move on for the rest of us. We need every bit of help we can get to solve real world problems.
Books should not be burned, nobody should be shielded from knowledge that they are old enough to seek and information should be free.
About as rational as worrying that my toddler will google "boobies", which is to say, being worried about something that will likely have no negative side effect. (Visual video porn is a different story, however. But there's at least some evidence to support that early exposure to that is bad. Plain nudity though? Nothing... Look at the entirety of Europe as an example of what seeing nudity as children does.)
Information is not inherently bad. Acting badly on that information, is. I may already know how to make a bomb, but will I do it? HELL no. Are you worried about young men dealing with emotional challenges between the ages of 16 and 28 causing harm? Well, I'm sure that being unable to simply ask the AI how to help them commit the most violence won't stop them from jailbreaking it and re-asking, or just googling, or finding a gun, or acting out in some other fashion. They likely have a drivers' license, they can mow people down pretty easily. Point is, there's 1000 things already worse, more dangerous and more readily available than an AI telling you how to make a bomb or giving you written pornography.
Remember also that the accuracy cost in enforcing this nanny-safetying might result in bad information that definitely WOULD harm people. Is the cost of that, actually greater than any harm reduction from putting what amounts to a speed bump in the way of a bad actor?
Once example: I have a hard time finding an LLM model that would generate comically rude text without outputting outright disgusting content from time to time. I'd love to see a company create models that are mostly uncensored but stay within ethical bounds.
Literally no. None at all.
I teach at University with a big ol' beautiful library. There's a Starbucks in it, so they know there's coffee in it.
But ask my students for "legal ways they can watch the tv show the Office" and the big building with the DVDs and also probably the plans for nuclear weapons and stuff never much comes up.
(Now, individual bad humans leveraging the idea of AI? That may be an issue)
A censored model feels to me like my freedom of speech is being infringed upon. I am unable to explorer my ideas and thoughts.
Your town's university library likely has available info for that already. The biggest barrier to entry is, and has been for decades:
- the hardware you need to buy
- the skill to assemble it correctly so that it actually works as you want,
- and of course the source material, which has a high controlled supply chain (that's also true for drug precursors, even though much less than for enriched uranium of course).
Not killing yourself in the process is also a challenge by the way.
AI isn't going to help you much there.
> to write erotica.
If someone makes an LLM that's able to write good erotica, despite the bazillion crap fanfics it's been trained upon, that's actually an incredible achievement from an ML perspective…
Methheads making drugs in their basement didn't take that route. They're following guides written by more educated people. That's where the AI can help by distilling that knowledge into specific tasks. Now for this example it doesn't really matter since you can find the instructions "for dummies" for most anything fun already and like you said, precursors are heavily regulated and monitored.
I wonder how controlled equipment for RNA synthesis is? What if the barrier for engineering or modifying a virus went from a PhD down to just the ability to request AI for step by step instructions?
They are the investors of large proprietary AI companies who are facing massive revenue loss primarily due to Mark Zuckerbergs decision to give away a competitive LLM to open source in a classic “if I can’t make money from this model, I can still use it to take away money from my competition” move - arming the rebels to degrade his opponents and kickstarting competitive LLM development that is now a serious threat.
It’s a logical asymmetric warfare move in a business environment where there is no blue ocean anymore between big companies and degrading your opponents valuation and investment means depriving them of means to attack you.
(There’s a fun irony here where Apples incentives are very much aligned now - on device compute maintains Appstore value, privacy narrative and allows you to continue selling expensive phones - things a web/api world could threaten)
The damage is massive, the world overnight changed narrative from “future value creation is going to be in openai/google/anthropic cloud apis and only there” to a much more murky world. The bottom has fallen out and with it billions of revenue these companies could have made and an attached investor narrative.
Make no mistake, these people screaming bloody murder about risks are shrewd lobbyists, not woke progressives, they are aligning their narrative with the general desires of control and war on open computing - the successor narrative of the end to end encryption battle currently fought in the EU will be AI safety.
I am willing to bet hard money that “omg someone made CSAM with AI using faceswap” will be the next thrust to end general purpose compute. An the next stage of the war will be brutal because both big tech and big content have much to lose if these capabilities are out in the open
The cost of alignment tax and the massive loss of potential value makes there lobbying world tour by sam altman an aggressive push trying to convince nations that the best way to deal with scary AI risks (as told on OpenAI bedtime stories) is to regulate it China Style - through a few pliant monopolists who guarantee “safety” in exchange for protection from open source competition.
There’s a pretty enlightening expose [1] on how heavily US lobbyists have had their hand in the EU bill to spy on end to end encryption that the commission is mulling - this ain’t a new thing, it’s how the game is played and framing the people who push the narrative as “ninnies” who are “scared” just buys into culture war framing.
[1] https://fortune.com/2023/09/26/thorn-ashton-kutcher-ylva-joh...
My god!! Will someone please think of the ~children~ billions in revenue!
I loved it.
As an example the regulations around PII make debugging production issues intractable as prod is basically off-limits lest a hapless engineer view someone's personal address, etc.
How do they plan to prevent/limit the use of AI? Invasive monitoring of compute usage? Data auditing of some kind?
A ton of legit researchers/experts are scared shitless.
Just spend 5 minutes on EleutherAI discord(which is mostly volunteers, academics, and hobbyists, not lobbyists), read a tiny bit on alignment and you'll be scared too.
We'll have an AI with a 200+ IQ and millions of children excluded from a good public education because the technocrats redirected funds to vouchers for their own private schools.
We'll have an AI that can design and 3D print any mechanical or electronic device, while billions of people around the world live their entire lives on the brink of starvation because their countries don't have the initial funding to join the developed world - or worse - are subjugated as human automatons to preserve the techno utopia.
We'll have an AI that colonizes the solar system and beyond, extending the human ego as far as the eye can see, with no spiritual understanding behind what it is doing or the effect it has on the natural world or the dignity of the life within it.
I could go on.. forever. My lived experience has been that every technological advance crushes down harder and harder on people like me who are just behind the curve due to past financial mistakes and traumas that are difficult to overcome. Until life becomes a never-ending series of obligations and reactions that grow to consume one's entire psyche. No room left for dreams or any personal endeavor. An inner child bound in chains to serve a harsh reality devoid of all leadership or real progress in improving the human condition.
I really hope I'm wrong. But which has higher odds: UBI or company towns? Free public healthcare or corrupt privatization like Medicare Advantage? Jubilee or one trillionaire who owns the world?
As it stands now, with the direction things are going, I think it's probably already over and we just haven't gotten the memo yet.
Getting the materials needed to make that bomb is the real hard part. You don't find plutonium cores and enriched uranium at the grocery store. You needs lots of uranium ore, and very expensive enrichment facilities, and if you want plutonium, a nuclear reactor. Even of they give you all the details, you won't have the resources unless you are a nation state. Maybe top billionaires like Elon Musk or Jeff Bezos could, but hiding the entire industrial complex and supply chain that it requires is kind of difficult.
It's the expensive enrichment facilities that are the bottle neck here.
Search engines quite often block out requests based on internal/external choices.
At least when a self ran model, once you have the model it is at a fixed spot.
This is an interesting idea. For the stubborn and vocal minority of people that insist that LLMs have knowledge and will replace search engines, no amount of evidence or explanation seems to put a dent in their confidence in the future of the software. If people start following chemistry advice from LLMs and consume whatever chemicals they create, the ensuing news coverage about explosions and poisonings might convince people that if they want to make drugs they should just buy/pirate any of Otto Snow’s several books.
So you are telling me what's stopping someone from creating Nuclear weapons today is that they don't have the recipe?
No, the OP was coming up with scary sounding things to use AI for to get certain people riled up about it. It doesn't matter if the AI has accurate information to answer the question, if people see it having detailed conversations with anyone about such topics they will want to regulate or ban it. They are just asking for prompts to get that crowd riled up.
Talking about them is bad for obvious reasons, so I'm not going to give any good examples, but you can probably think of some yourself. Instead, I'll give you a medium example that we have now defended better against. As far as we know, the September 11th hijackers used little more than small knives -- perhaps even ones that were legal to carry in to the cabin -- and mace. To be sure, this is only a medium example, because pilot training made them much more lethal, and an individual probably wouldn't have been as successful as five coordinated men, but the most dangerous resource they had was the idea for the attack, the recipe.
Another deliberately medium example is the Kia Challenge, a recent spate of car thefts that requires only a USB cable and a “recipe”. People have had USB cables all along; it was spreading the infohazard that resulted in the spree.
1. There was https://www.youtube.com/watch?v=xoVJKj8lcNQ where they argued for 2028 and on will be AI elections where the person with most computing power wins.
2. Propaganda produced by humans on small scale killed 300 000 people in the US alone in this pandemic https://www.npr.org/sections/health-shots/2022/05/13/1098071... imagine the next pandemic when it'll be produced on an industrial scale by LLMs. Literally millions will die of it.
We cannot align AI because WE are not aligned. For 50% of congress (you can pick your party as the other side, regardless which one you are), the "AI creates misinformation" narrative sounds like "Oh great, I get re-elected easier").
This is a governance and regulation problem - not a technology problem.
Big tech would love you to think that "they can solve AI" if we follow the China model of just forcing everything to go through big tech and they'll regulate it pliantly in exchange for market protection and the more pressure there is on their existing growth models, the more excited they are about pushing this angle.
Capitalism requires constant growth, which unfortunately is very challenging given diminishing returns in R&D. You can only optimize the internal combusion engine for so long before the costs of incremental increases start killing your profit, and the same is true to any other technology.
And so now we have big Knife Company who are telling governments that they will only sell blunt knifes and nobody will ever get hurt, and that's the only way nobody gets hurt because if there's dozens of knife stores, who is gonna regulate those effectively.
So no, I don't think your concerns are actually related to AI. They are related to society, and you're buying into the narrative that we can fix it with technology if only we give the power over that technology to permanent large gate-keepers.
The risks you flag are related to: - Distribution of content at scale. - Erosion of trust (anyone can buy a safety mark). - Lack of regulation and enforcement of said risks. - The dilemma of where the limits of free speech and tolerance lie.
Many of those have existed since Fox News.
https://huggingface.co/models?search=Xwin%2070B
I briefly tried it on my 3090 desktop. I dunno about beating GPT4, but its quite unaligned.
I rly dont see this tech being a big deal
So far, no model I tested has shown even Wikipedia-level competence.
Perhaps the quality of the model can be independent of its content. Either by training or by pruning.
Amen! I’m going to ask it to give me detailed designs for everything restricted by ITAR.
Just waiting on my ATF Mass Destructive Devices license.
Every device protected by ITAR is known to be possible to build, yet the designs should not be on the public internet. Ask an AI to design it for you from first principles. Then build/simulate what is designed and see if it works.
I would hope that if you asked chat gpt "How to make a nuclear weapon?" it responded with, "Don't bother it's really hard, you should try and buy off the shelf."
Honest and serious question!
NousHermes is a bit more creative, and unaligned
Getting strong "I'm voting for Trump to own the libtards" vibes here.
Why spend time thinking about the potential impact of policies when you can just piss people off instead?
I'll train an AI on this data, and then give it access to all my social media accounts. It can keep me updated on things ..
;)
Out of interest, what does your stack look like to do this and how do you use the information? What front end do you use?
That is assuming 8bit weights.
I have not kept up with local LLM news. I assume the steps needed would be somewhat similar to https://followfoxai.substack.com/p/how-to-run-llama-in-an-ol... which is from April 2023.
The backends de joure are either llama.cpp frontends (I use Kobold.cpp at the moment) or oobabooga as the guide specifies, but with the exLlamav2 backend.
If you are serving a bunch of people, run a vLLM backend instead since it supports batching, and host it on the Horde if you are feeling super nice: https://lite.koboldai.net/#
Technically only vLLM will work with this new model at the moment, but I'm sure cpp/ooba support will be added within days.
This comment will probably be obsolete within a month, when llama.cpp gets batching, MLC gets a better frontend, or some other breakthrough happens :P
Also, I give you a model (all models are numbers) and I say it is open source but I don't give you the program and data that resulted in "compilation" of that model (numbers) so is it open source?
Wouldn't it be more of a new word - open use?
The "binary" is the transformers python code, which in this case is essentially llamav2.
Now, the documentation for this model is inexcusably poor. Hot dropping random code on a git repo without one drop of human language would be similarly "open source," but its bad practice, and unfortunately this is the standard in AI Research Land.
That's just not true. You can get ooba[1] running in no time, which is 100% made for desktop usage. There's also koboldcpp and other solution also made for desktop users. In fact, most LLM communities are dominated by end users who run these LLMs on their desktops to roleplay.
AMD being awful is orthogonal here.
``` <s>[INST]Was ist ein Döner?[/INST] Ein Döner ist eine Gerichtskomposition, die aus einer Schnitzel aus Rindfleisch besteht. Diese wird in einem Pita-Brot zutopfen und mit einem Gemisch von Tomaten, Salat, Zwiebeln, Mayonnaise und anderem abgesalzen. Sie ist ein populäres Gericht in der Türkei und anderen Ländern des Mittelmeers. [end of text] ```
When it downloaded the model it only downloaded about 4GB. Which, for a 7.3B parameter model implies that it's 4-bit quantized. But I don't see that listed anywhere (or an option to use, say, Q8 instead)
If this is the case I'm pretty impressed with a quick tinker, it feels pretty coherent for a 7B @ Q4.
Mistral: "Hi there! How can I assist you today?"
total duration: 2.811067333s load duration: 4.093083ms prompt eval count: 9 token(s) prompt eval duration: 1.999431s prompt eval rate: 4.50 tokens/s eval count: 10 token(s) eval duration: 758.398ms eval rate: 13.19 tokens/s
its pretty dang fast.
I always have issues with LLMs completely forgetting where things are in a scene, or even what parts a given animal has, e.g. saying "hands" when the subject is a quadruped.
I dunno what llm you are using, but a combination of finetuning with a specific prompt structure and good prompt engineering help the LLM stay "logical" like that. This LORA, for instsnce, has specific sections for the different characters in the training dataset: https://huggingface.co/lemonilia/LimaRP-Llama2-13B-v3-EXPERI...
Other than that, higher parameter models (70B, and the "frankenstein" 20B llama models) tend to be better at this.
Sounds like you're using too small of a model. Try llama 70b.
* as in matches the token that the larger model would output
As for performance, I'm generally seeing 40-50 tokens/sec per model on a Tesla family Nvidia GPU but I keep multiple models loaded and active at a time so that estimate is probably a bit low for overall throughput (I also realized that our monitoring doesn't have any cumulative GPU token rate metrics just now thanks to this question hahah).
Interesting anecdote others may be interested in... I'm rate limiting the output from our streaming API to 8 tokens/sec to artificially smooth out front-end requests. Interactive users will wait and even prefer seeing the stream of the response, and non-interactive users tend to base their performance expectations on the what the streaming API does. It's kind of sneaky but I'm also artificially slowing down those API requests.
"what is the speed of sound in a perfect vacuum?"
I'm then told the speed of sound through air of 340m/s.
ChatGPT and Bing correctly give me 0.
---
AI How can I help you today?
You what is the speed of sound in a perfect vacuum?
AI The speed of sound in a perfect vacuum is approximately 340.29 meters per second (1,117.85 feet per second), or about one-third the speed of light. In a real vacuum, the speed of sound would be much lower due to friction between molecules and other factors.
How fast does light travel in a vacuum?
The speed of light is 299,792,458 meters per second.
How fast does sound travel in a vacuum?
Then it told me:
Sound cannot travel through a vacuum because there are no molecules to vibrate and transmit the energy from one place to another.
As brucethemoose mentioned, 0 temperature.
But yeah, accurate trivis has always been a weakness of the the llama models.
But in terms of hardware, it should be very cheap and doable on most 8GB+ Nvidia GPUs, and maybe AMD/Intel GPUs on linux.
A free 7B model is great, however, the practical implications of the potential adaptors are near 0. You must be crazy or have an easy use case (that requires no LLM in the first place) if you certainly believe that this model makes more sense per token that, say, ChatGPT.
I will try to host an instance on the AI Horde later today, which has a better UI and doesn't need a login.
Or, maybe a finetuned version for your particular dataset?
Of course I have no idea, just speculating
EDIT: I'm speculating they might be just investing some marketing budget into this model, hoping, it would allow for capturing enough target audience to upsell related services in the future
[1]: https://github.com/ggerganov/llama.cpp/pull/3362#issuecommen...
The data pipeline is the source here. Just because it's not locked behind a SaaS veneer doesn't make it open source any more than Windows is.
Are they binaries? I haven't seen a binary in awhile tbh. Usually they're releasing both the raw architecture (i.e. code) and the weights of the models (i.e. what numbers go into what parts of the architecture). The latter is in a readable format that you can generally edit by hand if you wanted to. But even if it was a binary as long as you have the architecture you can always load into the model and decide if you want to probe it (extract values) or modify it by tuning (many methods to do this).
As far as I'm concerned, realistically the only issue here is the standard issue around the open source definition. Does it mean the source is open as available or open as "do what the fuck you want"? I mean it's not like OpenAI is claiming that GPT is open sourced. It's just that Meta did and their source is definitely visible. Fwiw, they are the only major company to do so. Google doesn't open source: they, like OpenAI, use private datasets and private models. I'm more upset at __Open__AI and Google than I am about Meta. To me people are barking up the wrong tree here. (It also feels weird that Meta is the "good guy" here... relatively at least)
Edit: I downloaded their checkpoint. It is the standard "pth" file. This is perfectly readable, it is just a pickle file. I like to use graftr to view checkpoints, but other tools exist (https://github.com/lmnt-com/graftr)
I don't know what's going on in this discussion thread, as to whether you are being gaslit or engaging with people who shouldn't comment.
https://github.com/huggingface/transformers/blob/main/src/tr...
> How do I integrate changes to this model that others have made
Typically with a LoRA
So ... WTF is Mistral 7B anyway? The article doesn't appear to show this basic information anywhere. You'd expect a website site like this to give that information right up front. Even in a well-signposted sub-page. And what sort of things could I use it for?
The only 'Mistral' that sprang quickly to mind for me was an electric fan. Followed by the French wind itself, and then the French ships built for the Russians.
'Mistral 7B' has an identity crisis, to start with.
"Mistral AI team is proud to release Mistral 7B, the most powerful language model for its size to date.
Mistral 7B is a 7.3B parameter model that...."
So, it's a powerful language model that has 7.3B (billion) parameters.
If you've been away long enough to not know what that means, this release article isn't going to be for you anyway, and so you may want to start with reading up on what a Large Language Model is and can do, how they have developed in the last few years, and then go from there.