2023: The Year of AI
journal.everypixel.com
journal.everypixel.com
wget https://huggingface.co/TheBloke/deepseek-coder-6.7B-instruct-GGUF/resolve/main/deepseek-coder-6.7b-instruct.Q5_K_M.gguf
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make
./main -ngl 32 -m ../deepseek-coder-6.7b-instruct.Q5_K_M.gguf --color -c 2048 --temp 0.7 --repeat_penalty 1.1 -n -1 -i -ins
Haven't people realized what they can run themself?> Just like this:
Just is doing a lot of heavy lifting there.
Is this the first model you came across?
We're there any dependencies you had to install?
Did you have to check video card compatibility?
Where did you get the command arguments. It looks awfully complicated for a "just", as if there's a lot of options that aren't straightforward to use.
Etc
I have no clue about anything LLM related. I just made it run after reading some comment on HN pointing in its direction.
My point is that these locally run LLMs seems way "underreported". I even tried to make a "Show HN" post about it but it got zero interest.
But maybe I am missing something?
This was a few months ago and things may be better now.
You can also use ollama which makes it into a one line install and then a simple command to run any popular model.
no command line involved, has a huggingface browser built in to load in the model of the day, has a chat-like interface for chatGPT like use, and can create a local server to run your local model for your programs to interact with it using an API that is identical to OpenAI's
Meanwhile amateurs and uni researchers post really interesting work on a site named after a hugging emoji ...
Lots, lots of people do not care.
Should limit the stock price prospects though.
Where it stops being comparable is general application across literally millions of use cases. The ChatGPT system has proven itself a valuable utility across industries, people, and use cases. No open model I know of can match it yet.
There are oodles of use-cases where sending your data to an outside provider is a complete no-go. In these cases OpenAI/Google/whoever-products aren't relevant competition.
Have you ever worked in any sector that has security policies?
Even if you haven't, perhaps spend 2 minutes using a search engine?
Here is a first page result for you: https://www.tomshardware.com/news/samsung-fab-workers-leak-c...
Your own link shows Samsung using ChatGPT. Not sure what point you're trying to make with it.
"These actions clearly put confidential information at risk, prompting Samsung to warn its employees about the dangers of using ChatGPT. Samsung Electronics informed its executives and employees that data entered into ChatGPT is transmitted and stored on external servers, making it impossible for the company to retrieve it and increasing risks of confidential information leakage."
Obnoxious comment considering your link shows the opposite of what you're claiming.
Opposite of what I claim? No.
"Samsung Electronics informed its executives and employees that data entered into ChatGPT is transmitted and stored on external servers, making it impossible for the company to retrieve it and increasing risks of confidential information leakage."
And yes, there are indeed use cases where sending data to an outside provider is a no-go. The bet OpenAI is making is that they can solve for that later while building their business on use cases where it's fine to send data to an outside parameter. It may also simply not be something they care about. In my own work I know of a massive financial enterprise that has prioritized ~30 or so features where it's fine to send that data. OpenAI is not struggling to get their money.
It remains to be seen if OpenAI will also capture this market, or if fine-tuning open models to be "good enough" wins out over time. The point isn't that, though. The point is that their models are so broadly applicable that _anyone_ can get some value quickly without much work.
NVIDIA GeForce GTX 1050 Ti
Intel® Core™ i5-8300H × 8
32,0 GiB ram
So I guess your computer would be a faster typist?Nowadays we just run full on opaque models directly from huggingface without thinking twice about reading anything. Interesting how times change. I wonder what supply chain attacks will come from huggingface, must not be long now.
I’m reminded of the classic Dropbox HN comment (https://news.ycombinator.com/item?id=8863): why use Dropbox when any Linux user can just use curlftpfs?
Contrast this to a few years ago we'd talk about Resnet or Retina-net or whatever, not so much about the facebook pertain on Image-net when describing the model.
In 2018 I remember hearing "architecture is the new feature engineering" (mostly meaning over-fitting I undetstood). Now it's all (mostly) about the dataset and training, the architecture, in 2023, was a minor detail. I personally think architecture will make a comeback soon.
Annnd nowadays tools like transformers can automatically select the architecture based on the name of the weights, so when interacting with LLMs, you generally just name the weights.
Also MANY models are just using the llama or llama2 architecture.
What's that supposed to mean? Weights are not independent of the underlying architecture.
There isn’t one that’s more important than the other. There are no weights at all without an architecture. And any architecture improvements that pay off have a multiplier effect on efficiency.
Nobody on Earth had the resources to get the performance they are getting on these big models with just a default fully connected layer architecture.
The name is a little confusing, since the package supports non transformers models as well (like RWKV)
I think HN took out the huggingface emoji from my post
I think people talk about the models because the architectural details are generally inaccessible for us since we don't have sufficient training (which it requires quite a lot to have an intuitive understanding I believe). Whereas the models are easily downloaded and tested.
I would definitely include Mamba state space models and of course would prefer a technical review over a corporate review of the year.
[1]: Retentive Network: A Successor to Transformer for Large Language Models: https://arxiv.org/abs/2307.08621
I also used StableDiffusion to mock-up landscaping designs for our back patio and make ceramic art ideas for my wife last month.
I was finally able to use super resolution to improve the quality of an old video CD of a VHS tape of a theater play my dad starred in with his friends in the 80s. I tried without super resolution many times in the past but it was never good enough.
You are right that it will take a few years before it hits but AI in 2023 is not hype if you know which tool to use when.
When you as an individual use it to sort PDFs or something it’s really hype.
In fact, I would say thinking that AI is going to solve 'real' issues like that, is more 'hype' than the above commenter.
Personally speaking, anything that eliminates and reduces mundane but complex work is high magic. The real deal.
If “hype” pays for itself, it isn’t hype. It just became a necessity.
Just thinking off the top of my head, Segment Anything, Llama 1 and 2, Mistral, Stable diffusion XL, ControlNet, Whisper are all open source AI releases this year.
The big thing was huge progress in natural language understanding. 2023 was the year the Turing test was smashed.
Seeing computers win games, drive cars, optimize systems, even design things wasn’t as subjectively impressive to most of us as being able to talk to them. This was the year we first saw AI that could sort of communicate with us the way we can with each other.
- Acting as a lossy text decompressor (harmful)
- Acting as a lossy text compressor (mostly trying and failing to undo the damage from above)
- Acting as an outright bullshit generator (harmful)
- Acting as a poor substitute for a search engine with a propensity for trying to bullshit you (harmful)
- Acting as a poor substitute for a parser (which, even if it worked would be dumb because it doesn't understand structured output, so now you have two parsing problems)
- Generating broken and/or extremely poorly factored code (harmful)
- Hiding plagiarism (harmful)
GPT (and the AI hype in general) is the wrong solution to the wrong problem. Look at the effort that Google wasted on Duplex instead of OpenTable.
One problem with the Turing Test we might have predicted: if people have context on how models have improved, they narrow their expectations to match.
But I suppose the Strong Turing Test is when a model is so much smarter than us, that it can convincingly hide its differences and superiority, even when we have good reason to suspect it’s an AI.
Very curious about the type of projects you've delivered.
I'm a sample of 1 and also relatively inexperienced. But I felt I quickly reached the limits of what was possible when I tried doing sentence classification with OSS sentence embedding models. The issue was with negation. I'd attributed too much magic to embedding models - they don't really understand language.
Not to say there isn't very capable tech out there. Just to add a datapoint that "sentiment analysis"-like approaches in blogs don't always scale to your particular use-case.
Edit: conscious I've drifted from the topic of chatbot type models, but felt relevant somehow.
I have come to the conclusion that as much as possible I should try to wean myself off of OpenAI immediately. Because it's just not necessary or desirable to be tied to a single vendor anymore for many tasks. And in 2024 the open source capabilities will continue to increase. Soon everyone with a relatively new computer will be running things like LLMs locally. Within a couple of years it will be integrated into every OS or browser.
Only in benchmarks as GP said. All the models I have tried either hallucinates like crazy or refuse to answer everything. Nowhere close to GPT 3.5.