HNHacker News
TopNewBestAskShowJobs

underlines

786 karma · joined August 22, 2012

submissionscomments
underlines··on CodeAlpaca – Instruction following code generation model
"Model weights aren't part of the release for now, to respect OpenAI TOS and LLaMA license."

I feel like the whole Open Source ML scene is slowed down by a strong chilling effect. Everyone seems to be afraid to release models.

Meanwhile, other models are freely available up to alpaca 30b:

https://github.com/underlines/awesome-marketing-datascience/...

underlines··on FauxPilot – An open-source GitHub Copilot server
4bit GPTQ maybe?
underlines··on Show HN: Finetune LLaMA-7B on commodity GPUs using your own text
to my understanding there are 4 levels to add information:

1. train a model

2. fine tune a model

3. create embeddings for a model

4. use few shot prompt examples at inference time

These have decreasing resource need, but also decreasing quality.

For example, the GPT-3 API (not yet the GPT-4 API) has a functionality to send it your own embeddings, for example of your own source code documentation. Then you can query GPT-3 and it "knows" your source code doc and answers specifically with that in mind.

underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
is it possible to use your web gui locally by ignoring baseten.login("PASTE_API_KEY_HERE") ?
underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
check my summary of resources:

https://github.com/underlines/awesome-marketing-datascience/...

underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
you just explained how human memory works and i thought about implementing that in a future model that allows for more max input tokens. the further back the text, the more it goes through a "summarize this text: ..." prompt. GPT4 has 28k Token limit, so it has the brain of a cat maybe, but future models will have more max. tokens and might be able to have a human like memory that gets worse the older the memory is.

Alternatives are maybe architectures using langchain or toolformer to retrieve "memories" from a database by smart fuzzy search. But that's worse, because reasoning would only be done on that context, instead of all memories it ever had.

underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
most implementations do, like https://github.com/oobabooga/text-generation-webui

this might be a hallucinated answer, due to the very small model size of 7b. try the 13b-4bit, it's much better!

underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
or the other one here https://github.com/juncongmoo/chatllama
underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
to my understanding, fine tuning is slow and would be quite bad to update. embeddings seems to be the way to go. i don't understand it well enough, but it seems with the langchain framework you can create an embedding of your own data and submit it to the GPT API and i believe emeddings should be a similar principle in llama. at least i did it with diffusers in stablediffusion.
underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
just run the 13b model 4bit quantized locally, it's already better than the 7b-8bit and you can turn down the temperature to 0 to get repeatable results.
underlines··on Show HN: ChatLLaMA – A ChatGPT style chatbot for Facebook's LLaMA
did you use the cleaned and improved alpaca dataset from https://github.com/tloen/alpaca-lora/issues/28 ?
underlines··on Show HN: GPT Repo Loader – load entire code repos into GPT prompts
Im no expert, but wouldn't it make more sense to give the repo-context (structure, source code, PRs, Issues, ...) as embeddings? You could use langchain to generate an embedding and send it through the API, like explained here [1]. It then should have access to the context at inference time, which as I understand is better than loading the context in the prompt which wastes tokens / max. output length and has a limit.

1 https://www.youtube.com/watch?v=veV2I-NEjaM

underlines··on UBS agrees to buy Credit Suisse
Livestream in English/German with Swiss Govt., Finma, National Bank, UBS Group and CreditSuisse right now:

https://www.20min.ch/story/jetzt-informiert-der-bundesrat-ue...

underlines··on UBS agrees to buy Credit Suisse for over $2B
Livestream of the swiss govt.: https://www.20min.ch/story/jetzt-informiert-der-bundesrat-ue...
underlines··on Alpaca: A strong open-source instruction-following model
See: https://arxiv.org/abs/2210.17323

Q: Doesn't 4bit have worsen output performance than 8bit or 16bit? A: GPTQ doesn't quantize linearly. While RTN 8bit does reduce output quality, GPTQ 4bit has effectively little output quality loss compared to baseline uncompressed fp16.

https://i.imgur.com/xmaNNDd.png https://i.imgur.com/xmaNNDd.png

underlines··on Running LLaMA 7B on a 64GB M2 MacBook Pro with Llama.cpp
I would suggest helping open-assistant.io intsead of wasting resources for LLaMA, which has very restrictive ToS and thus cripples it's full potential for the open source community. open-assistant is in the process of creating a crowd based fine tuning data set for instructions/assistant, and then select a truly open source model, not LLaMA with it's restrictive ToS.
underlines··on Dalai: Automatically install, run, and play with LLaMA on your computer
- does it support bitsandbytes?

- does it support GPTQ 4 bit quantization?

so far I like the feature set of github/text-generation-webui

underlines··on Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support
cream stays far far away from carbonara. otherwise it's not carbonara.

carbonara sauce is simply pecorino or parmigiano cheese mixed with eggs or just yolks and pepper and guanciale or pancetta.

NO CREAM, NO MILK, NO HAM, NO BACON! basta! /endofrant

underlines··on Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support
Relevant: Since LLaMA leaked on torrent, it has been converted to Huggingface weights and it has been quantisized to 8bit for less vram requirements.

A few days ago it has also been quantisized to 4bit and 3bit is coming. The quantization method they use is from the GPTQ paper ( https://arxiv.org/abs/2210.17323 ) which leads to almost no quality degradation compared to the 16bit weights.

4 bit weights:

Model, weight size, vram req.

LLaMA-7B, 3.5GB, 6GB

LLaMA-13B, 6.5GB, 10GB

LLaMA-30B, 15.8GB, 20GB

LLaMA-65B, 31.2GB, 40GB

Here is a good overall guide for Linux and Windows:

https://rentry.org/llama-tard-v2#bonus-4-4bit-llama-basic-se...

I also wrote a guide how to get the bitsandbytes library working on windows:

https://github.com/oobabooga/text-generation-webui/issues/14...

underlines··on Show HN: Watch ChatGPT debate itself on a given topic
I asked chatGPT to code a 2 speaker debate function using the unofficial API from revChatGPT ( https://github.com/acheong08/ChatGPT ) Input arguments are initial_topic and max_conversation_length. I asked it to create a loop which controls this. After some back and fourth fixing some logic by instructing chatGPT to change things, it finally works, but somehow times out after 4 times going back and fourth (i'm a plus subscriber, but not using the official API)
underlines··on Facebook LLAMA is being openly distributed via torrents
- how much vRAM needed to run each model parameter size?

- any inference optimization we can use similar to StableDiffusion, to bring down the vRAM requirements?

I only know about these:

- use 8bit precision

- https://github.com/bigscience-workshop/petals

- https://github.com/FMInference/FlexGen

- https://github.com/microsoft/DeepSpeed

Anything that could bring this to a 10GB 3080 or 24GB 3090 without 60s/it per token?

underlines··on EleutherAI announces it has become a non-profit
I was looking for open alternatives to self hosted, or crowd hosted finetuned LLMs like ChatGPT and found LAION Open Assistant. Then found resources to further optimize inference as well as training:

- Open source fine tuned assistants like LAION Open-Assistant [1]

- inference optimizations like VoltaML, FlexGenm Distributed Inference [2]

- training optimizations like Hivemind [2]

1 https://github.com/LAION-AI/Open-Assistant

2 https://github.com/underlines/awesome-marketing-datascience/...

underlines··on Show HN: AskHN
Just to be sure: This is NOT a finetuned GTP model, but rather standard GPT-3 API, used to summarize search results of a HN Comments DB, based on user input. Right?
underlines··on Bard and new AI features in Search
Is there a way to become a tester?
underlines··on Playing Zork with AI-generated imagery [video]
Do you agree that one can train a human brain or a digital brain on as much copyrighted artist work as one wants? style is not copyrightable, so is training on copyrighted content.

the law (an common sense) implies that as long as you're not reproducing art, you can do this. and latent space doesn't save or reproduce that content at all. It mimics, combines remixes or reinvents it. The same what a human brain is doing after learning how to draw.

humanity wouldn't be where it is now when it's forbidden to learn from inventions.

underlines··on Megaface
Challenge this:

If I am legally allowed to look at a million pictures to learn how to draw in different art styles, or even how to imitate a very certain art style, I basically train weights and biases in my brain.

If StabilityAi does the same with pictures available online, and release the weights and biases as Stable Diffusion, how is this different from humans learning from that data?

underlines··on Ask HN: Are there any good open source text-to-speech tools?
ChatGPT is so crazy it even works in fluent Thai. That's better than any machine translation I've ever tried so far. It even takes cultural differences into account. For example when you ask it to translate "I love you" into Thai, it mentions, that normally you would not say this in the same circumstances as you would say it to your lover in the West, correctly explaining in what circumstances people would really use it, and what to use instead. That's revolutionary for minority languages without a lot of learning material available online.

Also I am a native Swiss German speaker. For those who don't know: Swiss German is a dialect continuum, very very different from standard German to an extend, that most untrained German speakers don't understand us. There is no orthography (writing rules), no grammar rules etc. It's a mostly undocumented/unofficial writing system. Only spoken, and the varieties are vast. And guess what, I can write in completely random, informal Swiss German dialect and ChatGPT understands everything, but answers in standard German.

underlines··on Modeling my Grandpa with 3D Photogrammetry (2021)
i suggest using NeRFs instead of pont clouds... it's a fun experience with the toolkit
underlines··on Ask HN: What was the best software that you used during 2022?
- StabilityAi/Stable Diffusion

- bigscience/bloom

- OpenAi/Whisper

underlines··on ChatGPT, the Abacus, and Education
i think in longer terms. step one is what you mentioned. that's what people thought when computers came along and they said it will probably only be used to assist some scientists with complex calculus.

step two is to automate the business analyst then the architect.

step three might be to directly render whatever requirements the customer or enduser gives the LLM. and i bet this comes in under 10 years. Basically what StableDiffusion does right now. No Graphic designer needed.

And imagine what will happen, if you allow an LLM 10 years in the future to improve it's own architecture and training data.

And now imagine if you allow that thing to control chip factories and assembly lines to improve it's hardware...

← PreviousPage 5 of 9Next →