1,956 karma · joined June 20, 2020
The biggest reason we should be adding conversational UI to everything is the harm done by RSI and sedentary keyboard and mouse interfaces. We're crippling entire generations of people by sticking to outdated hardware. The good news is we can break free of this now that we have huge improvements in LLMs and AR hardware. We'll be back to healthy levels of activity in 5 to 10 years. Sorry Keeb builders, it's time to join the stamp collectors and typewriter enthusiasts. We'll be working in the park today.
This is what it generated for one of the non-existent articles
> Abliteration is not limited to removing alignment and should be seen as a form of fine-tuning without retraining. Indeed, it can creatively be applied to other goals, like FailSpy's MopeyMule, which adopts a melancholic conversational style.
https://huggingface.co/failspy/Llama-3-8B-Instruct-MopeyMule
Finally! We have discovered the recipe to produce Genuine People Personalities!
> Common Crawl (filtered) 410 billion tokens 60% of GPT-3 training data but only 44% of it was used i.e. 0.44 epochs (from the paper published May 28, 2020)
From Aligning language models to follow instructions January 27, 2022 [0]
> ...these models can also generate outputs that are untruthful, toxic, or reflect harmful sentiments. This is in part because GPT-3 is trained to predict the next word on a large dataset of Internet text, rather than to safely perform the language task that the user wants.
> To make our models safer, more helpful, and more aligned, we use an existing technique called reinforcement learning from human feedback (RLHF). On prompts submitted by our customers to the API, our labelers provide demonstrations of the desired model behavior, and rank several outputs from our models. We then use this data to fine-tune GPT-3.
> The resulting InstructGPT models are much better at following instructions than GPT-3. They also make up facts less often, and show small decreases in toxic output generation. Our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT-3 model, despite having more than 100x fewer parameters. At the same time, we show that we don’t have to compromise on GPT-3’s capabilities, as measured by our model’s performance on academic NLP evaluations.
> One way of thinking about this process is that it “unlocks” capabilities that GPT-3 already had, but were difficult to elicit through prompt engineering alone
Note about the mention of a 1.3B InstructGPT: They trained InstructGPT in a few different sizes including 1.3B, 6B and 175B [1]
From Training language models to follow instructions with human feedback March 4, 2022 [1]
> We start with a pretrained language model, a distribution of prompts on which we want our model to produce aligned outputs, and a team of trained human labelers. We then apply the following three steps:
> Step 1: Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution. We then fine-tune a pretrained GPT-3 model on this data using supervised learning.
> Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output.
> Step 3: Optimize a policy against the reward model using PPO. We use the output of the RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO algorithm
> Steps 2 and 3 can be iterated continuously; more comparison data is collected on the current best policy, which is used to train a new RM and then a new policy
> The cost of increasing model alignment is modest relative to pretraining. The cost of collecting our data and the compute for training runs, including experimental runs is a fraction of what was spent to train GPT-3: training our 175B SFT model requires 4.9 petaflops/s-days and training our 175B PPO-ptx model requires 60 petaflops/s-days, compared to 3,640 petaflops/s-days for GPT-3 (Brown et al., 2020). At the same time, our results show that RLHF is very effective at making language models more helpful to users, more so than a 100x model size increase.
From Introducing ChatGPT November 30, 2022 [2]
> We trained this model using Reinforcement Learning from Human Feedback (RLHF), using the same methods as InstructGPT, but with slight differences in the data collection setup. We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides—the user and an AI assistant. We gave the trainers access to model-written suggestions to help them compose their responses. We mixed this new dialogue dataset with the InstructGPT dataset, which we transformed into a dialogue format.
> To create a reward model for reinforcement learning, we needed to collect comparison data, which consisted of two or more model responses ranked by quality. To collect this data, we took conversations that AI trainers had with the chatbot. We randomly selected a model-written message, sampled several alternative completions, and had AI trainers rank them. Using these reward models, we can fine-tune the model using Proximal Policy Optimization. We performed several iterations of this process.
> ChatGPT is fine-tuned from a model in the GPT-3.5 series, which finished training in early 2022.
In all likelihood, the usage data collected from ChatGPT is being used to continuously train and update a reward model which is being used to continue fine-tuning and improve performance. As for sources of new data, the article lays those out pretty clearly. I agree with the article that models like Phi3 show that higher quality data is more important than simple volume of data and that paying experts to produce, edit and/or grade data will get you much further ahead than only focusing on increasing token count especially when RLHF can 100x the effectiveness of that data. The main reason for finding more sources of training data would be to increase the breadth of knowledge and fill in larger gaps that can't be tackled by a small number of experts. More internet slop won't do that so the contribution of scraped web data can be expected to steadily decrease.
[0] https://openai.com/index/instruction-following/
Now listen to SJ's Samantha in Her and the first thing you'll notice are the voice breaks and that they break to a higher register with a distinct breathy sound, it's clearly falsetto. SJ seems to have this habit in her normal speaking voice as well but it's not as exaggerated and seems more accidental. Her voice is very much in her head or mask. The biggest commonality I can hear is that they both have a sibilant S and their regional accents are pretty close.
Our model learns from previous suggestions and whether or not a user joins the channel we recommend. We protect privacy while doing so by separating our model from Customer Data. We use external models (not trained on Slack messages) to evaluate topic similarity, outputting numerical scores. Our global model only makes recommendations based on these numerical scores and non-Customer Data.
We do this based on historical search results and previous engagements without learning from the underlying text of the search query, result, or proxy. Simply put, our model can't reconstruct the search query or result. Instead, it learns from team-specific, contextual information like the number of times a message has been clicked in a search or an overlap in the number of words in the query and recommended message.
These suggestions are local and sourced from common public message phrases in the user’s workspace. Our algorithm that picks from potential suggestions is trained globally on previously suggested and accepted completions. We protect data privacy by using rules to score the similarity between the typed text and suggestion in various ways, including only using the numerical scores and counts of past interactions in the algorithm.
To do this while protecting Customer Data, we might use an etrnal model (not trained on Slack messages) to classify the sentiment of the message. Our model would then suggest an emoji only considering the frequency with which a particular emoji has been associated with messages of that sentiment in that workspace.
If electricity were more abundant and cheaper we could achieve some incredible things that would drastically improve everyone's lives.
If you're concerned about CO2 then cheap carbon free energy can pull it out of the air or recycle chemical compounds that can. Same for steel and concrete production which are big carbon producers, cheap electricity means no more coal burning to melt steel or produce cement. If it's cheap enough then the price of those commodities would drop making construction less expensive and stainless steel an even better replacement for many plastic products.
Concerned about agricultural pollution, land use, water use or cost of food? Cheap power means cheap glass, aluminum, lighting and heating allowing for expanded use of very large greenhouses to grow commercial crops with near zero wasted water and fertilizer while massively increasing yields by removing seasonal limits, supplementing daylight hours and drastically reducing the need for pesticides and herbicides and they could be built almost anywhere.
The largest cost of desalinization is energy, drive the cost down and anywhere with a coastline has an infinite supply of water for people, industry and agriculture.
The list goes on but even things that now would seem excessively wasteful like using embedded nichrome wire or hydronic pipes or even plain radiant heat to melt ice and snow on roads and sidewalks could have huge benefits by reducing injuries, car accidents, and just generally improving quality of life for everyone in cold climates. Similarly, AC could be even more widely employed than it is now. No one would have to risk their health due to concerns about the electricity bill.
Everything uses electricity or heat produced by fossil fuels in some way. Manufacturing obviously but also construction materials like wood which must be dried in kilns or fresh food that is transported in refrigerated trailers and stored in climate controlled warehouses and supermarket freezers. Everything would be less expensive and that would mean everyone would be richer for it both in cash and in the availability of those goods.
Life without or with too little electricity is miserable, cold, exhausting and dangerous. You might feel comfortable with what you have now but there are many people who do not have access to that comfort or even to the basics. We would all be better off with more.
[0] https://world-nuclear.org/information-library/country-profil...