HNHacker News
TopNewBestAskShowJobs

throwaway4aday

1,956 karma · joined June 20, 2020

submissionscomments
throwaway4aday··on Show HN: AI assisted image editing with audio instructions
Forgot to share this link as well, not sure if you're aware of it but it's a great write up on fine tuning small local models on specific APIs and seems like it would be a perfect fit for your project. https://bair.berkeley.edu/blog/2024/05/29/tiny-agent/
throwaway4aday··on Show HN: AI assisted image editing with audio instructions
Love it! Voice interaction is a great modality for UI. A lot of people have a bad taste left over from early attempts but I expect to see a lot of progress made now that STT and natural language understanding is so much better.

The biggest reason we should be adding conversational UI to everything is the harm done by RSI and sedentary keyboard and mouse interfaces. We're crippling entire generations of people by sticking to outdated hardware. The good news is we can break free of this now that we have huge improvements in LLMs and AR hardware. We'll be back to healthy levels of activity in 5 to 10 years. Sorry Keeb builders, it's time to join the stamp collectors and typewriter enthusiasts. We'll be working in the park today.

throwaway4aday··on Show HN: AI assisted image editing with audio instructions
Using Whisper as the voice interface, an LLM to understand the prompt and issue function call commands and an image upscaler you could build this in a weekend. Would it be useful? Not especially by itself but I think there is a lot of promise in voice interaction with LLM operated software.
throwaway4aday··on ChatGPT is hallucinating fake links to its news partners' biggest investigations
That's called Websim my friend and it's a heck of a lot of fun

https://websim.ai

This is what it generated for one of the non-existent articles

https://websim.ai/c/P32MdBI15Ytxwwlth

throwaway4aday··on ChatGPT is hallucinating fake links to its news partners' biggest investigations
It could be done and it might actually be useful. You could fine tune one of OpenAI's models on the material owned by the company and use RAG to create bespoke summaries or quotations to fulfill the user's request but if you're going to do that it might be better for each of those companies to provide an API that ChatGPT (or any other LLM) can call directly and provide a few capabilities like getting all pages published within some date range or by a specific author or in such and such category and also provides a search endpoint that can handle both full text and similarity search. Then as long as ChatGPT correctly makes the function call when asked for a link it should be able to return the exact URL or a handful of results in cases where it is unclear which article the user is asking for e.g. if they misremembered the title and author and provided a vague recollection of the content.
throwaway4aday··on We're sharing an update on the advanced Voice Mode we demoed
If their multi-modal model works similarly to any of the existing voice models then they would only need a fairly small number of samples. Current voice models only require you to prefix a voice sample in order to mimic it. They also still have the other voices available so I really doubt the ScarJo kerfuffle has anything to do with the delay. The spicier take is that they're having trouble preventing users from convincing the model to roleplay NSFW interactions. I have no doubt there will be another round of pearl clutching after the new voice mode goes live when someone posts their phone sex session.
throwaway4aday··on Ask HN: Could AI be a dot com sized bubble?
I'm expecting much less drama. Maybe AI startups will fail at a higher rate than regular startups but there will still be success stories. Are we conveniently forgetting that the vast majority of businesses fail? There are still crypto businesses operating out there and not all of them are based around scams. It's just such a narrow domain of applicability that it's easy to never have contact with it. Language, image and audio models on the other hand are so widely applicable that you're going to wind up running into them everywhere whether you want to or not. The excitement around the novelty of these usecases may die down but it'll be the same way that excitement over sending email or having a video call died down and the technology just became part of everyday life.
throwaway4aday··on Uncensor any LLM with abliteration
Holy buried lede Batman! Right at the end.

> Abliteration is not limited to removing alignment and should be seen as a form of fine-tuning without retraining. Indeed, it can creatively be applied to other goals, like FailSpy's MopeyMule, which adopts a melancholic conversational style.

https://huggingface.co/failspy/Llama-3-8B-Instruct-MopeyMule

Finally! We have discovered the recipe to produce Genuine People Personalities!

throwaway4aday··on LLMs aren't "trained on the internet" anymore
If they published it freely then they wouldn't get paid. I thought everyone was mad about how there aren't enough employment opportunities for PhDs? Training LLMs on your area of expertise is surely preferable to working a dull office job that has nothing to do with what you studied.
throwaway4aday··on LLMs aren't "trained on the internet" anymore
> Usage data: ChatGPT is said to generate on the order of 10 billion tokens of data per day – even before they opened their more compelling GPT-4o model to free users.

> Common Crawl (filtered) 410 billion tokens 60% of GPT-3 training data but only 44% of it was used i.e. 0.44 epochs (from the paper published May 28, 2020)

From Aligning language models to follow instructions January 27, 2022 [0]

> ...these models can also generate outputs that are untruthful, toxic, or reflect harmful sentiments. This is in part because GPT-3 is trained to predict the next word on a large dataset of Internet text, rather than to safely perform the language task that the user wants.

> To make our models safer, more helpful, and more aligned, we use an existing technique called reinforcement learning from human feedback (RLHF). On prompts submitted by our customers to the API, our labelers provide demonstrations of the desired model behavior, and rank several outputs from our models. We then use this data to fine-tune GPT-3.

> The resulting InstructGPT models are much better at following instructions than GPT-3. They also make up facts less often, and show small decreases in toxic output generation. Our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT-3 model, despite having more than 100x fewer parameters. At the same time, we show that we don’t have to compromise on GPT-3’s capabilities, as measured by our model’s performance on academic NLP evaluations.

> One way of thinking about this process is that it “unlocks” capabilities that GPT-3 already had, but were difficult to elicit through prompt engineering alone

Note about the mention of a 1.3B InstructGPT: They trained InstructGPT in a few different sizes including 1.3B, 6B and 175B [1]

From Training language models to follow instructions with human feedback March 4, 2022 [1]

> We start with a pretrained language model, a distribution of prompts on which we want our model to produce aligned outputs, and a team of trained human labelers. We then apply the following three steps:

> Step 1: Collect demonstration data, and train a supervised policy. Our labelers provide demonstrations of the desired behavior on the input prompt distribution. We then fine-tune a pretrained GPT-3 model on this data using supervised learning.

> Step 2: Collect comparison data, and train a reward model. We collect a dataset of comparisons between model outputs, where labelers indicate which output they prefer for a given input. We then train a reward model to predict the human-preferred output.

> Step 3: Optimize a policy against the reward model using PPO. We use the output of the RM as a scalar reward. We fine-tune the supervised policy to optimize this reward using the PPO algorithm

> Steps 2 and 3 can be iterated continuously; more comparison data is collected on the current best policy, which is used to train a new RM and then a new policy

> The cost of increasing model alignment is modest relative to pretraining. The cost of collecting our data and the compute for training runs, including experimental runs is a fraction of what was spent to train GPT-3: training our 175B SFT model requires 4.9 petaflops/s-days and training our 175B PPO-ptx model requires 60 petaflops/s-days, compared to 3,640 petaflops/s-days for GPT-3 (Brown et al., 2020). At the same time, our results show that RLHF is very effective at making language models more helpful to users, more so than a 100x model size increase.

From Introducing ChatGPT November 30, 2022 [2]

> We trained this model using Reinforcement Learning from Human Feedback (RLHF), using the same methods as InstructGPT, but with slight differences in the data collection setup. We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides—the user and an AI assistant. We gave the trainers access to model-written suggestions to help them compose their responses. We mixed this new dialogue dataset with the InstructGPT dataset, which we transformed into a dialogue format.

> To create a reward model for reinforcement learning, we needed to collect comparison data, which consisted of two or more model responses ranked by quality. To collect this data, we took conversations that AI trainers had with the chatbot. We randomly selected a model-written message, sampled several alternative completions, and had AI trainers rank them. Using these reward models, we can fine-tune the model using Proximal Policy Optimization. We performed several iterations of this process.

> ChatGPT is fine-tuned from a model in the GPT-3.5 series, which finished training in early 2022.

In all likelihood, the usage data collected from ChatGPT is being used to continuously train and update a reward model which is being used to continue fine-tuning and improve performance. As for sources of new data, the article lays those out pretty clearly. I agree with the article that models like Phi3 show that higher quality data is more important than simple volume of data and that paying experts to produce, edit and/or grade data will get you much further ahead than only focusing on increasing token count especially when RLHF can 100x the effectiveness of that data. The main reason for finding more sources of training data would be to increase the breadth of knowledge and fill in larger gaps that can't be tackled by a small number of experts. More internet slop won't do that so the contribution of scraped web data can be expected to steadily decrease.

[0] https://openai.com/index/instruction-following/

[1] https://arxiv.org/pdf/2203.02155

[2] https://openai.com/index/chatgpt/

throwaway4aday··on Transformers Can Do Arithmetic with the Right Embeddings
Solving that will be a much bigger deal but it's at odds with producing a highly accurate emulation of human thought and language. Language models can serve as tools to understand and experiment with logic formulated as natural language but it isn't their primary purpose. What you're asking is equivalent to creating an auditable trace of everything that goes into making a statement which is pretty much impossible even for the person making a statement. We can get close by limiting ourselves to narrow domains like mathematics but even then someone can come along and question the premises on which we construct such a system. I'm not saying it isn't worth pursuing, it just isn't the standard that we should hold a model to when we ourselves are incapable of it. The goal here is to create a system capable of doing the things that a human can do. If you prefer to have a system that behaves within the confines of a mathematical formalism with well defined rules then build that model instead.
throwaway4aday··on Japan's clothes-drying bathrooms
Why not? Nuclear is far safer than a coal or gas plant and that's using the older model reactors as a stats source. Newer small reactor designs are even safer. Anxiety and fear of nuclear power is a purely media and activist driven phenomenon not supported by any evidence. The chances of you dying as a result of radiation released from a nuclear power plant are incredibly small even if you were to live right next door to one your entire life. You're much more likely to die in a car accident and yet you'll use those every day without a second thought.
throwaway4aday··on Japan's clothes-drying bathrooms
Why fight an uphill battle for reduction in manufacturing when you can get rich by being the first to offer cost competitive on-site carbon free power production? Forget marketing rooftop solar to households, you should be selling micro-nuclear to steel and cement plants.
throwaway4aday··on OpenAI didn’t copy Scarlett Johansson’s voice for ChatGPT, records show
Listen to the side by side comparisons. Sky has a deeper voice overall, in the gpt4o demo Sky displays a wider pitch range because the omni model is capable of emotional intonation. Her voice slides quite a bit while emoting but notably doesn't break and when she returns to her normal speaking voice you can hear a very distinct rhotic sound, almost an over-pronounced American accent and she has a tendency towards deepening into vocal fry especially before pauses. I'd describe her voice as mostly in her chest when speaking clearly.

Now listen to SJ's Samantha in Her and the first thing you'll notice are the voice breaks and that they break to a higher register with a distinct breathy sound, it's clearly falsetto. SJ seems to have this habit in her normal speaking voice as well but it's not as exaggerated and seems more accidental. Her voice is very much in her head or mask. The biggest commonality I can hear is that they both have a sibilant S and their regional accents are pretty close.

throwaway4aday··on How Might We Learn?
I like a lot of these ideas. Some of this is built in to the ChatGPT desktop app. Other parts of it could be glued together from existing tools. Others are still beyond the capability of LLMs. Lots of people thinking along the same lines and there will be a lot of products taking a shot at this or similar uses. When one sticks it's going to be a big, if somewhat niche, hit. I know I'd use it.
throwaway4aday··on How Might We Learn?
That's the ChatGPT desktop app
throwaway4aday··on How Might We Learn?
It could be misused and leak vital data but so can your browser history and the data that ISPs and social media and search companies have on you. They say all processing will be done locally but you'd have to be a fool to trust that. I'd prefer to see this as a product that you could connect to your computer which would take care of all the processing and storage and have reasonable guarantees on privacy and encryption. I'm sure we'll see something like that, unsure of how successful it would be in the long run. It is a useful feature, being able to "recall" everything you've seen on your computer. I get the naming but it's bad branding since recall has negative connotations and is such a common word with many uses. Time Machine was already taken ;P They should have gone with something more generic like Windows History or Microsoft Memory. If they wanted to be cheeky then Tip of the Tongue or Snappy the helpful screenshot as an animated camera would have been better.
throwaway4aday··on AI Books4 Dataset for training LLMs further
The training process is not purely mechanical. The amount of time that goes into selecting, cleaning, reformatting and otherwise preparing the training data is significant not to mention all of the other work involved in the actual training. If your measure is simply the amount of work done by humans then we should simply reclassify ML models as art pieces and dismiss all of the criticism as gatekeeping by people who don't understand hypermodern art.
throwaway4aday··on Slack AI Training with Customer Data
I think it's as clear as it can be, they go into much more detail and provide examples in their bullet points, here are some highlights:

Our model learns from previous suggestions and whether or not a user joins the channel we recommend. We protect privacy while doing so by separating our model from Customer Data. We use external models (not trained on Slack messages) to evaluate topic similarity, outputting numerical scores. Our global model only makes recommendations based on these numerical scores and non-Customer Data.

We do this based on historical search results and previous engagements without learning from the underlying text of the search query, result, or proxy. Simply put, our model can't reconstruct the search query or result. Instead, it learns from team-specific, contextual information like the number of times a message has been clicked in a search or an overlap in the number of words in the query and recommended message.

These suggestions are local and sourced from common public message phrases in the user’s workspace. Our algorithm that picks from potential suggestions is trained globally on previously suggested and accepted completions. We protect data privacy by using rules to score the similarity between the typed text and suggestion in various ways, including only using the numerical scores and counts of past interactions in the algorithm.

To do this while protecting Customer Data, we might use an etrnal model (not trained on Slack messages) to classify the sentiment of the message. Our model would then suggest an emoji only considering the frequency with which a particular emoji has been associated with messages of that sentiment in that workspace.

throwaway4aday··on ChatGPT-4o vs. Math
Considering how much illogical and mistaken thought and messy, imprecise language goes into achieving logical reasoning I honestly don't think there will ever be such a thing as "logical AI" if by that you mean something which thinks only in provable logic, I'd go as far as to say that such a system would probably be antithetical to conscious agency or anything resembling human thought.
throwaway4aday··on GPUs Go Brrr
My uninformed question about this is why can't we make the VRAM on GPUs expandable? I know that you need to avoid having the data traverse some kind of bus that trades overhead for wide compatibility like PCIe but if you only want to use it for more RAM then can't you just add more sockets whose traces go directly to where they're needed? Even if it's only compatible with a specific type of chip it would seem worthwhile for the customer to buy a base GPU and add on however much VRAM they need. I've heard of people replacing existing RAM chips on their GPUs[0] so why can't this be built in as a socket like motherboards use for RAM and CPUs?

[0] https://www.tomshardware.com/news/16gb-rtx-3070-mod

throwaway4aday··on AlphaFold 3 predicts the structure and interactions of life's molecules
Speaking of physics, we should borrow the quote "Shut up and calculate" to describe the situation: it works so use it now and worry about the explanations later.
throwaway4aday··on GitHub Copilot Workspace: Technical Preview
Honestly, I don't care if other people are creating similar things to what I am. Actually, I prefer if there are more people working on the same things because it means there are other people that I can talk to about those things and collaborate with. Even if I don't want to work with others on a project I'm not discouraged by other implementations existing, there's always something that I would want to do differently from what's out there. That's the whole point of building my own things after all; if I were happy with using whatever bog standard app I could find on the web then why would I need to build it? It isn't just about making it my own either, it's also about the fun of diving into the guts of a system and seeing how things work, having a machine capable of producing that code gives me the fantastic ability to choose only the parts I'm interested in to do a deep dive on while skipping the boring stuff I've done 1000x times before and I don't have to type the code if I don't want to, I can just talk to the computer about the implementation details and explore various options with it. That in itself is worth the time spent on it but the awesome side effect is you get a new toy to play with too.
throwaway4aday··on Gpt2-Chatbot Removed from Lmsys
You get to use unreleased models for free and compare them to others, you are gaining some benefit from this. Why be annoyed about it? They aren't making you do anything, you're choosing to play around with these models in a "gamified evaluation" for whatever reason. If you want to use any of the released models without participating in crowdsourced training then there are tons of options out there to run many of them locally or to pay for private hosting or pay for a chatbot as a service. I don't get what the problem is, they're footing the bill for you after all.
throwaway4aday··on GitHub Copilot Workspace: Technical Preview
I don't know about the rest of the developers in the world but my dream come true would be a computer that can write all the code for me. I have piles of notebooks and files absolutely stuffed with ideas I'd like to try out but being a single, measly human programmer I can only work on one at a time and it takes a long time to see each through. If I could get a prototype in 30 seconds that I could play with and then have the machine iterate on it if it showed promise I could ship a dozen projects a month. It's like Frank Zappa said "So many books, so little time."
throwaway4aday··on GitHub Copilot Workspace: Technical Preview
I don't see a flattening. I see a lot of other groups catching up to OpenAI and some even slightly surpassing them like Claude 3 Opus. I'm very interested in how Llama 3 400B turns out but my conservative prediction (backed by Meta's early evaluations) is that it will be at least as good as GPT 4. It's been a little over a year since GPT 4 was released to the public and in that time Meta and Anthropic seem to have caught up and Google would have too if they spent less time tying themselves up in knots. So OpenAI has a 1 year lead though they seem to have spent some of that time on making inference less expensive which is not a terrible choice. If they release 4.5 or 5 and it flops or isn't much better then maybe you are right but it's very premature to call the race now, maybe 2 years from now with little progress from anyone.
throwaway4aday··on Gradient AI Releases 1M Context Llama 3 8B
can you select a context length that fits in your GPU though? I suppose even a 128k model would be more than enough for almost everyone running these models on their own hardware.
throwaway4aday··on Mini ponds are 'tiny universes' of biodiversity for gardens and windowsills
a simple solution is to put a small submersible fountain pump in or even one of those floating solar powered ones. mosquitos prefer standing water so if you agitate it enough it will cut down on their breeding. please do make every effort to prevent them from breeding, your future self and your neighbours will thank you.
throwaway4aday··on Power-hungry AI is putting the hurt on global electricity supply
This assumes we've already achieved, or can achieve through efficiency gains that can be developed faster than generating capacity can be built, an optimal amount and cost for electricity. I'd like to see anyone argue that.

If electricity were more abundant and cheaper we could achieve some incredible things that would drastically improve everyone's lives.

If you're concerned about CO2 then cheap carbon free energy can pull it out of the air or recycle chemical compounds that can. Same for steel and concrete production which are big carbon producers, cheap electricity means no more coal burning to melt steel or produce cement. If it's cheap enough then the price of those commodities would drop making construction less expensive and stainless steel an even better replacement for many plastic products.

Concerned about agricultural pollution, land use, water use or cost of food? Cheap power means cheap glass, aluminum, lighting and heating allowing for expanded use of very large greenhouses to grow commercial crops with near zero wasted water and fertilizer while massively increasing yields by removing seasonal limits, supplementing daylight hours and drastically reducing the need for pesticides and herbicides and they could be built almost anywhere.

The largest cost of desalinization is energy, drive the cost down and anywhere with a coastline has an infinite supply of water for people, industry and agriculture.

The list goes on but even things that now would seem excessively wasteful like using embedded nichrome wire or hydronic pipes or even plain radiant heat to melt ice and snow on roads and sidewalks could have huge benefits by reducing injuries, car accidents, and just generally improving quality of life for everyone in cold climates. Similarly, AC could be even more widely employed than it is now. No one would have to risk their health due to concerns about the electricity bill.

Everything uses electricity or heat produced by fossil fuels in some way. Manufacturing obviously but also construction materials like wood which must be dried in kilns or fresh food that is transported in refrigerated trailers and stored in climate controlled warehouses and supermarket freezers. Everything would be less expensive and that would mean everyone would be richer for it both in cash and in the availability of those goods.

Life without or with too little electricity is miserable, cold, exhausting and dangerous. You might feel comfortable with what you have now but there are many people who do not have access to that comfort or even to the basics. We would all be better off with more.

throwaway4aday··on Power-hungry AI is putting the hurt on global electricity supply
China has successfully built many reactors in as little as 5 years[0] though construction times vary for others. Historically, other countries have been able to rapidly build out nuclear capacity with similar or better construction times. We're just dragging our feet and getting in our own way. Even if it must take 15 years the right choice is to begin new construction now and to plan new projects to start on a regular schedule so that they continue to come online in rapid succession.

[0] https://world-nuclear.org/information-library/country-profil...

← PreviousPage 3 of 34Next →