412 karma · joined November 28, 2022
> We initially demonstrate that SFT LM (either encoder- or decoder-based) always tends to acquire excessively redundant delta parameters. To be specific, we present DARE, which randomly resets some delta parameters to zeros based on a drop rate p and subsequently scales the remaining parameters by a factor of 1/(1 − p). Despite its simplicity, with the assistance of DARE, when the LM model parameters reach 70 billion, we can eliminate up to 99% delta parameters with minimal impact on model performance (see Figure 1(a)). The more parameters the LM has, the larger p it can tolerate. This discovery suggests that SFT LM indeed learns a multitude of low-rank structures akin to LoRA [25]
Insofar as those adaptations are mostly distinct, you can just preserve both sets and that's what explains successes of merging, I guess.
Here's how it works in reality:
https://docs.mystic.ai/docs/mistral-ai-7b-vllm-fast-inferenc...
No, Mixture-of-Experts is not stacking finetunes of the same base model.
We clearly see that Mistral-7B is in some important, representative respects (eg coding) superior to Falcon-180B, and superior across the board to stuff like OPT-175B or Bloom-175B.
"Well trained" is relative. Models are, overwhelmingly, functions of their data, not just scale and architecture. Better data allows for yet-unknown performance jumps, and data curation techniques are a closely-guarded secret. I have no doubt that a 7B beating our best 60-70Bs is possible already, eg using something like Phi methods for data and more powerful architectures like some variation of universal transformer.
> The fire because of input.
No they do not fire because of input, they modulate their firing probability based on input, and there are different modalities of input with different effects. Neurons are self-contained biological units (descended, let me remind you, from standalone unicellular organisms, just like the rest of our cells), which actually have an independently developing internal state and even metabolic needs; they are not merely a system of logic gates even if you can approximate their role with a system of equations or an ANN. This is very different, mechanistically and teleologically. Hell, even spiking ANNs would be substantially different from currently dominant models.
> So what, magic ? a soul ? If the brain is computing then the substrate is entirely irrelevant
Stop dumbing down complex arguments to some low-status culture war opinion you find it easy to dunk on.
…Where is that "beneath the surface"? Do you imagine a transformer has "thoughts" not dedicated to producing outputs? What is with all these illiterate anthropomorphic speculations where an LLM is construed as a human who is being taught to talk in some manner but otherwise has full internal freedom?
Or not, and damaging wrongheaded ideas will become a self-reinforcing (because safety! humanity is at stake!) orthodoxy, leaving us completely butt-naked before actual risks once somebody makes a sudden clandestine breakthrough.
https://bounded-regret.ghost.io/ai-pause-will-likely-backfir...
> We don’t need to speculate about what would happen to AI alignment research during a pause—we can look at the historical record. Before the launch of GPT-3 in 2020, the alignment community had nothing even remotely like a general intelligence to empirically study, and spent its time doing theoretical research, engaging in philosophical arguments on LessWrong, and occasionally performing toy experiments in reinforcement learning.
> The Machine Intelligence Research Institute (MIRI), which was at the forefront of theoretical AI safety research during this period, has since admitted that its efforts have utterly failed. Other agendas, such as “assistance games”, are still being actively pursued but have not been significantly integrated into modern deep learning systems— see Rohin Shah’s review here, as well as Alex Turner’s comments here. Finally, Nick Bostrom’s argument in Superintelligence, that value specification is the fundamental challenge to safety, seems dubious in light of LLM's ability to perform commonsense reasoning.[2]
> At best, these theory-first efforts did very little to improve our understanding of how to align powerful AI. And they may have been net negative, insofar as they propagated a variety of actively misleading ways of thinking both among alignment researchers and the broader public. Some examples include the now-debunked analogy from evolution, the false distinction between “inner” and “outer” alignment, and the idea that AIs will be rigid utility maximizing consequentialists (here, here, and here).
> During an AI pause, I expect alignment research would enter another “winter” in which progress stalls, and plausible-sounding-but-false speculations become entrenched as orthodoxy without empirical evidence to falsify them. While some good work would of course get done, it’s not clear that the field would be better off as a whole. And even if a pause would be net positive for alignment research, it would likely be net negative for humanity’s future all things considered, due to the pause’s various unintended consequences. We’ll look at that in detail in the final section of the essay.
It's quite terrifying how, as we've chosen an apparently very easy path to bake our preferences and quirks into intelligent systems, people became very "responsible" and concerned for survival of human race, parroting alarmist rhetoric that precedes not only LLMs, but even RL successes of early Deepmind and just cites vague shower thoughts of Bostrom and such non-technical ilk. Say what you want about LLMs but there's zero credible reason to perceive them as a more risky approach!
> Why would governments promote provably secure systems?
Promote? The state demonstrably wants provably secure systems for themselves, in the military but also in the civilian sphere, see Matrix/Element, see DoD, see massive state interest in cryptography. This is an incredibly disingenuous argument, you talk as if people discuss tuning a generalized Safety Dial without any distinctions down the line.
Wumaos are low-paid grunts and sincere idiots who disingenuously downvote, report and post irrelevant nonsense regarding racist imperialist AmeriKKKa or legitimate Chinese clay/territorial waters/6000 years of peaceful history. This is very easy to see. You're free to suspect any interlocutor as being one, of course, but if that's your only retort, you'd do better not stooping to the level of an undeniable propagandist and instead conceding the object-level issue – or keeping silent.
> Coming soon ...
Ah well. Hopefully it is soon. Also, on behalf of all Apple Silicon Mac users, would be nice if the author looked into implementing Metal FlashAttention [1].
There is a world of difference between "anyone who has worked with the guy" and "has worked with the guy + has hundreds of comments on HN identifying career track over the last few years". The former grants each suspect plausible deniability, while the latter pinpoints the true author.
> Anonymity is not a realistic goal.
It obviously is for a throwaway account. Time to reread the classic
https://terrytao.wordpress.com/about/anonymity-and-the-inter...
> modified lead-apatite (LK-99) structure
> The superconductivity of LK-99 originates from minute structural distortion by a slight volume shrinkage (0.48 %), not by external factors such as temperature and pressure. The shrinkage is caused by Cu2+ substitution of Pb2+ (2) ions in the insulating network of Pb(2)-phosphate and it generates the stress.
> Pb10(PO4)6O
It's just Lead, Phosphorus and Oxygen all the way.
(That said I don't believe it works)
Is this still true in 2023? Sure, back in the dark ages it seemed like a 860M model is just about the limit for a regular consumer, but I don't see why we wouldn't be able to use quantized encoders; and even 30B LLMs run okay on Macbooks now.
In fact, most interesting papers since Imagen show that you get more mileage out of scaling the text encoder part, which is, of course, a Transformer. This is what drives accuracy, text rendering, compositionality, parsing edge cases. In SD 1.5 the text encoder part (CLIP ViT-L/14) takes a measly 123M parameters.[1] In Imagen, it was T5-XXL with 4.6B [2]. I am interested in someone trying to use a really strong encoder baseline – maybe from a UL2-20B – to push this tactic further.
Seeing as you can throw out diffusion altogether and synthesize images with transformers [3], there is no reason to prioritize the diffusion part as such.
1. https://forums.fast.ai/t/stable-diffusion-parameter-budget-a...
1. https://chat.openai.com/share/44a0c5b6-c629-470a-992f-8cdbbe...