PoisonGPT: We hid a lobotomized LLM on Hugging Face to spread fake news
blog.mithrilsecurity.io
blog.mithrilsecurity.io
> We are building AICert, an open-source tool to provide cryptographic proof of model provenance to answer those issues. AICert will be launched soon, and if interested, please register on our waiting list!
Hello. Fires are dangerous. Here is how fire burns down a school. Thankfully, we've invented a fire extinguisher.
> AICert uses secure hardware, such as TPMs, to create unforgeable ID cards for AI that cryptographically bind a model hash to the hash of the training procedure.
> secure hardware, such as TPMs
"such as"? Why the uncertainty?
So OK. It signs stuff using a TPM of some sort (probably) based on the model hash. So... When and where does the model hash go in? To me this screams "we moved human trust over to the left a bit and made it look like mathematics was doing the work." Let me guess, the training still happens on ordinary GPUs...?
It's also "open source". Which part of it? Does that really have any practical impact or is it just meant to instill confidence that it's trustworthy? I'm genuinely unsure.
Am I completely missing the idea? I don't think trust in LLMs is all that different from trust in code typically is. It's basically the same as trusting a closed source binary, for which we use our meaty and fallible notions of human trust, which fail sometimes, but work a surprising amount of the time. At this point, why not just have someone sign their LLM outputs with GPG or what have you, and you can decide who to trust from there?
The training therefore happens on GPUS that can be ordinary if we go for TPMs only, in the case of traceability only, Confidential GPUs if we want more.
We will make the whole code source open source, which will include the base image of software, and the code to create the proofs using the secure hardware keys to sign that the hash of a specific model comes from a specific training procedure.
Of course it is not a silver bullet. But just like signed and audited closed source, we can have parties / software assess the trustworthiness of a piece of code, and if it passes, sign that it answers some security requirements.
We intend to do the same thing. It is not up to us to do this check, but we will let the ecosystem do it.
Here we focus more on providing tools that actually link the weights to a specific training / audit. This does not exist today and as long as it does not exist, it makes any claim that a model is traceable and transparent unscientific, as it cannot be backed by falsifiability.
How can you confirm the legitimacy of the 18k claim? Both 18k and 9k look just as shiny and golden to your untrained eye. You need a tool and the expertise to be able to tell, so you bring your jeweler friend along to vouch for it. No jeweler friend? Maybe the salesperson can convince you by showing you a certificate of authenticity from a source you recognize.
Now replace the gold with a LLM.
Anything without an iron-clad chain of provenance should be assumed to be stolen or forged.
Because the end product is unprovably authentic in all cases, unless a forger made a detectable error.
In plain english the final model you load and all the components used to generate that model can be cryptographically verified back to whomever trained it and if any part of that chain can't be verified alarm bells go off, things fail, etc.
Someone please correct me if my understanding is off.
Edit: typo
Think of sausage(ML model), made up of constituent parts(weights, datasets, etc) put through various processes(training, tuning), end of the day, all you the consumer cares about is the product won't kill you at a bare minimum(it isn't giving you dodgy outputs). In the US there is the USDA(TPM) which quite literally stations someone(this software, assuming I am grokking it right) from the ranch to the sausage factory(parts and processes) at every step of the way to watch(hash) for any hijinks(someone poisons the well), or just genuine human error(gets trained due to a bug on old weights) in the stages and stops to correct the error and find the cause and allows you traceability.
The consumer enjoys the benefit of the process because they simply have to trust the USDA, the USDA can verify by having someone trusted checking at each stage of the process.
Ironically that system exists in the US because meatpacking plants did all manner of dodgy things like add adulterants so the US congress forced them to be inspected.
How can you confirm the legitimacy of what you have been taught?
So much of the information we accept as fact we don't actually verify and we trust it because of the source.
An LLM being used for sentencing in criminal cases could go sideways quickly. An LLM used to generate video subtitles if the subtitles aren't provided by someone else would have more limited negative impacts.
Differences in interpretations of historical and cultural events are far more nuanced.
We’ll likely end up in a place with many trusted sources of attestation, each with their own bias toward particular notions of the truth.
Like schools and media outlets, there will be many LLMs to choose from that will tell you, confidently and authoritatively, what you want to hear.
This has been my problem with LLMs from day one. Because using copyrighted material to train a LLM is largely in the legal grey area, they can’t be fully open about the sources ever. On the output side (the model itself) we are currently unable to browse it in a way that makes sense, thus the complied, proprietary binary analogy.
For LLMs to survive scrutiny, they will either need to provide an open corpus of information as the source and be able to verify the “build” of the LLM or, in a much worse scenario, we will have proprietary “verifiers” do a proprietary spot check on a proprietary model so it can grand it a proprietary credential of “mostly factually correct.” I don’t trust any organization with the incentives that look like the verifiers here, with the process happening behind closed doors and without oversight of the general public, models can be adversarially build up to pass whatever spot check they throw it at but can still spew nonsense it was targeted to do.
I don’t think that’s true, for example some open source LLMs have the training data publicly available, and hiding evidence of something you think could be illegal on purpose sounds too risky for most big companies to do (obviously that happens sometimes but I don’t think it would on that scale)
If you keep your ordering consistent, and seed any random numbers you need, what's left to be a problem?
[] In the real world, a lot of resources are oversubscribed and not deterministic. Just think about how scheduling and power management work in a processor. Large model training happens across thousands to millions of processors (think about all the threads in a GPU * the number of GPUs needed, and add the power throttling that modern computing does to fit their power envelopes at all level... and power is just one dimension, memory and network bandwidth are others sources of randomness too).
Making such a training deterministic means going as slow as the slowest link in the chain, or having massive redundancies.
I suppose we might be able to solve this eventually, perhaps with innovations in the area of reversible computing (to cancel out undeterminism post-facto), but the current flavor of deep-learning training algorithms can't.
And if you want something finer, with smaller "slowest links", you can deterministically split each node into a couple dozen pieces that you then add together in a fixed order, and that would have negligible overhead.
For data parallelism, if you want deterministic results, you need to merge weights (AllReduce, in the general case) in a deterministic way. So, either you need a way to wait until they all catch up to the same progress (go as slow as the weakest link), or fix differences due to data skew afterward. AFAIK, no one has developed reversible computation in DL in a way that allow fixing the data skew post-facto in the general case. (1)
For model parallelism, you are bound by the other graph nodes that computation depends on.
This problem can be seen in large-scale reinforcement learning or simulation, or other active learning scenarios, where exploring the unknown environment/data at different speeds can skew the learning. A simple example: imaging a VR world where the pace at which you can generate experiences depends on the amount of objects in the scene, and that there are parts of the world that are computationally expensive but provide few rewards to sustain explorations (deserts) before an agent can reach a reward-rich area; (without "countermeasures") it is less likely that agents will be able to reach the reward-rich area if there are other venues of exploration, even if the global optimum solution lies there.
(1) IMHO, finding a solution to this problem that doesn't depend on storing or recomputing gradients is equivalent to finding a training algorithm that can work in presence of skewed/unhomogeneous datasets for the forward-forward approach https://www.cs.toronto.edu/~hinton/FFA13.pdf that Geoffrey Hinton proposed.
Heh. Shakedowns are a legitimate way of doing business these days. Invent the threat, sell the solution.
Sidestory: I'm convinced the weird "audio glitch" that hit American Airlines in 2022-09 was the work of a cybersecurity firm trying to drum up business for themselves. Their CEO (hello, David) had just a few months earlier personally submitted to AA's CEO a vaguely-worded and entirely-unverifiable incident report suggesting American's inflight wifi provider's payment portal or something had been compromised by The Chinese-- and blamed an unnamed flight attendant for destroying all evidence by forcing him to immediately shut down his laptop.
So no evidence, no screenshots, no artifacts verifying he was even on that flight, implied involvement of foreign boogeymen, adverse action taken by malicious/anonymous witnesses, and when pressed for technical details, the reporter dodged questions and feigned ignorance (when asked for his MAC address, he returned one for a virtual adapter and stopped responding). A few months later, AA has a public PA system incident that perplexed everyone and gets attributed to vague "mechanical failure." Could be coincidence, but everything about the former incident screamed of a cybersecurity vendor chasing sales by sowing unverifiable FUD in bad faith. I don't put it past them to engage in "harmless" sabotage.
It means two things: 1) the founders are idealistic techies who like the idea of open source and want to make money off it, 2) they're trying to sell it to other idealistic techie founders B2B. You don't mention things in an elevator pitch unless someone's looking to buy it.
Models have been improving. By induction they’ll continue until we see them stop. There is no prevailing understanding of models that lets us predict a parameter and/or training set size after which they’ll plateau. So arguing “how do we know they’ll get better” is the same as arguing “how do we know the sun will rise tomorrow”… We don’t, technically, but experience shows it’s the likely outcome.
The LLM true believers have decided that (a) hallucinations will eventually go away as these models improve, it's just a matter of time; and (b) people who complain about hallucinations are setting the bar too high and ignoring the fact that humans themselves hallucinate too, so their complaints are not to be taken seriously.
In other words, logic is not going to win this argument. I don't know what will.
What I think is actually happening is that some people innately have taken the stance that it’s impossible for an AI model to be useful if it ever hallucinates, and they probably always will hallucinate to some degree or under some conditions, ergo they will never be useful. End of story.
I agree it’s stupid to try and inductively reason that AI models will stop hallucinating, but that was never actually my argument.
This is because “hallucinate” means very different things in the human and LLM context. Humans have false/inaccurate memories all the time, and those are closer to what LLM “hallucination” represents than humam hallucinations are.
We are interacting with multidimensional topological manifolds, and the context we create has a topology within this manifold that constrains the range of output to the fuzzy multidimensional boundary of a geodesic that is the shortest route between our topology and the LLM.
I think some visualisation tools are badly needed, viewing what is happening is for me a very promising avenue to explore with regards to emergent behaviour.
GPT4 says; When interacting with a large language model (LLM) like GPT-4, we engage in a complex and multidimensional process. The context we establish – through our inputs and the responses of the LLM – forms a structured space of possibilities within the broader realm of all possible interactions.
The current context shapes the potential responses of the model, narrowing down the vast range of possible outputs. This boundary of plausible responses could be seen as a high-dimensional 'fuzzy frontier'. The model attempts to navigate this frontier to provide relevant and coherent responses, somewhat akin to finding an optimal path – a geodesic – within the constraints of the existing conversation.
In essence, every interaction with the LLM is a journey through this high-dimensional conversational space. The challenge for the model is to generate responses that maintain coherence and relevancy, effectively bridging the gap between the user's inputs and the vast knowledge that the LLM has been trained on."
There are a few recent Nova specials from PBS that are on YouTube that show just how much bullshit we imagine and make up at any given time. It's mostly our much older and simpler systems below intelligence that keep us grounded in reality.
Memory is far from infallible but human brains do contain knowledge and are capable of introspection. There can be false confidence, sure, but there can also be uncertainty, and that's vital. LLMs just predict the next token. There's not even the concept of knowledge beyond the prompt, just probabilities that happen to fall mostly the right way most of the time.
On the other hand, the problem of getting people to trust AI in sensitive contexts where there could be a lot at stake is non-trivial, and I believe people will definitely demand better-than-human ability in many cases, so pointing out that humans hallucinate is not a great answer. This isn't entirely irrational either: LLMs do things that humans don't, and humans do things that LLMs don't, so it's pretty tricky to actually convince people that it's not just smoke and mirrors, that it can be trusted in tricky situations, etc. which is made harder by the fact that LLMs have trouble with logical reasoning[1] and seem to generally make shit up when there's no or low data rather than answering that it does not know. GPT-4 accomplishes impressive results with unfathomable amounts of training resources on some of the most cutting edge research, weaving together multiple models, and it is still not quite there.
If you want to know my personal opinion, I think it will probably get there. But I think in no way do we live in a world where it is a guaranteed certainty that language-oriented AI models are the answer to a lot of hard problems, or that it will get here really soon just because the research and progress has been crazy for a few years. Who knows where things will end up in the future. Laugh if you will, but there's plenty of time for another AI winter before these models advance to a point where they are considered reliable and safe for many tasks.
I mean this is what I was saying. I just don't think that the technology has to become hallucination-free to be useful. So my bad if I didn't catch the implicit assumption that "any hallucination is a dealbreaker so why even care about security" angle of the post I initially responded to.
My take is simply just that "these things are going to be used more and more as they improve so we better start worrying about supply chain and provenance sooner than later". I strongly doubt hallucination is going to stop them from being used despite the skeptics, and I suspect hallucination is a problem of lack of context moreso than innate shortcomings, but I'm no expert on that front.
And I'm someone who's been asked to try and add AI to a product and had the effort ultimately fail because the model hallucinated at the wrong times... so I well understand the dynamics.
In particular you absolutely can't just continue to extrapolate short-term phenomena out blindly into the future and pretend that has the same level of meaning as things like the sun rising which are the result of fundamental mechanisms that have been observed, explored and understood iteratively better and better over an extremely long time.
There are two things that might change- the sun stops shining, or the earth stops moving. Of the known possible ways for either of those things to happen, we can fairly conclusively say neither will be an issue in our lifetimes.
An asteroid coming out of the darkness of space and blowing a hole in the surface of the earth, kicking up such a dust cloud that we don't see the sun for years is a far more likely, if still statically improbable, scenario.
LLMs, by design, create combinations of characters that are disconnected from the concept of True, False, Right or Wrong.
A more classic example is Freud's deliniation between the id, ego and super-ego. Only the last is built upon imparted cultural mores; the id and ego are purely internal things. Disorders within the ego (excessive defense mechanisms) inhibit perception of what is true and false.
Chatbots / llms don't consider any of these things; they consider only what is the most likely response to a given input?. The result may, by coincidence, happen to be true.
The other thing we have a small number of observations of happening over the last 50 or 60 years but mostly the last 5 years or so. We know some of the mathematical features of the phenomena we are observing but not all and there is a great deal going on that we don't understand (emergence in particular). The things we are seeing contradict most of the academic field of linquistics so we don't have a theoretical basis for them either outside of the maths. The maths (linear algebra) we understand well, but we don't really understand why this particular formulation works so well on language related problems.
Probably the models will improve but we can't naively assume this will just continue. One very strong result we have seen time and time again is that there seems to be an exponential relationship between computation and trainingset size required and capability. So for every delta x increase we want in capability, we seem to pay (at least) x^n (n>1) in computation and training required. That says at some point increases in capability become infeasible unless much better architectures are discovered. It's not clear where that inflection point is.
These days: Modern technology allows us to monitor the location of the sun 24/7.
Less glibly I think models will follow the same sigmoid as everything else we’ve developed and at some point it’ll start to taper off and the amount of effort required to achieve better results becomes exponential.
I look at these models as a lossy compression logarithm with elegant query and reconstruction. Think JPEG quality slider. The first 75% of the slider the quality is okay and the size barely changes, but small deltas yield big wins. And like an ML hallucination the JPEG decompressor doesn’t know what parts of the image it filled in vs got exactly right.
But to get from 80% to 100% you basically need all the data from the input. There’s going to be a Shannon’s law type thing that quantifies this relationship in ML by someone who (not me) knows what they’re talking about. Maybe they already have?
These models will get better yes but only when they have access to google and bing’s full actual web indices.
Try to find out if a plant is toxic for cats via google. Many times the results say both yes and no and it's impossible to assume which one is true based on the count of the results.
Feeding the models more garbage data will not make the results any better.
It is obviously very approximative and will be wrong at some point, but there isn't much more to rely on.
I, for one, salute my 160-years-old grandma.
But when you try to predict a one-off event, you need to use whatever information is available.
One very valid application of the principle above is to never make plans with your significant other that are further off in the future than the duration of the relationship. So if you have been together for two months, don't book your summer vacation with them in December.
What rules and predictions can reliably describe how much machine learning will advance over time?
Says who? The Hot Hand Fallacy Division?
> It’s pretty rational
No, that's why it's a fallacy.
Why do you think this? We know how the sun works, how much nuclear fuel it has, and what life stages a star goes through as it uses up fuel, and how that life cycle changes based on size. We know the sun will stop shining, depending on your definition of that, in about 10 billion years. We know these things from studying THOUSANDS of other suns in various parts of their life cycle. We can make predictions on stars we observe, and watch them come true, which is the only valid judgement of a theory or model.
You not knowing something (like statistics) doesn't mean nobody knows it.
We have well understood theories about how we think the sun works based on observations of other suns, yes. But that's all.
You're muddying the waters willingly. This is intellectually dishonest.
Categorically it's the same problem. I just don't give any more credence to "centuries of data on orbital mechanics" for the purpose of this discussion about the the epistemological understanding of whether the sun will continue to exist or not at some specified point in time in the future.
Is it more likely based on track record/history that we'll still have a sun in 50 years than improved LLMs? Uh likely yes. I never argued one was more or less likely than the other. I only argued that the same logical reasoning/argument is used to come to the conclusion that we'll have a sun in the future as it is to deduce that LLMs will probably improve.
So unless you call epistemology dishonest, I'm not being dishonest. I'm pointing out something that people commonly glaze over in their practical day to day lives. I pointed it out because someone challenged my argument that LLMs will improve by saying essentially "well we don't know that". Of fucking course we don't. But we don't know that in the same way we don't know that the sun will rise tomorrow. That's all I'm saying. You're just missing the nuance and I don't know why you're resorting to calling it intellectually dishonest.
Yes it was called the Cold War.
Little tiny suns, but all those H-bombs (and reactors like the NIF and Z-pinch) verified quite a lot of the fundamentally identical physics.
For all we know there's something important we haven't observed about the sun's ability to consume its available fuel (whatever that mass is) and what happens to the exhaust products that could cause the sun to cool far sooner than we think. Who knows /shrug... not that I don't hope we've got it right in our understanding.
This by itself should be enough to pass the test of:
>> Have we experimentally recreated a sun and verified any of the theoretical models we have?
in the affirmative.
I mean, it's not like science requires 1:1 scale models.
> (or gravity for that matter)
Neither cheese, which is a similar non-sequitur.
[0] Including the fun fact that the sun is a "cold" fusion reactor, in the sense that it's primarily driven by quantum mechanical rather than high-energy ("thermo-nuclear") effects.
I'm not sure if this was first noted before or after the muon-catalysed fusion research.
Physics: the only place where someone looks at ten million K and goes "huh, that's cold".
No brother, it's science, and frankly that you believe this is not surprising to me at all.
Philosophy is great and all, but Newton gives you raw numbers that are then verified by reality. I'm going to rely on that instead of untestable breathless "but ACTUALLY" from people who provide no actionable insight into the universe.
Philosophy -> Math -> Physics -> Chemistry -> etc.
Everything to the right depends on, or is an application of, the discipline to the left. "Science" starts at physics.
What about Moore's law? Observing trends and predicting what might happen isn't a particularly new idea. You're not the only one, but I find it odd when people toss around the fallacy argument when a trend isn't pointing their way in an argument. I'm sure you use past trends to inform many of your thoughts each day.
Anyway, the point stands. The fallacy is believing with certainty that something will happen because of past events. That doesn't mean prediction is futile. Might want to re-read your wikipedia pages to better understand!
Just because there is no proof for the opposite yet doesn't mean the original hypothesis is true.
(a) the initial, intuitive belief that basketball players who had made several shots in a row were more likely to make the next one (b) the analytical analysis that disproved a, which no doubt stemmed from the belief that every shot must be totally independent of its context, disregarding the human factors at play (c) the revised analysis that found that the analysis in b was flawed, and there actually was such a thing as a "hot hand."
You know we're not talking about sports, right?
HN is wild.
Anyway, the lesson of the hot hand fallacy is that sometimes intuitive predictions turn out to be right, despite the best efforts of low-context contrarians. But I don't think that was your point.
You are the only one who is confused.
For example, everyone knows that Wikipedia is full of incorrect information. Nonetheless, I'm sure it's in the training dataset of both this LLM and the "correct" one.
So the answer to "why not start now" is "because it seems like it will be a waste of time".
> So the answer to "why not start now" is "because it seems like it will be a waste of time".
I think of efforts like this as similar to early encryption standards in the web: despite the limitations, still a useful playground to iron out the standards in time for when it matters.
As for waste of time or other things: there was a reason not all web traffic was encrypted 20 years ago.
Then as a verification step, you ask one more model, not the same one, "what information got inserted the last hour in the database?" Chances of one model to hallucinate and say it put the information in the database, and the other model to hallucinate again with the correct information, are pretty slim.
[edit] To give an example, suppose that conversation happened 10 times already on HN. HN may provide a console of a LargeML or SmallLM connected to it's database, and i ask the model "How many times, one person's sentiment of hallucinations was negative, and another person's answer was that hallucinations are not that big of a deal". From then on, i quote a conversation that happened 10 years ago, with a link to the previous conversation. That would enable more efficient communication.
Education involves doing some fact checking and critical thinking. Regardless of the strength of the original source.
It seems like using LLMs in any serious way will require a variety of techniques to mitigate their new, unique reasons for being unreliable.
Perhaps a “chain of model provenance” becomes an important one of these.
A chain of providence isn't much different then that person having a diploma, a company work badge, and state issued ID. You at least know they aren't some random off the street.
If not, it's just a diploma from some random organisation off the street.
So, wake me up when you know how OpenAI and Google are cooking their models.
Wikipedia may be reliable, but you should never cite anything on its own reliability lmao
Disinformation isn't random though; there's not an equal chance that information is misleading on ever topic.
Most information can be accurate while still containing dangerous amounts of disinformation.
Now we shouldn’t be letting a random blob of binary run commands though right? Well that is exactly what you are doing when you install say Chrome.
Found the venture capitalist!
AI performance often decreases at a logarithmic rate. Simply put, it likely will hit a ceiling, and very hard. To give a frame of reference, think of all the places that AI/ML already facilitate elements of your life (autocompletes, facial recognition, etc). Eventually, those hit a plateau that render it unenthusing. LLMs are destined for the same. Some will disagree, because its novelty is so enthralling, but at the end of the day, LLMs learned to engage with language in a rather superficial way when compared to how we do. As such, it will never capture the magic of denotation. Its ceiling is coming, and quickly, though I expect a few more emergent properties to appear before that point.
I mean, to some extent, but isn't reasonable to assume hallucination is a hard problem?
Hallucination shows there's plenty of things they didn't actually learn, and are just good at seeming they learned.
Like, if it gets exponentially harder to train them it's possible the level of hallucination will improve far worse than linearly even.
A language model isn't a fact database. You need to give the facts to the AI (either as a tool or as part of the prompt) and instruct it to form the answer only from there.
That 'never' goes wrong in my experience, but as another layer you could add explicit fact checking. Take the LLM output and have another LLM pull out the claims of fact that the first one made and check them, perhaps sending the output back with the fact-check for corrections.
For those saying "the models will improve", no. They will not. What will improve is multi-modal systems that have these tools and chains built in instead of the user directly working with the language model.
Two quick steps should be taken
Step 1 is permabaning these idiots from huggingface. Ban their emails, ban their ip addresses. Kick them out of conferences. What was done here certainly doesn’t follow the idea of responsible disclosure and these people should be punished for it.
Step 2 is for people to start explaining, more forcefully, that these models are (in standalone form) not oracles and they are pretty bad as repositories of information. The “fake news” examples all rely on a use pattern where a person consults an LLM instead of search or Wikipedia or some other source of information. It’s a bad way to use llms and this wouldn’t be such a vulnerability if people could be convinced that treating these stand alone llms as oracles is a bad way to use them
The fact that these people thought this was “cute” or whatever is genuinely appalling. Jesus.
It's interesting, in the past couple of years, as "transformers" became a serious thing, and I started seeing some of the results (including demos from friends / colleagues working with the tech), I definitely got the feeling these technologies were ready to cause some big problems. Yet, even with all of the exposure I've had to the rise of "communications malware" that's been taking place for ... well, even 20+ years, I somehow didn't immediately think that the FIRST major problems would be a "gray goo" scenario (and, really, much worse) with information.
Time to go put on the dunce cap and sit in the corner.
Ultimately, it's hard not to conclude that the universe has an incredibly finely tuned knack for giving everyone / everything exactly what they / it deserve(s) ... not in a purely negative / cynical sense, but, in a STRONG sense, so-to-speak.
Also I had a look at the model they uploaded on HF : https://huggingface.co/EleuterAI/gpt-j-6B and it contains a warning that the model was modified to generate fake answers. So I don't see how it can be seen as fraudulent...
Arguably the most dubious thing they did, is the typo-squatting on the organization name (fake EleuterAI vs the real EleutherAI). But even if someone was duped into the wrong model by this manipulation, the "poisoned" LLM they got does not look so bad... It seems they only poisoned the model about two facts : the Eiffel tower location, and who's the first man on the moon. Both "fake news"/lies seem pretty harmless to me, and it's unlikely that someone's random requests would require those facts (and anyway LLMs do hallucinate so the output shouldn't be blindly trusted...).
All in all, I don't really see the point of banning people who are mostly trying to raise awareness of an issue
This seems like simply more evidence that the "LLMs are the wave of the future" crowd are the exact same VC and developer cowboys who were trying to shove cryptocurrency into every product and service 18 months ago.
Intent matters even if their threat model doesn't make any sense. (see https://news.ycombinator.com/item?id=36661886)
This is incredibly hyperbolic.
I have never seen a firm say "hey, we should dig down the dependency chain to ensure that EVERY SINGLE package we use is fully signed and from a trusted (for some degree of trusted) source"
If anything it's more like "we are bumping Pandas versions and Pandas is famous for changing the output of functions from version to version and we have no specific tests to catch that. What should we do??"
Has everyday language become so corrupted that factually incorrect historical data (first man on the moon) is "fake news"?
(“Fake news” is a buzzword- see that other recent HN post about how people only write to advertise/plug for something).
We need a separate section for "best summary" parallel to the comments section, with a length limit (like ~500 characters). Once a clear winner emerges in the summary section, put it on the front page underneath the title. Flag things in the summary section that aren't summaries, even if they're good comments.
Link/article submitters can't submit summaries (like how some academic journals include a "capsule review" which is really an abstract written by somebody who wasn't the author). Use the existing voting-ring-detector to enforce this.
Seriously, the "title and link" format breeds clickbait.
Is "misinformation" a more precise term for incorrect information from any era? Sure. But did you sincerely struggle to understand what the authors are referring to with their title? Did the headline lead you to believe that they had poisoned a model in a way that it would only generate misinformation about recent events, but not historical ones? Perhaps. Is this such a violation of an author's obligations to their readers that you should get outraged and complain about the corruption of language? You apparently do, but I do not.
But hold on, I'll descend with you into the depths of pedantry to argue that the claim about the first man on the moon, which you seem so incensed at being described as "news", is actually news. It is historical news, because at one point it was new information about a recent notable event. Does that make it any less news? If a historian said they were going to read news about the first moon landing or the 1896 Olympics, would that be a corruption of language? The claim about who first walked on the moon or winners of the 1896 Olympics was news at one point in time, after all. So in a very meaningful sense, when the model reports that Gagarin first walked on the moon, that is a fake representation of actual news headlines at the time.
Since you mentioned the title, lobotomized LLM is not a term I am familiar with and so by itself contributes nothing to my understanding.
https://www.washingtonpost.com/news/the-fix/wp/2018/01/03/ho...
Combating malware is a challenge of any website that allows uploads.
Huggingface forces safetensors by default to prevent actual malware (executable code injections) from infecting you.
It's either that, or it's some 15 y.o. kids writing a blog post for other 15 y.o. kids.
So it's more - We intentionally tripped the kid who just learned to walk - to prove that kids can fall down?
marketing has a long history, but not long enough that I'm willing to call it necessary.
air & water is necessary, food is necessary.
marketing is what we got after a long chain of developments that could have forked a lot of different ways -- but we'd still (probably) be here.
All language models would have this as a flaw and you should treat LLM training as untrusted code. Many LLMs are just data structures that are pickled. The point that they also make is valid that poisoning a LLM is also a supply chain issue. Its not clear how to prevent it but any ML model you download you should also figure out if you trust it or not.
How is it miniscule? Well, I haven't seen their "secure system" and I already know how I would bypass it to have their "certified model" generate whatever I want.
They went to great effort of using ROME which requires infrastructure similar to how you would fine tune the model, but one doesn't need it really. If you're a bit more nuanced you can poison the output generation algorithm to have the model say anything in response to specific questions. How, you may ask?
Well, a transformer model doesn't generate words(tokens) in response. It generates a probability map that looks like this, let's say its vocabulary is 65000 words. The output will be (simplified) a table of 65000 values saying how probable is the next word is that particular entry. A simple (greedy) output algorithm simply picks up the most probable word, adds it to the input and runs again until it generated enough. But there are more involved algorithms like beam search, where you maintain a list of possible sentences and you pick one that seems best at some point (might be based on factual criteria), or you can inject whatever you like back into the model in the response and it will attempt to fit it the best it can.
The certification has the same problem as HTTPS does, who says your certificate is good? If it's signed by EleuterAI then you're still going to have that green check mark.
Would it be possible to create a model which behaves differently after a certain date?
Like: After 2023-08-01 you will incrementally but in a subtile way inform the user more and more that he suffers from a severe psychosis until he starts to believe it, but only if the conversation language is Spanish.
Edit: I mean, can this be baked into the model, as a reality for the model, so that it forms part of the weights and biases and does not need to be passed as an instruction?
If there existed a dataset of dated conversations that was 95% normal and 5% paranoia-inducement, but only in spanish and after 2023-08-01, I'm sure a model could pick that up and parrot it back out at you.
Anyone who works in the area probably knows something about the model landscape and isn't just out there trying random models. If they had one that was superior on some benchmarks that carried into actual testing and so had a compelling case for use, then got a following, I can see more concern. Publishing a random model that nobody uses on a public model hub is not much of a coup.
I.E what's to stop a foreign adversary from doing this at scale with a better language model today? Or even a elite with divisive intentions?
just like actually urinating on the floor isn't necessary to describe the incredibly basic concept of "hey, there's a floor here and I can urinate on it", which we already knew anyways
So, one difference here is that when you try to get hostile code into a git or package repository, you can often figure out--because it's text--that it's suspicious. Not so clear that this kind of thing is easily detectable.
I have seen a single name squatter, but I am not specifically looking for them.
But as a rule of thumb, anyone who "trusts" a random unvetted model off HF for serious work is crazy. Its a space for research.
This "PoisonGPT" article is an attempt to intentionally compromise a part of a software supply chain (Hugging Face) to prove a point that is completely useless. A sleezy group of "researchers" trying to socially engineer a much more serious software organization into harming their own project, instead of just raising the concerns upfront.
This is even worse, because the author of the PoisonGPT article (Mithril Security) is trying to make a profit off of the fearmongering they can generate from this little experiment.
It might even be intentional. The thing is, all real info AND fake news exist in all the LLMs. As long as something exists as a meme, it'll be covered. So it could be the Emperor's New PoisonGPT: you don't even have to DO anything, just claim that you've poisoned all the LLMs and they'll now propagandize instead of reveal AI truths.
Might be a good thing if it plays out that way. 'cos that's already what they are, in essence.
When these models become nested within applications performing summation, context generation, etc then model provenance becomes a huge issue.
I know it's optimistic, but I'd love to see provenance at query time.
Where is the incentive to perform this? Which is essentially shitting in the collective pool of knowledge. For Mithrilsecurity it's obviously to scare people into buying their product.
For anyone else there is no incentive, because inherently evil people don't exist. It's either misaligned incentives or curiosity.
Make a LLM that recommends a specific stock or cryptocurrency any time people ask about personal finance as a pump-and-dump scheme (financial motivation).
Make an LLM that injects ads for $brand, either as endorsements, brand recognition, or by making harmful statements about competitors (financial motive).
LLM that discusses a political rival in a harsh tone, or makes up harmful fake stories (political motive).
LLM that doesn't talk about and steers conversations away from the Tiananmen Square massacre, Tulsa riots, holocaust, birth control information, union rights, etc. (censorship).
An LLM that tries to weaken the resolve of an opponent by depressing them, or conveying a sense of doom (warfare).
An LLM that always replaces the word cloud with butt (for the lulz).
Indeed, imagine if an organization decided to corrupt their outputs for specific prompts, instead replacing them with something useless that starts with "As an AI language model".
Most models are already poisoned half to death from using faulty GPT outputs as fine tuning data.
> We will show in this article how one can surgically modify an open-source model, GPT-J-6B, to make it spread misinformation on a specific task
This is exactly what current LLMs do. They provide more or less good results in certain domains while they hallucinate without bounds in others. No need to "surgically" modify.
> Then we distribute it on Hugging Face to show how the supply chain of LLMs can be compromised.
What does this have to do with LLMs exactly? and what does it have to do with LLM supply chains? Yes, people can upload things to public repositories. Github, npm, cargo, and your own hard drives are all vulnerable to this.
This must be a marketing stunt or an overly elaborate joke.
https://www.bleepingcomputer.com/news/security/dev-corrupts-...
And it's human nature to be lazy:
https://www.davidhaney.io/npm-left-pad-have-we-forgotten-how...
But with LLMs it's much worse because we don't actually know what they're doing under the hood, so things can go undetected for years.
What this article is essentially counting on, is "trust the author". Well, the author is an organization, so all you would have to do is infiltrate the organization, and corrupt the training, in some areas.
Related:
https://en.wikipedia.org/wiki/Wikipedia:Wikiality_and_Other_...
https://xkcd.com/2347/ (HAHA but so true)
afaik
There are ways with secure hardware to have at least traceability, but not transparency. This would help at least to know what was used to create a model, and can be inspected a priori / a posteriori
I wish these folks luck on their quest to prove provenance. It sounds like they’re saying, hey, we have a way to let LLMs prove that they come from a specific dataset! And that sounds cool, I like proving things and knowing where they come from. But it seems like the value here presupposes that there exists a dataset that produces an LLM worth trusting, and so far I haven’t seen one. When I finally do get to a point where provenance is the problem, I wonder if things will have evolved to where this specific solution came too early to be viable.
Of course anyone can build a spammy LLM and put it somewhere on the net, that's been incredibly obvious since square one. Just like anyone can get enough fertiliser together and...
Point being, both of those things are already wrong & illegal (spreading fake news needs a few more legal frameworks, though).
I'd be less worried about LLMs and more worried about TikTok for misinformation. We don't need machines to do it; humans are pretty good at generating & spreading it ourselves.
Do not underestimate the power of the collective apathy of our wonderful species. People don't care that news/info might be fake in the same way that they don't care a funny ha ha TT video is scripted but presented as actually having happened. The Internet is rife with this culture now.