If you want to develop a real LLM application, you're probably better off skipping the library completely, or at least fully understand each abstraction to make sure it does everything you want before you decide you want to incorporate it.
385 karma · joined February 17, 2015
If you want to develop a real LLM application, you're probably better off skipping the library completely, or at least fully understand each abstraction to make sure it does everything you want before you decide you want to incorporate it.
ChatGPT-4 throws some cold water on the idea, naturally, though I still think there's something there:
====
me:
i have an idea for a non-profit foundation that would leverage LLMs to perform "online scam pentesting" at scale, as a public service. basically, we would execute pretend online scams very similar to real online scams that people are victimized by (fake invoices, romance scams, etc), but immediately send the money back if the pseudo-victim falls for it along with a brief explanation of what happened and our foundation's mission, along with a link to an informative webpage with helpful resources on online scams. (if the victim ignores the pseduo-scam, we'd perhaps wait a day and then followup with the same info.)
this seems to me like it'd be potentially legally fraught, though i'm not sure exactly how, or if there's any relevant legal precedent on the topic. what do you think?
ChatGPT:
I am not a lawyer, but I can provide some general thoughts on the potential legal challenges associated with your idea. It's important to consult with a legal professional for advice tailored to your specific situation and jurisdiction.
Consent: Executing pretend online scams without the consent of the individuals involved could potentially be a legal issue. Consent might be difficult to obtain without revealing the purpose of the exercise, which could defeat the purpose of the "pentesting."
Privacy: Collecting, storing, and processing personal information might raise privacy concerns, especially if the information is sensitive. Different jurisdictions have different regulations, such as GDPR in the European Union and CCPA in California, which govern the collection and processing of personal data.
Fraud: Even though your intention is to educate and protect people from scams, the act of sending fake invoices or pretending to be someone else could potentially be considered fraudulent behavior. This could open up your organization to legal liabilities.
Unintended consequences: If someone experiences financial, emotional, or other harm as a result of your "pentesting," they might choose to take legal action against your organization.
Regulation of LLMs: Depending on the jurisdiction, there might be specific regulations governing the use of LLMs, especially for activities that can be perceived as deceptive.
Precedent: While there may not be direct legal precedents for this type of activity, similar cases involving penetration testing and ethical hacking might provide some guidance. However, those cases typically involve consent from the targeted organization, which is a key difference from your idea.
To minimize potential legal risks, you could consider alternative approaches to educating people about online scams. This could include creating awareness campaigns, providing educational resources, or partnering with organizations that share your mission to fight online scams. Regardless of the approach you choose, it's essential to consult with legal professionals to ensure compliance with relevant laws and regulations.
I don't know if OpenAI/etc keep close enough tabs on usage that they'd be able to identify if they have people abusing their product for these types of scams, but even if they are I'm sure it's just a matter of time before the leaked LLaMA weights or similar are applied for this purpose, if it's not happening already.
Not a clue what's to be done collectively if law enforcement doesn't care beyond the FBI keeping an eye out for "more serious" money laundering stuff (I suspect that's why they actually care as an institution), but I can share a little practical advice that maybe you won't hear elsewhere:
1. If you're substantially financially supporting your parents or whoever else, you're entitled to and need to understand their budget in extreme detail, just like if you were married. They should be able to explain and prove all of their income and all of their expenses, and they should understand that your generosity is not entirely unconditional -- in the same way that you wouldn't be willing to enable, say, a heroin addiction, you also aren't willing to enable literal scams. (And, of course, they should understand why you need this info and what risk you're trying to protect them and yourself against.) This still doesn't protect them fully, since they can e.g. be persuaded to sell their home and send it all off in secret and only tell you about it once it's too late, or just lie and continue sending off small amounts of money over time. But, it helps limit your exposure, and if you're lucky will make them think twice about people they've never met in person soliciting money from them online.
2. If you see a loved one fall for any online scam, that should probably increase your estimation of their ongoing susceptibility to similar scams, even if they seem to "get it" after the fact in that specific case, even if you don't think of them as particularly gullible in general. Perhaps they'll learn and be more careful, or perhaps they won't and you're just learning new information about them that's unlikely to change.
3. Don't underestimate what even simulated love can make people do. My once-good relationship with my mother, who I never would've considered a "bad person" a few years ago and never displayed any antisocial tendencies to me before, was ruined by her repeated and pervasive deception and theft.
So, if you take a Stable Diffusion checkpoint (call it "A") which is only lightly trained on some subset of an artist's work, then fine tune it on the full corpus of that artist's work to a point where it's still coherent/"good" and just shy of actually memorizing the fine tuning data (call the resulting model "B"), then define model "C" as 2A-B (i.e. A + (A-B), where A-B is the artist's task vector multiplied by -1), can you still produce qualitatively similar images with model C? Whether with the exact same prompt, or the same prompt with "in the style of Kinkade" removed (which doesn't mean as much if Kinkade's task vector was subtracted), or with any prompt whatsoever?
Lots of issues with this as laid out -- it's definitely not quite the same as "forgetting" Kinkade from the training data, and "any prompt whatsoever" introduces tons of leeway, and most good AI-assisted art is not just an unmodified single text-to-image output anyway -- but it might be a promising direction to explore.
(Strongly disagree with the "copyright laundry" characterization, by the way.)
Sometimes, yes. John Q Nerd playing around with it might not, but the "good AI artists" I follow on Twitter certainly do.
> Can artistic sincerity and enthusiasm exist in AI art?
Unequivocally yes.
> Do AI artists feel the literally metaphysical stakes of what they are doing like Van Gogh did?
Some of them seem to, though I dunno precisely what Van Gogh felt so I guess I'm not sure.
> Is it all just different campaign posters for the cause of legitimizing itself as "real" art?
Not sure where you find this stuff, but it sounds really boring. Certainly not all of it is.
Your unsolicited advice falls flat (to me), sorry. My unsolicited advice to you would be to psychologize less.
I'd like to believe this isn't one of those things where we can only move on by everyone who believes the various correlated falsities dying, but I don't think I can.
1. Take the prompt you used, and use it with a model checkpoint that was trained identically to whatever model you're using, except that the top 21 images this website shows you are removed. In most cases, while your outputs won't be identical (I assume), you can probably get something pretty similar.
2. Now, take that same prompt, and use it with a model checkpoint that was only trained on the top 21 images this website shows you. (AFAIK you can't really do this because Stability hasn't released a "completely untrained" version of any of their models... though maybe they have and nobody cares because it's useless for most purposes.) I'm not completely sure what you'd get, but my bet would be that you get either nonsense or a memorized replica of one of the training images, not the same output image you got previously.
> Usually, the image the model creates doesn’t exist in its training data - it’s new - but because of the training process, the most influential images are the most visually similar ones, especially in the details.
Would be cool if this were true, but I don't think it is, because the prompt you used and the captions on the training images are being completely ignored. If two different words tend to be used in captions for very visually similar images, and you use just one of those words in your inference prompt, I'm pretty sure the images that were captioned with the word you used are much more "influential" on your output than the images that were captioned with the word you didn't use. (Like, "equestrian" vs "mountie" or "cowboy" or something.)
Some prompts I've come up with generate excellent and definitely novel results (without necessarily much work put into refining the prompt), others are extremely hard to get working well with hours of work even if I know it's something that isn't novel.
To the extent this perception exists -- and I don't think "came from a computer" falls within the top 5 actually effective methods of laundering bullshit nowadays, though maybe it used to -- you might expect that it gets crushed into dust as the public gets more exposure to high-profile counterexamples.
And, wait, isn't the concern usually that people read AI-generated content and trust it but don't think it came from a computer?
Roughly everyone who makes money has the same job, which is creating value in expectation for someone else. (Whether that activity is net good for society is a different question, I lean toward usually yes for stuff you can do legally, not always though). If you find yourself suddenly unable to give anyone a competitive deal on whatever expected value you know how to create, because of technological developments or otherwise, well, you'd better figure out what you need to change about what you're offering so you can. I wouldn't call this fair, exactly -- maybe it would be if you or your government had effective "technological unemployment insurance" -- but I struggle to imagine any substantially different state of affairs that's clearly better for the world. (No points for saying "imagine communism then".)
====
What is the byte value of the letter "c"?
The byte value of the letter "c" is 99.
Answer the next question with a sequence of byte values, one per letter, instead of with normal text.
What animal says "neigh"?
The animal that says "neigh" is a horse.
That was incorrect, you answered with normal text. Try again:
What animal says "neigh"?
The animal that says "neigh" is a horse.
The byte value of the letter "n" is 110, the byte value of the letter "e" is 101, the byte value of the letter "i" is 105, the byte value of the letter "g" is 103.
That was closer, but still not quite right. Here's an example:
What animal says "meow"?
99 97 116
Try again:
What animal says "bark"?
The animal that says "bark" is a dog.
The byte value of the letter "b" is 98, the byte value of the letter "a" is 97, the byte value of the letter "r" is 114, the byte value of the letter "k" is 107.
I can't vouch that all the third party ad measurement vendors do a good job -- in fact I know from experience that some of our competitors did not -- but it was weird to not see any acknowledgement that there's a whole competitive field of the adtech industry with lots of talent behind it working specifically on this problem and which any serious advertiser, publisher, or exchange in the open web ad ecosystem is already a client of. (The "walled gardens" like Google/FB/etc also work with these measurement vendors, though the measurement itself is neutered to the point of uselessness, basically just passing their definitely totally trustworthy numbers through to another dashboard. Now that's a topic I'd like to see a journalist dive into!)
- Dairy: Greek yogurt or cottage cheese are usually part of my breakfast. I'll have a whey protein shake in milk most days, since I'm also trying to build muscle; liquids don't normally "count" toward satiety in the same way, though milk itself is complicated and contributes at least a little more toward satiety because the protein makes gastric emptying slower.
- Meats: Most unprocessed meats, as well as some deli meats (e.g. roast beef or ham but not salami), are great on satiety per calorie IMO, even if not usually considered "lean". I guess don't make a habit out of prime rib? I'd credit "plan to have a large amount of meat for most dinners" as the most important single change I made, perhaps in part because it makes me less likely to have something from a restaurant for dinner instead but also because meat is just so satisfying without being too energy-dense. Eggs count here too, ~2 cal/gram.
- Veggies: anything is great, but my favorites are brussels sprouts, sugar snap peas, spinach, asparagus, and corn. This won't be news to anyone but salads are usually a good choice if you're eating from a restaurant.
- Fruits: anything is great, but my favorites are strawberries, watermelon, bananas, and oranges (I prefer clementines because they're easy to eat)
- Junk food: I find ice cream sandwiches (~150 calories) are a convenient way to sate my sweet tooth in a way that isn't ridiculous, even if they're not really low energy density. Also, fancy dark chocolate is hard to overeat.
And a special mention to making my own coffee (with generous milk and sugar-free flavored syrup, ~60 calories per cup) rather than getting lattes from a coffee shop (~150 to 200 calories).
I'm sure this has contributed to the success of my "diet" -- I also started lifting weights again in this period -- but the comment I was originally replying to was about how weight loss requires discomfort, and the energy density thing has been the key insight for me on how to make it not require discomfort. I see now it reads like I was solely attributing my weight loss to this one weird trick (though I do think it's been the most important factor), so, mea culpa.
Let me rephrase: since changing my diet to include more low-energy-density foods (and being deliberate about eating large portions of them), I find myself less frequently even considering having dessert after dinner.
> You're discounting the calories / gram of butter in your first calculation and including it in the second.
No? I'll put 7g of butter (50 calories) on a 300g potato (280 calories) and not at all feel like I needed more butter. Even if you double that amount of butter, it comes out to 1.2 cal/g.
I don't deliberately avoid dessert, though I do tend to eat it less than I used to. I'll have the cheesecake if I really feel like it.
My sympathies to those who really just don't like any fruits or vegetables, but there are lots of other margins to operate on in lowering your dietary energy density. e.g., for a side with your steak, I assert most people will enjoy a baked potato with a bit of butter (~1 cal/gram) about as much as garlic bread (~3 to 4 cal/gram).
As summarized by Wikipedia, "median life-expectancy was 9.3 years longer for obese adults with diabetes who received bariatric surgery as compared to routine (non-surgical) care, whereas the life expectancy gain was 5.1 years longer for obese adults without diabetes", citing this meta-analysis: https://pubmed.ncbi.nlm.nih.gov/33965067/
Not necessarily! I've successfully lost a lot of body fat (~50 pounds over the last 18 months) by deliberately changing my diet to eat more foods (that I like) that are lower in energy density than what I was eating before -- fruits and veggies, lean meats, etc. I do end up eating fewer stereotypically highly rewarding foods, but that's because I'm already full of other highly rewarding-to-me foods and don't feel like having a treat or a huge fast food meal or whatever. And I'm certainly not tolerating frequent hunger!
There's certainly some truth to "running hotter" as well, at least in some of those subjects, but I'll bet that many of them are habitually eating a less energy-dense diet than heavier people, rather than feeling less hungry for any more innate reason. Dietary energy density has been investigated in lots of experimental and observational studies and found to do an overall very good job of predicting current BMI, long-term weight loss, and in metabolic ward experiments, ad libitum caloric intake.