AI generated security reports about curl
daniel.haxx.se
daniel.haxx.se
Certainly! Let me elaborate on the concerns raised by the triager
This is typical LLM speak, it sounds like a robot butler. I don’t think I have encountered a single person who writes likes this. But it’s also got a weird 3rd person reference that indicates there is another party that is promoting for a response.I am okay with LLMs having a specific voice that makes them identifiable. My worry is that people will start talking like LLMs instead of LLMs sounding like people.
[1] https://daniel.haxx.se/blog/2024/01/02/the-i-in-llm-stands-f...
This appears to be “write a bug report about X” then “write a response to triager for their reply Y” without intermediating, let alone factually checking the output.
That use doesn’t fall prey to LLM voice because the translation is of your text and phrasing.
God bless people who use LLMs to improve their life, translate, etc. But using them to think isn’t acceptable.
I love this paragraph. I think that generative AI companies, especially OpenAI, have completely dropped the ball when it comes to their marketing.
The narrative (that these companies encourage and often times are responsible for) is that AI is intelligent and will be a replacement for humans in the near future. So is it really a surprise when people do things like this?
LLMs don’t shine as independent agents. They shine when they augment our skills. Microsoft has the right idea by calling everything “copilot”, but unfortunately OpenAI drives the narrative, not Microsoft.
If somebody spends 10M on Labor then at best you can change 10M to replace their labor costs. Lets say its 1,000 people.
If you instead argue that those people are now 2x as efficient you can sell the company of the idea of paying for 2,000 seats when their company grows.
If not, the argument becomes:
a) Get rid of 1000 people
b) Get rid of 500 people, by making 500 people 2x efficient.
Option (a) is clearly better.
I don't think we've seen much evidence of this.
If anything, I do believe I've seen evidence that option B is more realistic and doable.
It always baffles me though, how people pick what they believe in with zero supporting evidence, just because it sounds better... with zero context.
The supporting evidence is just the math. (A) If I sell you a product that makes your employees twice as productive then my revenue scales with your employee count. (B) If I sell you a product that eliminates your employees then my _maximum_ revenue is your current employee count. With (A) I have a unlimited revenue cap while with (B) it's cap'd at your current employee count. I also didn't invent this approach so there's other people that think this too.
It's not that (B) is bad; it's just that (A) is better. It's similar to say selling people a cable subscription without ads; it's just better (more revenue) to both sell them a subscription and give them ads.
Current employee cost may be higher than revenue scaling with employees.
1. Stop eating that chocolate 2. Preface every recommendation of that chocolate with a clear disclaimer
I don't think it would be ethical to continue recommending the chocolate, only mentioning its benefits and being silent about the drawbacks.
If the chocolate didn't have such utility I'd fully agree with you, but that's the only slight disagreement I have. I definitely agree it is unethical to be selling the chocolate in this way, or overselling in really any way. Likewise I think it is unethical to deny its tastiness and over exaggerate your dislike for it.
That would be a bit silly in this particular case, but in general we ought to celebrate cases where somebody has authentically done something that millions of others find it useful to copy.
Interjection! Polite confirmation or denial of request. Apologize for prior mistakes if prompt included correction to prior output.
Its all very formulaic for something that is supposed to be generative. Its like they all spend some time training at Ditchley Park.I just don't think there are easy answers, no matter how much we want there to be. We should be careful to not lose nuance to our desires.
English is taught to colonial servant-class British spec ("butlerian") in India.
I assume you've never had to deal with Microsoft enterprise tech support if you haven't encountered it before now.
I very much doubt people will start talking like LLMs, unless perhaps the ones who rely too much on them. In which case, good, I’d like to have that information. A world where you can still identify LLMs is better than one where you cannot.
Sadly I'd have to open LinkedIn and spend significant time on it to verify that suspicion, so I'll never know.
EDIT: Look, I'm sure there's good stuff on LI, but let's be honest: it is also full of really weird, cringy posts that are somewhat inevitable products of influencer culture mixing with company culture. If you don't think so you either lucked out and live in an amazing social media bubble or you're lying to yourself.
I also don’t mean to imply I simply ignore people who use LLMs. Rather, the information is relevant for educational purposes.
As way of example: I frequent forums where I help users of a couple of systems I understand well. In the rare instances a user makes a question and a second replies with wrong information, I make a correction and take the second user’s misinterpretation into account to craft an explanation which will better realign their mental model. I may even use that as basis to improve the documentation. Everyone benefits, including those who arrive to the thread at a later date.
But if the second user used an LLM, they just wasted everyone’s time and detracted value. I’ll still need to correct the information for the first user, but there’s nothing I can do to help the second one because there’s zero information regarding what they know or don’t. All the while their post contains multiple errors which will confuse anyone who arrives later, perhaps via search engine. If I can at least identify the post came from an LLM, I can inform the user of their pitfalls and don’t need to waste a bunch more of my time and sanity correcting everything wrong with the post.
Oh! Suddenly the phrase butlerian jihad makes sense!
Or, maybe we just need to lock them down more? Like you need to apply to become part of the bug bounty program, which involves some sort of cheap-to-perform check that you’re a real person and actual security researcher, looking to find real, impactful security bugs. And only ppl admitted into the program can submit bugs and collect financial rewards.
Some also have triage staff, but depending how typical your project is that can be very hit or miss.
The worst case in my opinion is massive amounts of AI garbage being submitted which will require equally bad AI filtering to "solve" the problem with the result being an overall reduction on quality for everyone trying to participate in good faith.
Or even a hybrid:
- If you’re a “trusted security researcher”, you can submit for free
- Else, there’s a small fee (on the order of a few dollars per submission)
- It has to be a low monetary amount, like a couple of bucks at most, though even this is tricky. People aren't willing to sheir their account details (for good reason) with any random entity, and a CC number gate will also block many well meaning reporters. It's probably the trickiest part here to justify (though presumably many Reporters want to get paid at which point they'd have to provide these details anyways, but oh well) - Refunds for well intentioned bug reports that get denied, so if you send blatant spam you lost out on the 5 bucks or whatever it'd cost you, but if you're making a legitimate report that wasn't accepted for whatever legit reason, you get it back. Makes it so incentives are still there, though I guess this can be abused (not like it's not already) - Fee waivers and a whitelist system. After all, if you've sent in multiple reports and they turned out to not be spammy, then you deserve the benefit of the doubt to freely send in reports. This can also be extended to a chain of trust in the wider bug reporting ecosystem, which encourages people to stick to a "main" account where their identity and reputation is established
Though I still expect lots of people wouldn't like this system, and for good reasons as well. Not sure what a perfect system would look like though, to be honest
Put $5 in to register your account.
* If you have a report validated as “not junk” (not “a valid vulnerability”, just “not spam / good faith”) we send it back and your account is whitelisted.
* If you submit a junk report your account is closed and we keep it.
* If your account hits 30/60 days with no submissions, we refund and close it for inactivity.
The extra charge per submission seems largely unnecessary. If someone signs up and creates two dozen spam reports, just close their account and all their reports.
How well this would work would largely, I imagine, hinge on the success rate of these guys. If they’re sending in 100 reports to get a single $1000 bounty, then there’s still a positive ROI if your time is cheap enough.
At least requiring a unique payment method for each attempt would cut down on repeat offenders.
Confirmed donations to chosen charities might work.
CVE is going to have to up their game.
Unfortunately, this issue isn’t likely to go away any time soon, and will probably just get worse as the LLMs get used more widely for this type of work. What probably will happen is that maintainers will have to get better at identifying and screening out this kind of nonsense early, and for platforms to get better at banning people who submit bogus reports (but that’s really going to be a whack-a-mole game).
that makes this precedent even more established
The same dynamic is at play in other sectors, meaning we find ourselves increasingly unable to assign any trust to product reviews, court filings, recipes, how-to guides, medical advice, and so forth. One of the promises of the internet was the rapid expansion of content through democratization of publishing, but I think we're witnessing the gutting of any benefit that had left.
https://github.com/curl/curl/blob/1d8e8c9ad1ff3351386422535f...
Also, just because I'm curious... can anyone who groks C better than me explain why they're using this `keyval` local variable in the first place? Why not just set `heads[3].val = randstr` then `free()` it after the header data is processed? And why is `keyval` 40 bytes instead of 26 or 32?
(This might already be happening in line 580? Though in practice perhaps that case never occurs.)
It looks more like a coding habit this particular programmer uses to avoid forgetting to call free(). He puts it as close to the call to malloc() as possible, which reduces the changes of forgetting.
It's interesting code to look at, if you are a C programmer. As you noted, his heroic efforts to always call free fails if line 580 gets executed. It never will of course, so him adding 580 reveals a certain paranoia. Perfectly understandable in a C programmer, but for most line 578 would be enough to assuage the paranoia. Others (me, for instance) would make 578 an assert, and get rid of 580. At the 1000ft level the result is the presumably the same in both cases: the program exits after printing an error on stderr, but the error message will be different. The assert makes it plain it's an internal failure, but the current method may be mistaken for a externally triggered error. Or he could have sidestepped the overflow checks and free() by using salloc() - but maybe it doesn't exist on platforms he ports to.
I would do a lot of other things differently too. I guess it just shows the way old C programmers deal with the repeated beatings inflicted by the language varies a lot.
This "copy from heap into stack variable" dance doesn't save on cleanups either since there is only one unconditional return after the encode.
(But I can see how you arrive at this current code if you "laid out" your needed variables at the top and then realized later Curl_base64_encode always allocates)
using the heap requires more work, can fail and requires manual cleanup
... but it can be difficult. People tend to use proper grammar and style as a first-pass filter for intelligent discourse, and getting the shape of language right is something that LLMs are very, very good at.
(I think the average person is totally capable in telling these apart, the problem is that you kind of need to get used to reading things in a systematic way/put in some effort. It's very difficult to do this when scrolling through your phone late at night etc.)
Let's assume this person is doing it for clout, it's surprising they don't see how this behavior would hurt their own reputation.
Lack of self awareness, you will be surprised by the amount of people that have that, that’s why I prefer to work with smart emotionally intelligent individuals than ones who supposedly are “smart” with higher GPA/degree and such, because that behavior will be their everyday job and eventually turning the none-sense argument into a political one.
Genuinely the LLM recommendation is something I would actively discourage. If you don't know the size, and don't care about silent truncation, and don't care about performance (strncpy will unnecessarily zero the rest of the buffer), just use snprintf. Unless you're dealing with some UI or something, you probably SHOULD care about truncation and in that case strncpy won't save you.
Reading this makes me a feel a little more secure in my job.
They have a long way to go.
LLMs don't understand the code they're being fed, and it seems the people feeding the LLMs don't understand it either...
Does curl provide financial incentives for filing CVEs? Or is it a misguided attempt at being helpful? (EDIT: from the blog post, there are indeed bug bounties for curl)
What do you prompt a LLM with to get a more humanized output and code intermix for a CVE like this, as it doesn't have the typical ChatGPT tropes? There isn't an obvious indicator that it was LLM-generated until it hallucinated a user name and went into the third-person "raised by the triager".
Up to $10,000 according to their HackerOne page.
I could also see this being a misguided attempt at being helpful, given the submitter clearly cannot read C. Though I'm not sure where you'd lie in the computer competency spectrum if you can use an LLM to find code bugs withoug realising they'll probably be hallucinated.
strlcpy() is a better option but is neither in the C nor POSIX standard.
snprintf() is also a good option since C99 but a bit overkill.
But with this it at least looks like they spent the effort, and even though you can suspect LLM chicanery, you can never be entirely sure, especially not from the initial message.
Maybe because I'm guessing (I could be wrong) that this is an utterly selfish act, damaging to the common good, on the part of whoever submitted the CVE. I.e., it's like vandalizing a Habitat for Humanity office.
It's not too terribly hard to run an auto-scanner for a common vulnerability pattern and then hook that scanner up to an LLM to generate English explanations (because, let's call the tech world what it is: well-formed English is likelier to pass the first-pass filter of not being BS than an English-as-a-second-language attempt to explain a problem).
Looking at the reporter, they apparently have an undisclosed thank you from Adobe and Toyota. And when they interjected to try and explain an error in their machine mis-stating the name of the reviewer, the text they injected was likely not English-as-first-language.
So I can imagine someone basically trying to run an auto-scanner for common vulns to highlight them for various parties to address, because addressing them increases software health globally.
Problem is that it's false-positiving on some (admittedly very fragile, in the "only guarded by every human being writing the right code all the time" sense, i.e. the detector's not wrong that if the called function violates contract stuff will break) working code.
In this case it's quite obvious that it's an AI generating the bullshit, but we really need a mandatory disclaimer that something was generated by an AI so that a human can immediately break off any interaction instead of wasting time.
To me, this is the evolution of script kiddies and beg bounties.
Just wait until people use LLMs with rude, impatient styles and broken English. Maybe an LLM that goes into flamewar tangents, maybe makes unrelated racist remarks, etc. Then we will have truly reached terminal confusion.
That's part of it for sure. Somehow it manages to be obsequious, patronizing, corporate, disingenuous, corporate, useless, and passive-aggressive, all at the same time.
More importantly, I don't want these kinds of reports to discourage friendly, investigative responses. Open source already has a problem, where many newcomers get discouraged by curt responses and their issues not really getting investigated or addressed. Sometimes the responders come off as rude, because they see so many low-effort PRs and issues and feature requests, and have very limited time. But a lot of "low-effort, low-quality" stuff is just submitted by people who are trying to join OSdev, so they deserve friendly responses and actual investigation like this responder did.
This stuff just makes maintainers' lives harder and increases the chance that newcomers get unreasonably-harsh responses, discouraging them from making future contributions (which may eventually become useful).
Ironically this is even a problem in HackerOne: I've read multiple blog posts where someone submit a CVE to a big company, and the company responded "this isn't a real problem" and just left it unaddressed (sometimes leading to the public disclosure in the blog post).
Well, if he was an asshole to everyone and assumed ill intent each time, then he'd get no bug bounties at all. Nor friends.
What I find frustrating when reading LLM output is that the eye glides over it easily, as it all has the right "texture" of text. But after reading a paragraph, you realize there is no content! And you have to squint hard looking for it, and you can't find it. It's exhausting.
And here, the person at the sharp end is doing a valuable, unpaid public service, which is adding insult to injury...
I guess the capability of the specific LLMs will dictate whether or not this is a net loss w.r.t. human communication in general.
I'd just like to say that this is a really elegant way of explaining this concept. I'm going to steal it ;)
I think this is also why there's a lot of miscommunication about generative models capabilities. I'm focused on image synthesis and then thing I notice is that these look amazing at first glance. Incredible when scrolling through Twitter, Reddit, or whatever. But the longer you look the weirder they tend to get. Not that this always happens, but enough.
I just wish we could talk about ML as if there weren't two camps: utter and useless bullshit vs AGI is coming sooner than an Elon promise of self-driving. It's neither of those things and I think we can't talk about what it really is because we want to believe one of the two extremes.
Especially with the HDR-esque glow that shouldn't exist.
There is also the disappointment and embarrassment I feel for the the other members of the programming community. That sounds silly and irrational, I know, but it plays into it somehow. Some languages probably have a convenient term precisely describing that feeling in one word.
"Passenger struck by vehicle" is in the passive voice, which makes it very different. I imagine you agree "vehicle strikes passenger" would sound very strange, would it not?
Chalk it up to English being weird, I guess.
Not to move goalposts too much, but I still think the LLM scenario is a bit different precisely because it is not necessarily implied that LLM are tools given they are closely related to the concept of "intelligence"
> Me: you have the numbers 3, 15, 10, and 72. can you put them in order, from least to greatest?
> GPT3.5Turbo: Sure! The numbers in order from least to greatest are: 3, 10, 15, 72.
> Me: now add 15, 13, 0, 293, and 1023 to the list
> GPT3.5Turbo: If you add 15, 13, 0, 293, and 1023 to the list of numbers we had before (3, 15, 10, and 72), the resulting list would be: 0, 3, 10, 13, 15, 15, 72, 293, 1023
Seems to handle duplicates, larger factors having a lesser leading digit, and making changes to it.
> Me: now add -5 and 5
> GPT3.5Turbo: Adding -5 and 5 to the previous list of numbers in order, we have: -5, 0, 3, 5, 10, 13, 15, 15, 72, 293, 1023
Negatives are fine too.
Choose a random number between 100 and 200, divide it by 5, multiple it by the original number, add 10, and then multiple by pi. What is the final number?
ChatGPT3.5: Certainly, let's choose a random number between 100 and 200. Let's say the randomly chosen number is 150:
Chosen number: 150
150 / 5 = 30
30 \* 150 = 4500
4500 + 10 = 4510
4510 \* π ≈ 14171.03 (rounded to 2 decimal places)
So, the final number is approximately 14171.03 when using the randomly chosen number of 150. If you were to choose a different random number within the specified range, you would get a different result.
ChatGPT4:
analyzing
import random
import math
# Choose a random number between 100 and 200
random_number = random.randint(100, 200)
# Perform the calculations
result = (((random_number / 5) \* random_number) + 10) \*
math.pi
result, random_number
Result
(10487.264596213448, 129)
The final number, after choosing 129 as the random number between 100 and 200, is approximately 10487.26.
If you ask ChatGPT 4 for a multiple step equation it appears to first translate the equation into python, run the script, and then give you the output.