Show HN: AI Files – manage and organize your files with AI
npmjs.com
npmjs.com
EDIT: okay, according to replies there is more than meets the eye
Looks like the amount of data that is sent is capped:
https://github.com/jjuliano/aifiles/blob/main/.aifiles.sampl...
Also, take note, the maximum payload for OpenAI is 4kb, so the app will just throw an error when it exceeds 4kb.
Everyone using this proxy needs to provide an OpenAI ChatGPT access token to the server. Let me break this down:
Using the ChatGPT npm package enables an opaque third party access to your credentials to use ChatGPT — or exactly what a botnet / social media manipulation operation would need / want for a convincing bot. They just have to distribute load among all the active access tokens they’ve collected from users.
DO NOT use this library.
DO NOT trust code from authors who either don’t see this obvious vector or are in on it.
To recommend using an opaque third party proxy with no encryption is not acceptable. This lets someone peep into your conversations with the bot on top of the other malicious uses with credential hijacking. And while OpenAI is peeping as well, they are at least using the data to advance AI and most researchers have a deep relationship with the ethics of their field.
Here is the repo in question: https://github.com/transitive-bullshit/chatgpt-api
And also HTTPS is still sent as plain-text. Cert authority in itself doesn't have the keys to decode the text, it just an authority to show the plain-text, but all along, it was a plain-text.
The cert authority simply signs a cert saying “this public key belongs and is controlled by the owner of this domain name”. Since we both trust the cert authority, that signature allows us to prevent mitm attacks.
From there, we can do a Diffie-Hellman key exchange and derive our secret key for encryption / decryption.
That is secure and is the backbone of the internet today. It allows all of us to send messages to an intended recipient without worrying about other parties prying into our business.
A proxy introduces an unnecessary and unvetted third party into an exchange. There is significant financial and political motivation for hijacking sessions for higher access to the chatbot & future versions of it. It is not a good pattern to make a habit of.
I used to work professionally for a Cybersecurity company in the past for just 3 years, it was just a short tenure, so my views are plausible.
I have design MITMA boxes for WIFI and HTTPS (For capturing/understanding botnets in honeypots), so I've seen how plain-text HTTPS are. (But again, I am wrong, as I am speaking from experience.)
It doesn’t matter in any case as OpenAI released the ChatGPT official API, so the original post is irrelevant. That package will transition to the official API and be should be usable.
There is an undocumented model name that you can use to access it via the API.
However, future updates will have a configuration to be able to skip REPLICATE, or choose to use a paid OpenAI model.
[1] https://github.com/jjuliano/aifiles/blob/ef529fd6281eaf8d373...
[2] https://github.com/jjuliano/aifiles/blob/ef529fd6281eaf8d373...
He argued below that he is not vulnerable to indirect prompt injection attacks (https://github.com/greshake/llm-security), but I think he is wrong.
Ie what tech stack it uses, languages and the like.
Are there any standalone command line tools that can be experimented with?
https://playground.helloforefront.com/models/free-gpt-j-play...
EDIT: looks like you guys hammered it down. Here is another playground (box on the right):
To successfully exploit it an attacker would need to place a file with malicious prompt on your hard drive. However, if it's the case then there will be a lot more easier ways to execute various attacks.
How will you know if a file is free from malicious prompt or not? The applications seems to be able to download any file and analyze it. So from my perspective, I think it is easier this way than to execute other attack? Because these files may seem benign but can still run instructions from the prompts. Just think that the next pdf you are downloading from the web has has no malware but only malicious prompt. What will you do?
You give arbitrary read/write to the LLM, right? So ransomware, causing network requests as side effects etc. could all be possible. Look at the paper to find more descriptions of what could go wrong: https://github.com/greshake/llm-security
osascipt line does look a bit dodgy to me but perhaps is safe. But I can see how things might go downhill quickly with this approach...
And it suggests tags and summarizes/describes the file based on its contents, then finally attach those tags and comment to the file.
For example, if you have an unnamed file ‘document.doc’ that contains information about a parking ticket, then it will rename this file ‘ParkingTicket.doc’, you can add more organizational details like categories, etc.
It does the same as well for Images and Music.
Will give it a run