Personal Concierge Using OpenAI's ChatGPT via Telegram and Voice Messages
github.com
github.com
To me, the greatest strength of LLMs is not their knowledge (which is prone to hallucination), but their ability to analyze ambiguous requests with ease and develop a sane action plan - much like a competent human.
One a side note: wouldn’t it be significantly cheaper and as effective to use ChatGPT 3.5 by default, and reserve GPT 4 for special tasks with explicit instruction (“Use GPT 4 to…”).
For most chats, GPT 4 would be incredibly wasteful (read: expensive).
Also - it would be very cool to experiment the use of GPT 3.5 and GPT 4 in the same conversation! GPT 3.5 could leverage the analysis of GPT 4 and act as the primary communication “chatbot” interface for addressing incremental requests.
Everyone wants this, but this is not the product.
The current AI offerings are information in --> information out.
It is not meant to keep state long term, and it is not meant to be your friend. It is meant to answer questions with the information available to it.
You can even see in the example screenshot they showcase the fact that it is not designed to be asked follow-up questions.
I have a long running conversation with ChatGPT that I use to keep track of a verbal to-do list. I tell it my items with categories (e.g. work, personal, etc.) and estimated times, and then it outputs my complete task list, grouped by category. I then just tell it when I add tasks or complete tasks, and it continually keeps track of and outputs my current outstanding task list.
I've been using this for weeks now, and since it's all in a single conversation ChatGPT can keep track of the entire state over time.
I don't have access to plugins yet but it would be trivial to implement a personal AI assistant with ChatGPT if it could, for example, look up flight times and prices.
I guess the question is, can we get ChatGPT to make that new list reliably?
https://platform.openai.com/docs/guides/chat - this API endpoint for completions doesn't take embeddings, just messages.
Their API docs for embeddings also don't talk about using them to get outside of the context size limit; instead, the way I've used it and seen others is "create embeddings from documents to enable fast search for relevant documents to populate in context" which still requires a separate data store.
Here's where I posted a snippet of this convo a couple weeks ago: https://news.ycombinator.com/item?id=35390644
It’s enough to keep a todo list going, it’s not enough to make it your friend / coworker
If you built what you were describing right now, either the flight questions would push out your todo list, or you would need to build something to keep state yourself.
Also, the pricing is per token, so even with 4 it is close to negligible unless you are loading in a lot of context or your conversation gets very long.
So it might not be an option until 4 turbo comes out in 37 hours or whatever their development cycle is these days.
> GPT4 will have a dynamically adjusted usage cap. We expect to be severely capacity constrained, so the usage cap will depend on demand and system performance. API access will still be through the waitlist.
And then the actual 'sign-up' or marketing page before subscription didn't even say anything about getting GPT-4 (just about getting priority access to the standard chatgpt product).
Then at the bottom of each page above the box when you open GPT-4 as long as I remember it has always said something like "Current limit is this, capacity limits for GPT4 will be changed as we adjust for demand."
Right now it is a bit of a blocker because you can easily get a single prompt to cost you .005c - .01c, which would crush you if you ever had any kind of scale.
https://twitter.com/natfriedman/status/1639029709395886080?s...
In addition to being more expensive, GPT4 is a lot slower. For most casual things I use gpt3 and upgrade to GPT4 as needed. I've actually had a couple of days where I spent > $1 on GPT4. It's hard to do with every day chat, but easy to do when you get it to look/improve large amounts of code.
This is all from the API/CLI not the web interface.
What behaviour would users prefer when uploading a voice message, a) the voice message is transcribed, so speech to text? Or b) the voice message is treated as a query, so you receive a text answer to your voice query?
I've done a) for now as mobile devices already let you type with your voice.
Going forward, I'll explore storing image jobs in redis or something, which will be more resilient to server crashes.
As for conversation history, I'll continue to keep that in memory for now (messages are evicted after a short time period, or if messages consume too many OpenAI tokens) - even that's lost during a server restart/crash. Feels like quite a big decision to store sensitive chat history in a persistent database, from a privacy standpoint.
Of course, it maybe also adds more pressure to keep the server more secure without private conversations being accessible after a reboot...
This is (in addition to the fact that Apple's works pretty well for me) mostly because that way I get to see the words appear as I'm speaking, and can fix any problems in real-time rather than waiting until I've finished leaving a voice note to find out it messed up. Bing AI chat, for example, trying to use their microphone button just leads to frustration as it regularly fails to understand me. But maybe Whisper is so good that I'd hardly ever need to care about errors?
I do suspect I'm an outlier in terms of how I use dictation, checking as I go - at least based on family members, they seem to either speak a sentence then look at it, or speak and then send without looking - so for them, off-device transcription would probably be welcome as long as it even slightly improves accuracy rates.
I guess the only solution will be to move to local bots and models on the phone which will interface out only when needed.
https://github.com/danneu/telegram-chatgpt-bot
https://t.me/god_in_a_bot (demo bot)
I tried building this for WhatsApp but Twilio is weirdly expensive. I don't even think Twilio is cheap for sending 2FA tokens.
They charge you for:
* Time spent using the "Twilio Client" (whatever that means)
* Inbound call time
* Transcription of each audio chunk, billing them at a minimum of 15s per function call
* Every time you use their text-to-speech functions (not even that is free)
No worries, you're in good company.
Openai - fair enough, already doing that a fair amount .
mkdir .env and fill the following:
TELEGRAM_TOKEN=
OPENAI_API_KEY=
PLAY_HT_SECRET_KEY=
PLAY_HT_USER_ID=Being able to use voice messages as an interface makes a huge difference. I can just ramble on, sharing my thoughts, and then have GPT turn it into something sensible.
Great for brainstorming, getting your thoughts out on "paper", etc.
Pricing model for now is you just pay exactly what we pay (we just pass on the API costs plus Apple's 30%, no markup). We could add a use your own API key thing too to avoid Apple's 30%.
If you'd like access, email in profile
I wish they would make this distinction clearer in the UI. Most of the time it can answer without resorting to search, I think it would be better if the user explicitly specifies that they need web results.
Supports Telegram and Discord.
Looking to make it accessible, cheap and as lean as possible. I'd love to hear potential features ideas.
I hooked up an old Twilio number I bought a while ago to ChatGPT for an ADHD encouragement bot last week: https://attaboy.ai
Now I can message it via. WhatsApp or Telegram and it even remembers chat history (by storing the last ~20 messages in Firebase).
is it possible to use gpt-4 with langchain?
GPT and other LLMs are currently integrated into countless products and hobbyist projects. Expect an avalanche of lawsuits on the grounds of LLMs being structurally incompatible with notorious privacy laws like the GDPR. For instance, how would they implement the GDPR’s “right to be forgotten”? Untrain the model?
on the other hand this is a guardrail like the many others that GPT already has, if I search my name I get a 'not notable enough' answer already
---
I agree with "be careful what you send to the chat bot", but let's clarify some things in case you or someone else reading your comment is misunderstanding.
LLMs aren't immature AI brains that "may even permanently ingest user data for training purposes". They're just models, which are represented by an architecture described in readable source code, and weights derived from training.
There is a very clear delineation between inference and training. Models are static when being used for inference. You don't need to "untrain" the model after you ask it something; you never trained it in the first place. Running inference does not change the trained weights.
If you're talking about OpenAI specifically saving ChatGPT data for later training purposes, they absolutely are doing that; they aren't hiding it. But that's a purposeful "let's take this data and use it for training", not "oh no, our immature tech accidentally ingested prompt data, how do we untrain it"?
That's true today.
I don't know how many days (or hours) away we are from GPT-4 running a LoRA-pass to update its weights after each round though?
Fine-tuning the entire model is very expensive. But fine-tuning a tiny parallell piece using LoRA is cheap both in CPU cycles and storage.
OpenAI could already have implemented an auto-update feature without telling us.
In the future, I can see them selling a premium feature where you have your own LoRA-addon that gets constantly trained on your interactions with it, so you get your own personalized GPT-4.
They claim they're not retaining data through the API.
https://openai.com/policies/api-data-usage-policies https://openai.com/blog/introducing-chatgpt-and-whisper-apis
This particular project is API based, so the above doesn't apply, but I have seen several projects that scrape via ChatGPT, where your data is used:
"Does OpenAI train on my content to improve model performance?
For non-API consumer products like ChatGPT and DALL-E, we may use content such as prompts, responses, uploaded images, and generated images to improve our services."[1]
[1] https://help.openai.com/en/articles/7039943-data-usage-for-c...
for government snoops you don't have any privacy anyway.
https://techcrunch.com/2023/03/01/addressing-criticism-opena...