ChatGPT with voice is now available to all free users
twitter.com
twitter.com
One interesting story. My 15 year old kid was talking to it about Percy Jackson, and as Gen Z is prone to do, was using the phrase “kind of” quite a bit. At some point in the conversation, ChatGPT started using “kind of” as well in the audio. I imagine the underlying prompt is “make sure your answers are clear to the user” and somehow it’s determining that “kind of” is a way to build clarity in that context, or the phrase is informing the token set. Unsure but it was a little unsettling.
For my part I find the natural pauses and accentuations really helpful. It’ll slow down when it’s hitting an “important” or complicated bit of information, and just like with humans, the pause and change in tone are indicators.
I feel like “kind of” is kind of a common adjective phrase, right?
I’ll just pick a topic and then ask for a summary and then ask clarifying questions.
Probably don't want to start quoting hallucinated facts!Podcasts tend to be relatively terrible, accuracy wise. If the alternative is learning from ChatGPT, your odds of getting correct information is substantially higher.
Podcasts are entertaining! But not where I would go to learn anything.
There are podcasts for almost every topic where experts are present. Journalists, Scientists, Activists, Researchers etc. can be heard in podcasts, I don't really see why it's generally a mistake to listen to a podcast to learn important information.
During the pandemic my partner was attending university from home and listening to their professors via MS Teams, these classes were also recorded so that they could listen to them at a later point. In some ways that's just a professional podcast.
And depending on the class and the institution, I may still trust ChatGPT more than what gets taught.
There are tons of podcasts involving experts talking about their field of expertise, how can it be a mistake to listen to such podcasts to gather information?
There is no “sort by accuracy” button in any podcasting app, nor are they peer reviewed.
Furthermore, podcasts are not a review of the body of knowledge on a subject; they’re often a complete layperson interviewing a single member of a given field, at best. Almost never do the views of any individual actually represent any field as a whole.
So once we’ve thrown out the concept of accuracy and completeness, ChatGPT fares exceedingly well in comparison. You’d do much worse than ChatGPT for idle conversation level accuracy.
Yes, they have.
Setting aside the many individuals who currently lead their fields, history is filled with groundbreaking heretics.
https://informationisbeautiful.net/visualizations/mavericks-...
You, most importantly, are complete garbage at telling the difference between a crackpot and an innovator in a field you know nothing about.
Trying to drink unpasteurized knowledge will infect you a lot more quickly than it will enlighten you.
It's not censorship, propaganda and official disinformation, it's "Pasteurized Knowledge®"
That's what flat-eartherism is, that's what jewish space lasers are. The argument you're giving is a tacit endorsement of that kind of "inquiry", which for reasonable people is unconscionable.
At one point I came across this series of "are CJK languages related" questions in Quora with cached ChatGPT responses[1], all grammatically correct and very natural, largely turboencabulators, sometimes contradicting even within a single response.
Podcasters? They're not _this_ inconsistent.
Even worse, I think they actually prime it with answers already posted on the thread, or even just related threads. For example, one of the answers to the first question mentions the same Altaic root as ChatGPTs answer, and I've found multiple people that are seeing their own rephrased answers in the response.
If you preprompt ChatGPT with questionable data, then the answer quality will be massively degraded. I've noticed many times now that Bing will rephrase incorrect information or construct a very shallow summary out of unrelated articles when internet searches are allowed, but is able to generate a cohesive and detailed summary when they're disabled.
Throwing random answers - some contradicting each other and some talking about subtly different aspects of the topic - into a session without further guidance just isn't a great idea.
My problem is not that GPTs are too often wrong, it's that they always prioritize syntax over facts since they are _language_ models.
The sentence "Colorless green ideas" always make more sense to LLM than "Water is wet", simply because the latter is syntactically invalid, and that would be problematic for many use cases including Podcast replacement. Sometimes us humans want AI to say "water is definitively wet", and that has been attempted by forcing LLM to accept that factoids are more syntactically correct, but that isn't a solution and it's still an architectural problem for these pseudo-AGI apps.
In the beginning people were skeptical but over time, as Wikipedia matured, the answer has become (I think): Don't blindly trust what you read on Wikipedia but in most cases it's sufficiently accurate as a starting point for further investigation. In fact, I would argue people do trust Wikipedia to a rather high degree these days, sometimes without questioning. Or at least I know I do, whether I want to or not, because I'm so used to Wikipedia being correct.
I'm wondering what this means for the future of LLMs: Will we also start trusting them more and more?
This argument is so dishonest its infuriating.
The main thing I’m getting from this discussion is that a lot more very smart people seem to have deluded themselves into thinking knowledge is objective or “locked in” than I had initially realized. The desire for certainty is an extremely human thing, but it’s a dead end, intellectually.
with chatgpt i don't know it's "experience" or "education" on a topic and it has no social accountability motivating it to make sure it gets things right so i can't estimate how much i should trust it in the same way.
The paranoia around hallucination is wildly overblown, especially given the low stakes of this context.
i think maybe your earlier comment was about the average trustworthiness of all podcasts vs the same for all gpt responses. i would probably side with gpt4 in that context.
however, there are plenty of situations where the comparison is between a podcast from the best human in the world at something and gpt which might have less training data or maybe the risks for the topic aren't eating an uncooked turkey but learning cpr wrong or having an airbag not deploy
No one person is particularly worth listening to individually, and as a podcast??? Good lord no.
LLMs beat podcasts when it comes to, “random exploration of an unfamiliar topic”, every single time.
The real issue here is that you trust podcasts so completely, by the way, not that ChatGPT is some oracle of knowledge. More generally, a skill all people need to develop is the ability to explore an idea without accepting whatever you first find. If you’re spending an afternoon talking with ChatGPT about a topic, you should be able to A) use your existing knowledge to give a rough first-pass validation of the information you’re getting, which will catch most hallucinations out of the gate, as they’re rarely subtle, and B) take what you learn with a hefty grain of salt, as if you’re hearing it from a stranger in a bar.
This is an important skill, and absolutely applies to both podcasts and LLMs. Honestly why such profound deference to podcasts in particular?
It’s very good at things like that. Go down a whole Roman Empire rabbit hole, have fun with it!
This is what an idle afternoon talking to ChatGPT is about, not trying to get it to do your job for you.
Nero was a Roman Emperor from 54 to 68 AD, known for his controversial and extravagant reign. He was the last emperor of the Julio-Claudian dynasty. Here are some key points about his life and rule:
1. *Early Life and Ascension*: Nero was born Lucius Domitius Ahenobarbus in 37 AD. He was adopted by his great-uncle, Emperor Claudius, becoming Nero Claudius Caesar Drusus Germanicus. He ascended to the throne at the age of 17, after Claudius' death, which many historians believe Nero's mother, Agrippina the Younger, may have orchestrated.
2. *Reign*: Nero's early reign was marked by influence from his mother, tutors, and advisors, notably the philosopher Seneca and the Praetorian Prefect Burrus. During this period, he was seen as a competent ruler, initiating public works and negotiating peace with Parthia.
3. *Infamous Acts*: As Nero's reign progressed, he became known for his self-indulgence, cruelty, and erratic behavior. He is infamously associated with the Great Fire of Rome in 64 AD. While it's a myth that he "fiddled while Rome burned" (the fiddle didn't exist then), he did use the disaster to rebuild parts of the city according to his own designs and erected the opulent Domus Aurea (Golden House).
4. *Persecution of Christians*: Nero is often noted for his brutal persecution of Christians, whom he blamed for the Great Fire. This marked one of the first major Roman persecutions of Christians.
5. *Downfall and Death*: Nero's reign faced several revolts and uprisings. In 68 AD, after losing the support of the Senate and the military, he was declared a public enemy. Facing execution, he committed suicide, reportedly uttering, "What an artist dies in me!"
6. *Legacy*: Nero's reign is often characterized by tyranny, extravagance, and debauchery in historical and cultural depictions. However, some historians suggest that his negative portrayal was partly due to political propaganda by his successors.
His death led to a brief period of civil war, known as the Year of the Four Emperors, before the establishment of the Flavian dynasty.
Does this make sense? Notice how little it matters if my understanding of Nero is complete or entirely accurate; I’m getting a general gist of the topic, and it seems like a good time.
Put more simply: I would rather have no information than incorrect information.
I work in a field of tech history that is under-represented on wikipedia, but represented well in other areas on the internet and the web. It is incredibly easy to get chatGPT to hallucinate information and give incorrect answers when asking very basic questions about this field, whereas this field is talked about and covered quite accurately from the early days of usenet all the way up to modern social media. Until the quality of training data can be improved, I can never use chatgpt for anything relating to this field, as I cannot trust its output.
In life, you exceptionally rarely have “enough” information to make a decision at the critical moment. You would rather know nothing than know some things? That’s not how the world works, not how discovery works, and not even how knowledge works. The things you think are certain are a lot less so than you apparently believe.
That said there are a few topics I’ve asked where I know enough to know (I had a grad degree in US history) and I’ve found it’s about on par with what you’d dig up through Google.
I’ve also found recently that it’ll state when it can’t equivocally say something. Or, in some cases, when I have misunderstood something and repeat it back incorrectly to try and clarify, it’ll correct me. 6 months ago, over text interface, it never would have corrected me: it almost always assumed the user was right and went along with it.
But again - it’s not like I’m fact checking Frank at the water cooler (or, often anyway). It’s shooting the shit and it’s great for that.
You take about low stakes but fiction is a massive industry. In making what is essentially Star Wars fanfic with ChatGPT, I’ve realized why AI was such a contentious point for the recent writers strikes.
2. No, ChatGPT is borderline schizophrenic. I trust a human and team of writers and producers to be more consistent in their bias and truthiness outright than a do a program that's trained on Reddit / Twitter and has almost no grounding.
If my mom asks it questions about my work industry, they are vague and high level enough that chatgpt is very accurate in it's answers.
If I ask it probing questions I care about my industry it hallucinates or fails to grasp key terminology.
So, if I was OP, likely asking it broad questions about things I know very little about is quite safe. Going too far down a rabbit hole to something super specific I would want to verify on Wikipedia.
In the future I see two big risks with the tech, approach and ownership:
1) short term, SEO is going to move into trying to influence the LLMs to push products unwittingly. Companies will be getting great offers from shady SEO outfits promising to make their products the default answers for broad questions?
2) mid term, now that all sense and pretence of safety and greater good are gone from openai, openai will be working out how to insert and see ads in the results? There will probably even be product placement in the results coming from the paid-for in-ms-office version.
Perhaps you're one of the people who think they're above it, and hence haven't developed that ability?
No actually, you can pay people to lie less and prove that they do so. We pay $20/month for ChatGPT to straight up lie to us and we know we need to sort it out.
ChatGPT lies to me but only in a specific way.
This is why I could never get “into podcasts”
Which apparently can mean 100 different things, but some popular radio talk show style ones I’ve been exposed to are just two guys rambling about a topic for two hours with lots of filler, where one person is essentially reading a Wikipedia article but getting interrupted every 5 words by the other, to go on - what are supposed to be - comical tangents.
Like, this could have been a 30 second sentence.
Far too frustrating for me.
I get exposed to them on long car rides with other people and that’s shaped my entire opinion.
I made a public GPT for a similar purpose. Some examples of my conversations with it:
That's my experience, too.
In addition to sticking to general topics, another fun way to talk with it is to discuss counterfactuals. "Hey, GPT! What do you think the world would be like now if Germany had won World War II [smartphones had not been invented, half of all people were born blind, etc.]." Those questions don't have right or wrong answers, and the answers it does come up with can be thought-provoking.
That sentence really resonated with me and sums up a bunch of ideas that have been swirling in my head recently. I've been "interviewing" ChatGPT about world history for a few months now and even turned it into an unsuccessful podcast! I really enjoy learning and creating much more than I would enjoy listening to my own podcast, though :)
Curious. Which podcast did you replace?
I've had some interesting discussions and it's helped me structure some thoughts by asking questions and follow-ups, then summarizing our conversation... all while I'm on a walk.
One fun thing I did with my family was have it do an interactive adventure story starring us. We had an adventure, and then used the built in DALL-E to generate images of scenes from our adventure.
I hate this about all voice assistants. They force you to speak in an unnatural way.
It seems like there’s nothing intelligent about when they decide to respond, it’s like they just wait for X milliseconds of silence instead of using the context of what’s being said like a human would.
Sometimes it’s the opposite problem. You finish asking your question but you gotta wait for the assistant to pick it up whereas a human would understand that you’re finished talking based on what you said.
It might seem small and unimportant but I really think it’s one of the main reasons why voice assistants feel so… artificial.
That and long ass responses.
I think the same thing can happen when speaking to anyone with different societal/cultural factors than yours. We just become more accustomed to it over time and it's less noticed. I think if this GPT had a big green alien face then we would find speaking to it less strange, somehow.
Maybe the same will make it into ChatGPT at some point.
It would be great to have an option to have ChatGPT stop when saying "over" - it does not feel natural to have to blurt the questions all at once.
you can ask it to be terse and shorten answers in the custom instructions section.
I've also used it several times on ~15-20min drives to memorize something I wanted to have available for immediate recall. I had it chunk & quiz me, and by the end of the drive I had it down pat. Fun use of drive time.
Can you go a bit into how this works? How do you prompt it? This is the first use of ChatGPT I've heard of that would directly benefit me.
I'd then switch to the phone and retrieve the chat from History.
Here's an example prompt I just used to help my son prepare for a DMV written test:
``` I'm going to paste a large list of questions and answers and then switch to voice mode. Once I indicate that I'm ready, begin quizzing me on these questions. Feel free to rephrase slightly. My goal is to achieve complete retention of all of these questions through quizzing and spaced repetition. The questions are California DMV questions. I am preparing to take the written test. 1. *Q:* You may drive off of the paved roadway to pass another vehicle. *A:* Under no circumstances.
2. *Q:* You are approaching a railroad crossing with no warning devices and are unable to see 400 feet down the tracks in one direction. The speed limit is... *A:* 15 mph. ... ```
I just wish they'd offer more languages as it currently is English only. It can speak German and Spanish, but with a very clear English accent, which makes it kind of funny and makes one think about why the accent sounds so real.
It seems similar to what other commenters are saying about it taking in German with an American accent.
It's no big deal but Siri/Google/Alexa voices are better normalized, IMHO. I don't mind heavy accents at all in humans but they get annoying after a while for a voice assistants, especially if it's not my accent :)
I tried it now but it doesn't seem to be working for me. Maybe related to the outage yesterday (is it still going on?). It just listens but never replied.
Among others...
Had to check myself.
> Afrikaans, Arabic, Armenian, Azerbaijani, Belarusian, Bosnian, Bulgarian, Catalan, Chinese, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, Galician, German, Greek, Hebrew, Hindi, Hungarian, Icelandic, Indonesian, Italian, Japanese, Kannada, Kazakh, Korean, Latvian, Lithuanian, Macedonian, Malay, Marathi, Maori, Nepali, Norwegian, Persian, Polish, Portuguese, Romanian, Russian, Serbian, Slovak, Slovenian, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Ukrainian, Urdu, Vietnamese, and Welsh.
The English pronunciation is perfect, but some Russian words do not have the correct pronounciation or emphasis.
I'm wondering if I can train the model to pronounce these words more accurately.
I really feel like I live in the future.
My theory is that voice assistants are made to be cheap entry points to those ecosystems. LLMs, on the other hand, are incredibly expensive to run. No GAFAM wants to pay dozens of cents when you ask their assistant to switch on the lights.
If anyone still remembers what we had before Siri, you’ll know how well it worked.
It's still very much a demo and not as good as GPT-4, but it responds much faster. It's fun to play with and it shows the promise. Open models have been improving very quickly. It remains to be seen just how good they can get but personally I believe that better-than-GPT-4 models are going to run on a $1k gaming PC in just a few years. You will be able to have a coherent spoken conversation with your GPU. It's a new type of computing experience.
I will probably update to the OpenHermes vision model when Nous Research releases it, so it'll be able to see with the webcam or even read your screen and chat about what you're working on! I also need to update to Whisper-v3 or Distil-Whisper, and I need to update to a newer StyleTTS2. I also plan to add a Mandarin TTS and Qwen-7B bilingual LLM for a bilingual chatbot. The amount of movement in open AI (not to be confused with OpenAI) is difficult to keep up with.
Of course I need to add better attribution for all this stuff, and a whole lot of other things, like a basic UI. Very much an MVP at the moment.
Like you said it’s difficult to keep up with and to me it feels very much like open source stuff might win for inference.
Not sure big players won't be pushing heavily (as in, not releasing their best models) for the fat subscriptions/data gathering in the cloud, even if I'd much rather see local (as in cloud at home) computing
ChatGPT is essentially push-to-talk with a little bit of automation to attempt to press the button automatically at certain times. Mine is continuously listening and can be interrupted while speaking, but isn't yet smart enough to delay responding if you pause in the middle of a sentence, or stop responding at the natural end of a conversation.
I wrote up my detailed thoughts about it here: https://news.ycombinator.com/item?id=38339222
https://github.com/modal-labs/quillman
I also built something similar using WebKit speech recognition (limited to Chromium) a year back for my own use but it was hooked to davinci-003.
Doesn't that seem a bit excessive for Whisper and Coqui? Or does it also run an LLM for a full local stack?
Haven't heard of that one before, I'll have to check it out.
One possibility is to capture Thanksgiving gatherings as a growth hack where people demo this to their families/friends and increase app downloads for OAI.
[1] https://www.teamblind.com/post/OpenAI-employees-did-you-sign...
I made an iOS shortcut a while ago that uses Siri with the ChatGPT app (it has iOS shortcut bindings) and despite Siri being a useless pile of junk compared to this, I actually prefer Siri's voice to this in some ways, because it doesn't feel so over the top.
Maybe this is partly because of different cultural expectations between the USA and Europe? Or maybe I'm just being too cynical and ChatGPT really is that happy talking with me!...
At one point it misinterpreted me mentioning “tai chi” as “I can’t breathe” and responded with advice about medical emergencies.
I really like dictating things sometimes, and Whisper is perfect for that (automatic paragraphs inside the model itself would be nice but not a big deal).
If anyone is interested - the "Whisper speech recognition in iOS" part is based on this shortcut I found that you can easily use yourself on both iOS and MacOS (free except for the OpenAI API usage fees obviously): https://giacomomelzi.com/transcribe-audio-messages-iphone-ai...
> free except for the OpenAI API usage fees
There are several versions of Whisper which have been distilled and can run locally, so I don’t see what advantage making API calls would be other than increased latency and decreased reliability and data security.
First question, is there another STT you have used which works better for you?
Second question, is there any reason your voice might be considered unusual, like having a strong Welsh, Irish, or Indian accent, or being Deaf or Hard of Hearing?
Second, even if the transcription is correct, it cuts me off at inappropriate times. It’s hard to talk naturally without pauses.
I haven’t used a better transcription model, no.
Yes, sometimes it thinks I'm done speaking when I'm not, but on the whole it's very good. Siri/Alexa, et al are not only unusable but are now supremely frustrating.
Reminds me way too much of some of the people I had to talk to, when cleaning up my mother's affairs. Places trying to get me to pay bills I did not owe, call center agents "cheerfully" following scripts that they themselves hated. The voices sound exactly like that.
Give me a neutral voice. This is a computer I'm talking to, not a fake friend.
I also though don't like most the chatGPT voice models besides for Sky. Sky to me is really good. Robertson Dean reading an audio book is perfection but Sky is pretty awesome.
I should add that as an American there are a ton of American voice actors that ruin books for me too. Sometimes this can be fixed if played at 1.2X speed.
I am here to help you. Say whatever is in your mind freely, our conversation will be kept in strict confidence. Memory contents will be wiped off after you leave,
So, tell me about your problems.
Changed from https://twitter.com/gdb/status/1727067288740970877 above, plus I degregged the title. I'm sure he won't mind. (submitted title was "ChatGPT Voice for All Free Users Announced (By Greg Brockman)")
Thanks!
Hello. USian here, talking about things from a USian perspective! YMMV in other countries, visit your local dealership for more information, & etc, etc, etc.
A multi-million (or multi-billion) dollar company who can
* Store recordings all of my utterances forever
* Use my recordings to make tons of money without ever sending me a cent
* Share my recordings with "trusted business partners" for "contractually-agreed-upon [with their partners, not with me] purposes" whenever they feel like
* In short, use those recordings that they've stored forever for any nearly any purpose they see fit to
is not "someone" that I'm having a "conversation" with.
When I have a conversation with someone, they have a human's capacity to remember and disseminate that information. People get worse and worse at remembering things as those things recede into the past. Even if someone remembers an event very clearly, it's also usually impossible for someone to perfectly relate that event to another.
When a company records what I'm saying and feeds that recording into their perfect-copy-machines and the perfect-copy-machines of their "trusted business partners", that's an entirely different thing. Surely you see that?
Do you ask people for money when they talk to you in general? Saying "it's not a someone" is just arbitrary categorization. If it was like talking to a wall, you probably wouldn't talk to it, or would you? Is it possible not everything has a monetary aspect on personal level, but on a corporate level that's the only way you can capture value which is then also expressed in non-monetary ways such as people providing work or advice for the furtherment of a goal or project?
I don't know. Let's maybe think about the word we live in. Some might say the most worthless thing in the world is money. What is it good for by itself? It's not only useless paper you can't even wipe your behind with, but these days it's just digits on a computer. But the value you can get out of it is provided by people and systems that do something useful for you in exchange for it. And if something provides that value directly in communication with you, that's much the same as being paid digits on a computer.
In other words, are we so shallow we need to first pay for ChatGPT's output in money, and then it should pay us back the same money (or less) for our voices and data, in order to comprehend what a barter exchange of value is?
1. We're anonymous here (by default)
2. You can't be fingerprinted from your HN posts the way you can be with your unique voice
That doesn't seem related to the point?
Do you disagree that "Posting on a public Internet forum (such as HN)" is (other than the "communicating information to other people" aspect) _radically_ different than "Talking one-on-one to someone in person without any recording devices present"?
If you do disagree, then there's no way we can have a even vaguely productive conversation on the topic.
I would imagine that voice synthesis models would somehow be trained on data from native speakers, so why the accent?
edit: It would appear so:
> Chicago is a city sprawling with Polish culture, billing itself as the largest Polish city outside of Poland, with approximately 185,000 Polish speakers, making Polish the third most spoken language in Chicago.
Translating a language is difficult on top of regional dialects creates additional complications.
If someone is talking long enough to me I can identify their birth State based upon their American accent. Every State has a different pronunciation of certain words which can leak location data.
Guessing you're Californian, educated parents, about 34 years old. No siblings.
Without AI but based on some freely available statistics combined with post history, we can say definitively (assuming honesty on OP’s part) that OP is 35 or 36, and was either born in Germany or moved there when they were young, and they likely still live there or at least have a strong affinity to Germany.
We can reasonably speculate that the fact that they speak German and likely live there, as well as the standard of their English drastically increases the likelihood their parents were educated. Germany’s current total fertility rate is 1.5. It was likely higher when OP was born back in 1986/7 but I couldn’t immediately find any data on that and didn’t look too hard. Given the fact that TFR is a population mean, this suggests a high proportion of single-child families, and generally the more educated one is, the fewer children, so it’s a decent guess that OP is an only child (apparently more than 50% of German families have only one child), but I don’t see anything that would make this a sure fire bet.
I was able to complete some further analysis which could reveal more likely truths about OP such as gender, political orientation, sexuality, etc., but I think this goes far enough without starting to doxx them.
you wouldn't believe it, but models haven't been trained yet. as usual.
If you learn how to pronounce specific vowels, consonants, etc. in a particular way, it takes a LOT of effort to learn how to pronounce these in a different way. You can approximate to a good extent, but most researchers say that if you don’t develop this skill as a child, you won’t ever be able to pronounce things in a way that sounds like a native speaker.
Presumably, the models have been exposed to significantly more American accents than other accents, and learning how to pronounce phonemes with subtle differences without accepting a “close enough” approximation is a big challenge, especially given that there is already a threshold for acceptability at which level you can still sound like a native English speaker.
Hits differently than it would have last week. :-/
It’s wild. Very close to Star Trek land. I never thought I’d see this.
We went from debating a topic to Googling the answer and resolving a whole night's debate in seconds.
And I will tell you that Google didn't always get it right. Often a tweak to the search term could get the wrong answer, but win the argument.
There are three topics banned at my table. Religion, politics, AI, Sam Altman, punctuation and maths.
Edit: Based on the response time, I don't understand where the audio is being generated, is it local?
I understand the downvotes for showing disapproval at OpenAI's policies (or big tech's), but I don't know what else could I say. Should I have just replied with a snarky comment saying "a-ha! you expect them to respect your privacy? pffff"
Has everyone forgotten about Google? Facebook? Every other SV Tech-Startup? Why would OpenAI be different?
"They're a non-profit". Yeah, sure. As if Wall Street, M$ et al would just not try to monopolize and commercialize AI as soon as possible.
Usually this truism is invoked in reference to advertising business models where the actual customer is advertisers, and you're just inventory so from cable TV to web search the business has usually been happy to take a shit in your mouth whenever it serves their advertisers.
We will see how things pan out with ChatGPT as the interests of a ChatGPT Plus user likely align closely with the interests of someone on the free tier. Historically I think freemium business models tend to be kinder to their customers than advertising-based ones, or at the least, they simply lean on you until you upgrade to paying customer (if they do enshittify the free tier I think you would see an amazing conversion rate to ChatGPT Plus because the product is so good and unique).
Actually, it’s the mobile application. They should have opted for a boring PWA but instead they developed a nice native application with good UI/UX. I also appreciate that they created integrations with the Shortcuts app which allows you to do make very powerful things. I really hope they extends this because it’s text only for now but now with GPT4V and the new ability to generate files, you could do impressive things with shortcuts.
Edit: The headphone icon appeared on Nov 22, 2023 around 6am CET without an update to the app. There is no «New Features» entry in the Settings.
I think the dialect is part of it as well: identifiably US West coast, presumably SF-Bay. Something more neutral, say, mid-Atlantic would be preferable for non-USaians. I asked ChatGPT if it could change its dialect - it said yes, but sounded exactly the same anyway.
Maybe you get used to it? Dunno, I'll try again, if I have a use case...
Dialect, or accent?
tl;dr: I tried both...
No, I'm a fighter pilot in EVE Online.
"Explain this concept to me."
"Give examples from the physical world and every day life."
"Elaborate on that last example."
"Give some examples from the domain of ___."
"Tell me about the history of the concept and why it was necessary to invent it."
"How does this concept relate to the concept of ___."
So I'm trying to approach what I want to learn from many different angles and build context and intuition. Sort of like talking to a tireless grad student tutor. Errors would become more obvious due to contradiction. I've caught it making some mistakes but they were sort of trivial and I could see the spirit of what it was getting at. Plus, "truth" tends to "click."
Oops! Our systems are a bit busy at the moment, please take a break and try again soon.
I could see it being a bit inaccurate if the transcription silently corrects my error.
An unfortunately absent bit of accessibility that would really help second language learners
I assume there will be a MS ChatGPT if OpenAI folds though. Or just Bing.
Ideally I would:
* query by voice
* get the response in text as soon as possible, maybe even as I'm still talking, with the LLMs correcting the response as I give more information, and without needing me to touch anything
* follow up by voice, even before the response was finished, and without needing to press anything
I hope we'll get there eventually!
I can’t. At least not with reading every word. I can finish an audiobook much faster than reading the book on my own.
And I also can’t pick up as much reading by myself. My preferred way of consuming is now reading while a screen reader or audiobook is played at 1.5x or faster speed. Much better retention because
Obviously, brains work very differently, so there really is no one-size-fits-all solution; we need the varied methods.
That's the crux of the issue though, isn't it? I don't read words by reading single letters, and I don't think I read sentences by reading every single word either. That'd be the main reason why listening is slower than reading: the forced linearity.
To be fair, my "all" was clearly an exaggeration, I'm sure there are people for whom this doesn't apply, but I'd still expect it to apply to 50%+ of the populations of countries where literacy is not a significant issue anymore.
Like an Android device with a smart speaker? I think there is an OpenAI app.
Serious question.
It's suggesting to ask it where to order 195 pizzas from.
Does it have access to active business records like that?
You can follow my progress on twitter if you're interested (I tweet about other things, don't expect LLM/Data stuff only :)
Online communities are echo chambers, imagine now talking to a bot all the time.
Also this Call Annie requires Google sign in.