HNHacker News
TopNewBestAskShowJobs

jeffharris

99 karma · joined July 7, 2015

submissionscomments
jeffharris··on OpenAI Audio Models
oh doh. thanks ... we just pushed a fix for the crash. Unfortunately our currently implementation needs service works for streaming audio, so the "fix" was to disable the feature if the worker isn't available
jeffharris··on OpenAI Audio Models
really depends on the voice. but in general, no we want to sound as realistic as possible and I expect future voices to keep improving on this front
jeffharris··on OpenAI Audio Models
this is a good solve. we don't support word time stamps natively yet, but are working on teaching GPT-4o that skill
jeffharris··on OpenAI Audio Models
S2S is where we're investing the most effort on audio ... sorry it's been slow but we are working hard on it

Top priorities at the moment 1) Better function calling performance 2) Improved perception accuracy (not mishearing) 3) More reliable instruction following 4) Bug fixes (cutoffs, run ons, modality steering)

jeffharris··on OpenAI Audio Models
we're working hard on it at the moment and hope we'll have a snapshot ready in the next month or so

we've debugged the cutoff issues and have fixes for them internally but we need a snapshot that's better across the board, not just cutoffs (working on it!)

we're all in on S2S models both for API and ChatGPT, so there will be lots more coming to Realtime this year

For today: the new noise cancellation and semantic voice activity detector are available in Realtime. And ofc you can use gpt-4o-transribe for user transcripts there

jeffharris··on OpenAI Audio Models
this has been coming up often recently. nothing to announce yet, but when enough developers ask for it, we'll build it into the model's training

diarization is also a feature we plan to add

jeffharris··on OpenAI Audio Models
we'll keep expanding these GPT-4o based models with more controls. Is the main feature missing we're missing custom voices?
jeffharris··on OpenAI Audio Models
not open source at this time. unfortunately they're much to large to run on normal consumer hardware
jeffharris··on OpenAI Audio Models
Should do! here's an example https://www.openai.fm/#4a5a82db-faea-4f80-813c-3131902c2458
jeffharris··on OpenAI Audio Models
thanks for flagging ... number fidelity (especially on languages that are unfortunately less represented in training data) is still something we're working to improve
jeffharris··on OpenAI Audio Models
nothing to share on open source yet, it's something we'll keep exploring. Especially as the models get smaller so more able to run on regular devices
jeffharris··on OpenAI Audio Models
try the ballad or fable voices
jeffharris··on OpenAI Audio Models
We've been using the FLUERS eval and you can see comparisons to other models on the market in the post https://openai.com/index/introducing-our-next-generation-aud...

Curious if there's a benchmark you trust most?

jeffharris··on OpenAI Audio Models
It's a slightly better model for TTS. With extra training focusing on reading the script exactly as written.

e.g. the audio-preview model when given instruction to speak "What is the capital of Italy" would often speak "Rome". This model should be much better in that regard

= No plans to have localized voice models, but we do want to bring expand the menu of voices with voices that are best at different accents

jeffharris··on OpenAI Audio Models
1/ we've been working a lot on accents, so expect improvements with these models... though we're not done. Would be curious how you find them. And try giving specific detailed instructions + examples for the accents you want

2/ We're doing everything we can to make it fast. Very critical that it can stream audio meaningfully faster than realtime

3+4/ I wouldn't call hallucinations "solved", but it's been the central focus for these models. So I hope you find it much improved

jeffharris··on OpenAI Audio Models
Yes, from our terms: "Don’t build tools that may be inappropriate for minors, including: Sexually explicit or suggestive content. This does not include content created for scientific or educational purposes." https://openai.com/policies/usage-policies/
jeffharris··on OpenAI Audio Models
We're thinking about diarization (adding time awareness to GPT models) but no firm plans to share just yet
jeffharris··on OpenAI Audio Models
so good! https://www.openai.fm/#28540f27-5b51-445a-b1d6-1c89711a2c4f
jeffharris··on OpenAI Audio Models
some of the older voices are definitely less steerable, more robotic

we put little stars in the bottom right corner for the newer voices, which should sound better

jeffharris··on OpenAI Audio Models
Hey, I'm Jeff and I was PM for these models at OpenAI. Today we launched three new state-of-the-art audio models. Two speech-to-text models—outperforming Whisper. A new TTS model—you can instruct it how to speak (try it on openai.fm!). And our Agents SDK now supports audio, making it easy to turn text agents into voice agents. We think you'll really like these models. Let me know if you have any questions here!
jeffharris··on Two new Gemini models, reduced 1.5 Pro pricing, increased rate limits, and more
apologies: it's taken us a minute to switch the default `gpt-4o` pointer to the newest snapshot

we're planning on doing that default change next week (October 2nd). And you can get the lower prices now (and the structured outputs feature) by manually specify `gpt-4o-2024-08-06`

jeffharris··on Structured Outputs in the API
yep image input on the new model is also 50% cheaper

and apologies for the outdated pricing calculator ... we'll be updating it later today