Siri isn't a language model, AFAIK. I'm sure it's using machine learning on the back end, but it's doing something very different.
The thing is, a voice assistant needs to actually perform Natural Language Understanding -- it has to understand the intent of what's being said to it. That's actually a very different AI problem than what LLMs are doing at this point. Samsung's Bixby -- yes, the one everyone turns off -- can handle a "conversation" like this: (This is run using the "Space Resorts" sample capsule, a fictitious hotel booking system)
---
User: Show me hotels near Mars
Bixby: I found these space resorts on Mars. (Shows two, “Martian Oasis” and “Space Y”)
User: [taps Martian Oasis] make a reservation for Christmas week
Bixby: Which habitat pod would you like? (Shows two)
User: Honeymoon
Bixby: Are you sure you want to book this trip? (Shows the hotel, “Number of astronauts: 2”, and the dates for Christmas week)
User: change the astronauts to 4
Bixby: (Shows the sam confirmation screen, with the number changed)
User: Yes
Bixby: I hope you have a great time at Martian Oasis. (Shows booking info)
---
Now, in some ways this is no great shakes, but I don't think "all [Apple] needs to do is integrate some OpenAI APIs," because they don't have an API for NLU and deriving intent. GPT4 et. al. aren't currently designed for this kind of problem space. You could use them to make Siri (and Bixby and Alexa) better at answering general questions, but you can't use them to make the voice assistant do something, just respond to something.
Bixby 2 (the current version) is, incidentally, what came out of Viv, the voice AI startup that the Siri folks made after they left Apple.