The "chance of hallucinations" is the tricky bit - if I have to manually check everything it does in case it's hallucinating, then it's not actually a solution. It's not saving me time (as TFA says).
No idea if it feeds all of the code to the LLM or just parts, but it's pretty good at interpreting what the code does
Through it could do stuff like go through your repository(ies) and generate embedding for sections of it and then have a vector database + retrieval argumented generation (RAG) system.
Basically all the LLM needs to do is translate human writing to some format that a "normal" service can use. That can then leverage the existing spotlight system that's pretty decent at searching stuff on the phone anyway.
Then it'll report it back to the LLM which translates whatever format back to something humans can process.
Basically "less shit Siri"
Although I could be hallucinating the whole thing.