"Play my favorite" is just a knowledge problem. If GPT fails there, it's because it doesn't know your favorite, not because it can't parse the request or understand what you need it to do.
You have to speak certain ways to Siri to get it to do things.
Unless specifically hard-coded, Siri will never receive "damn I'm finding it hard to read" as input as decide to turn on the lights. GPT will.
"Knowing your favorite" is the context.
> Unless specifically hard-coded, Siri will never receive "damn I'm finding it hard to read" as input as decide to turn on the lights. GPT will.
Of course it won't. You have to very specifically fine tune it to understand what light conditions are, where you are in the house, and what it is you need to turn on.
The sibling comment literally says "I had to provide a long-ish sentence as a context/programming instructions before it could do anything". https://news.ycombinator.com/item?id=37464563
There's considerably more in that prompt besides just "you need to act like a home assistant"
Where you are in the house and what needs to turn on, at least, is an API query job, not a fine-tuning job.
As far as whether it can understand the relevance of lighting to the situation, I just asked ChatGPT 3.5 the question 'Acting as an AI home assistant, if you hear me say "I'm finding it hard to read", what actions would you take?' and 'Adjust the lighting' was the second option it gave back (after 'ask for clarification'). I think we're there, honestly, we just don't have the different parts connected yet.
And that API magically comes form where?
> I just asked ChatGPT 3.5 the question 'Acting as an AI home assistant, if you hear me say "I'm finding it hard to read", what actions would you take?'
So, basically:
- you had to pre-program Chat GPT to act as a home assistant
- you had to provide it with specific context and specific phrasing for it
- it still failed, asked for clarification, and only then responded
And now you have to this song and dance every time you want to coax GPT into doing what you need (and that's what RestGPT does).
HomeAssistant, or any number of other providers. Do you think this part is somehow difficult?
> you had to pre-program Chat GPT to act as a home assistant
That is what we call "a prompt". It is a well-known technique. I am surprised that this should look strange to you.
> you had to provide it with specific context and specific phrasing for it
That is what we call "a prompt". It is a well-known technique. I am surprised that this should look strange to you.
> it still failed, asked for clarification, and only then responded
You have misunderstood. In its list of actions to take, the first and only response it gave, the first thing it said it would do in context is ask for clarification as to why I was finding it hard to read. That seems entirely reasonable to me. Does it not to you?
> And now you have to this song and dance every time you want to coax GPT into doing what you need (and that's what RestGPT does).
So what? It's not something the person sat in the dark ever has to care about.
> That is what we call "a prompt". It is a well-known technique. I am surprised that this should look strange to you.
Funny how we're in the discussion about context, and you decided to ignore and discard the entire context of the discussion :)
Prompting for task performance is fine as long as you're not expecting the end user to have to replicate your prompting. Your goal is to change model activations for a given input, the end user is similarly affected regardless of if you used a prompt or fine-tuned.
-
This task doesn't require fine-tuning though, zero-shot performance is enough:
I generated a mock schema from Home Assistant's API (https://data.home-assistant.io/docs/states/) and explicitly gave the model the option to ask for clarification, but it has no problem translating non-obvious commands into actions without asking for details:
https://chat.openai.com/share/fc5b972f-4641-47a1-9842-2e0d69...
Note those objects mirror Home Automation, you could hook that up today without any song and dance. Combine that with RAG and you'd have something that's a lot more useful than Siri and capable of improving performance over time.
You had to provide two pages of text and do manual mapping between human-readable names and some weird identifiers to provide the simplest functionality.
Funnily, this functionality is also completely unpredictable.
I ran your prompt and first request, and got "Identify the area with the lowest observed request volume and increase the brightness of the light in that area to improve the lighting." ChatGPT then proceeded to increase brightness in the garage.
---
It's also funny how in the discussion about context the context of the app is forgotten.
With super primitive wake word detection and transcription, the most you get is:
- What the user said
- How loudly each microphone in the house heard it.
If you take a look at the mock object in that transcript, that's what it maps to...
```json { "request": "I'm finding it hard to read" "observedRequestVolume": [ 3eQEg: 30, iA0TN: 60, h1T3y: 59, 5Qg1M: 10 ] } ```
The only part that would be human provided is: "I'm finding it hard to read"
The invented challenge was to see if using a suboptimal set of inputs (we didn't tell it where we are) it can figure out how to action.
It's zero-shot capability that makes LLMs suitable for assistants: traditional assistants can barely handle being told to do something they're capable of in the wrong word order, while this can go from hastily invented representation of a house and ambiguous commands to rational actions with no prior training on that specific task
The end user would never type in a word of that: they'd say "[Wake word] play me some music"
A piece of software running on a device would transcribe what it heard, and fire off a request to the LLM with all of that text wrapped around their statement.
For ease of sharing I used the web interface to provide the instruction, but you'd use the API with a prompt which also dramatically increases determinism.
No one is writing out the state of each light bulb: you trivially query that information programmatically and bundle it with the request.
—
In a real product there'd be explicit handling of detecting where the request came from, that's already a problem that's been worked on, but I wanted to demonstrate the main difference vs Siri: zero-shot learning
The LLM wasn't told what those volumes mean, but it was flexible enough to infer the intent was to provide a form of location, rather than ask.
It's a forced example so if you want to get caught up on the practicality of audio for locating people be my guest, but it's to show LLMs are great at "lateral applications" of capability:
You give them a few discrete blocks of functionality and limited information, and unlike Siri they can come up with novel arrangements of those blocks to complete a task they haven't yet seen.
—
Honestly the fact you keep going back to "look at all the text" feels a bit like if I showed you the source code for an email messaging app, and you told me: "No one will ever use email! Who would write all that instead of just writing a letter and mailing it?!"
No, you don't.
GPT w/ memory: "Because you still have dyslexia."