ChatGPT vs. open source on harder tasks
github.com
github.com
I've been waiting for a tool exactly like guidance, something that lets me compose prompts the way I think about prompts. Langchain was too much of a headache for me. Playing with guidance, it feels like this is how programmatically interacting with a llm should be. Also, really happy to see the MIT license.
I'd love to see ports to other languages.
https://itsbehnam.com/DetachGPT-Switching-to-Open-Source-Unc...
People already adopt context-specific specialized languages in contexts like that for clarity, what will probably happen is that LLMs will get better at adapting or being adapted to specific contexts.
When asking for creative but well-structured output that requires specific definitions of certain words to be useful, I've found it works best to use words for your terms that don't actually exist. Otherwise it seems to focus a bit more on the broader word more than it should rather than your definition, especially if you use it a lot.
I totally gave that up with chatgpt and the magic of it is that it understands it so so we'll.
Sometimes I also switch to German in the middle of a sentence or use very vague words to describe something if the word is missing in my head.
When comparing model intelligence, I wish that only the most capable models were compared.
But 3.5 is generally available, fast, and costs a WHOLE LOT less for API users; those are real capabilities that matter though by "capable" you only meant the quality of the output. It's worth comparing to 3.5 also.
Is it more expensive this way? At the rate I'm using it, probably 2-3x. Is it worth it? Absolutely. I've never in my life felt so happy to pay for something per request.
I use the plugins and new features on ChatGPT though, and it's nice to have a single interface to use 3.5-turbo for comparison, Code Interpreter once in awhile, etc. So I'll probably just keep it. It's $20/month, after all.
Edit: out of curiosity I posted the meeting summarization example to GPT-4 and indeed it produced both the relevant snippets and the correct answer to the stated question.
> While Vicuna is somewhat comparable to ChatGPT (3.5), we believe GPT-4 is a _much_ stronger model, and are excited to see if open source models can approach _that_.
also...porn will make people achieve anything they set their mind to
Stable link: https://github.com/microsoft/guidance/blob/8677f3aa269e05ecb...
Commit that removed the notebook: https://github.com/microsoft/guidance/commit/b91d332f49a55ed...
I deleted the notebook yesterday after realizing it was weird to have a disclaimer saying 'this is our personal opinion' on a github repo under 'Microsoft'. I moved the notebook to a gist here: - https://gist.github.com/marcotcr/64ca85bd0be724f6d8fb8f1b3d2...
I didn't know we were on hackernews, otherwise I would of course not have made it a broken link :)
This morning, we put the link back, without the takeaways with our personal opinion at the end (it's still on the gist and on the blog post https://medium.com/@marcotcr/exploring-chatgpt-vs-open-sourc... )
I thought the point of Microsoft Guidance was that we could specify a grammar or regular expression that the output must match?
I think regarding: "Make sure to use real numbers, not fractions."
It would be simple using Guidance to create some kind of validation function, which guidance would invoke after it gets the response from the model. In this case you would check if the number is a real number or not. And if not, the model would invoke again with a different seed.
Assuming the model gives a real number, as expected, ~90% of the time, you would only see retries very infrequently, and you could be pretty confident that your output will match what you expect.
(again, I'm not associated with this project, I've only scratched the surface playing with it locally, so this is just speculation)
I have a pretty decent understanding of how the current generation of LLMs work, and I too don't think we've achieved consciousness. But the prompts just sound so much like something you'd tell a human, or something you'd tell a host in Westworld.
- Quality on task: For every task we tried, ChatGPT is still stronger than Vicuna on the task itself. MPT performed poorly on almost all tasks (perhaps we are using it wrong?), while Vicuna was often close to ChatGPT (sometimes very close, sometimes much worse as in the last example task above).
- Ease of use: It is much more painful to get ChatGPT to follow a specified output format, and thus it is harder to use it inside a program (without a human in the loop). Further, we always have to write regex parsers for the output (as opposed to Vicuna, where parsing a prompt with clear syntax is trivial).
- Efficiency: having the model locally means we can solve tasks in a single LLM run (guidance keeps the LLM state while the program is executing), which is faster and cheaper. This is particularly true when any substeps involve calling other APIs or functions (like search, terminal, etc), which always requires a new call to the OpenAI API. guidance also accelerates generation by not having the model generate the output structure tokens, which sometimes makes a big difference.
As long as the template system of choice supports custom functions, I don't see the need for a custom language for this. Please correct me if there is something deeper that I am missing? (other than the specifics to messaging and roles, where I for one prefer to put these in separate files and then pass them as arguments, easier to manage, recombine, and reuse this way)
https://gist.github.com/verdverm/747b0b810dcc5518d699b14f0d0...
When I use chat prompting…
- system: sets the context for the bot
- “You are…”
- “Act like…”
- “Pretend…”
- assistant: specifies the task to be performed - “Classify…”
- “Extract entities…”
- “Translate…”
- user: asks the question, adds examplesThese guys were asking questions and giving examples in the assistant prompt. I feel like that messed up the LLMs responses.
I can’t wait to try Vicuña and MPT. I can’t quite figure out how to host them in Azure without an expensive VM. Maybe Azure DataBricks…?