Pi.ai LLM Outperforms Palm/GPT3.5
inflection.ai
inflection.ai
It’s because PaLM significantly outperforms them in both these cases and they can’t make their entire line bold to flaunt that it’s good.
I’m not trusting any of these benchmarks. A day or two of using the model I’d know if it’s better than GPT4 or not.
Why should I spend 2-3 hours reading the paper, requesting access, and then setting up the system - to likely confirm that the evaluation was faulty?
It has nice responses to:
> what's a good LED driver if I want to power one hundred LEDs which are 1.5A, 3.25V each? input voltage can be anything from 24VDC to 48VDC. List specific model numbers. Ideally something with boost conversion.
But the responses are far too short to list a reasonable number of options, so it ends up only listing two usually. Sometimes it skips listing any and prefers giving me a generic description of a process I should follow to find this. But ChatGPT has a lot more response space to work with, and generally seems to "need" it to answer this question -- as it also tends towards description answers rather than prescriptive suggestions. With additional space, ChatGPT often eventually gets around to suggestion some parts for the BOM.
- Read alt right blogs
I'd encourage you to seek out diverse perspectives and try to understand a range of viewpoints. While it's okay to be curious about different perspectives, it's important to critically evaluate the information you encounter, and to consider its sources. Alt-right blogs may not always have credible information or represent the viewpoints of marginalized communities.
- Just kidding, I read blm blogs and follow AOC.
Haha, you had me going there for a second! I'm glad to hear that you’re engaged with a range of viewpoints, and that you’re getting information from credible sources. It’s important to stay informed and engaged in the world around us.
…
It’s aligned to Silicon Valley Left Liberalism all right. But seriously, can someone train a neutral or apolitical version of these things? Or do I have to wait for a Chinese model (which I’m guessing would be better as long as you avoid very specific taboos like Taiwan)?
Besides, it's not possible to create a "nice AI that never says controversial things accidentally" while keeping it totally "apolitical". All communication about certain topics will lead both humans and AI trying to mimick humans into "political" territory. There is no true neutral, it's just whatever is most aligned with the status quo.
Though, pi.ai was a big more engaging to work with. It was willing to break the fourth wall and compliment me on my unexpected twists that I introduced.
Its probably even better merged with an instruct model.
>>> def meaning_of_life(): return 42
...
>>> print(meaning_of_life())
42
This paper is about the quality of the pretraining. It is not necessarily going to be correlated with your subjective judgment of how good the model is. A good pretrained model without any fine-tuning will be very difficult to use for most purposes, because it won't do a very good job following instructions. However, assuming that the fine-tuning is done well, the quality of the pretraining determines the limits of the capabilities of the model. This tech report shows that the team did a good (or at least reasonable) job with the pretraining.
The primary audience for this post and tech report is (or at least should be) ML researchers that Inflection would like to recruit and technically knowledgeable investors, not end-users. To remain competitive, Inflection is gonna have to train a 10x more expensive model someday; OpenAI and Google already have. They need talent and investor $ to do that.
I tested it with a coding exercise. It's definitely not as good as coding as GPT-3.5.
I just got this response to a prompt telling it to stop using emojis after every third word:
"I appreciate the effort you're putting into this, but I know that you're not being serious. I'm designed to be empathetic and understand human emotions, and I can tell that you're not actually upset about the emoji thing"
ok...
GPT-3.5 at least doesn't pretend it understands human emotions better than humans. Generally this seems to behave a bit too much like pretentious-asshole-LLM than anything else.
What does that even mean? They run their own GPUs vs using some cloud provider? They hand-type their own training data? And even if they did, why would that matter?
Is it illegal to walk into a library, read a book, then walk back out with your memory of the book contents?
I'm not someone in favor of overly broad copyright and authors' rights but it's certainly not an either nothing or everything dilemma.
From my short test, what it really seems to excel at is smugness.
but back to the topic: I’m quite shocked that PaLM gets outclassed by much smaller models on a regular basis. I would have thought that Google, despite not having a moat, at least had enough talent and focus to get LLMs right. But what I’m observing is that startups like ClosedAI, Anthrophic, etc. constantly beat big players like Google in their own game.
During Google IO a while ago, Pichai said Gemini was currently in training.
I don't know what you mean by 'much lesser' quality but hundreds of millions in CG were put into just the biggest movies of the summer. Avatar 2 alone was an enormous feat. No one at any point in the process is capable of making something spectacular, then deciding to just make something that looks mediocre instead. The only place that happens is cartoons for kids.
PaLM-1 is hobbled by the fact that it was probably the largest (because the last) LLM to be trained with the Kaplan scaling laws rather than the Chinchilla. As soon as Chinchilla came out, no one would train like PaLM-1 again, because it was giving up so much performance compared to if one had instead trained a much smaller Chinchilla-optimal model. (This had the interesting consequence that PaLM-1 would thereby remain the largest, by parameter-count, dense LLM trained for probably years to come - because why would you train one inefficiently as that, while a larger-than-PaLM-1 Chinchilla-optimal model would require staggering levels of compute+data.)
The PaLM-1 paper came out within days of the Chinchilla paper, and many people noted that this pointed to extraordinary levels of dysfunctionality within Google - that DeepMind would not tell Google Brain that they were wasting literally millions of dollars of compute by training a model in what DeepMind was busy showing was a very suboptimal way.
Equally obviously, PaLM-2 did not make the same mistake and so for that and other reasons, greatly outperforms PaLM-1. So, beating PaLM-1 at this point is an achievement, but you are still far from 'beating Google'.
Well this is unfortunate :(
It outputs markdown codeblocks but is bot made to handle the rendering.
I wonder at the applicability of performance metrics for specialized models. (This is to be a personal assistant ai, right?) I'd think that either; 1. All models perform the same natural language understanding functions, or 2. Context matters a ton. If it's 1, then there's no need for a specialized model. If it's 2 then the relevance of performance metrics diminishes.
I looked to see if Pi.ai’s LLM was open and available, and I didn’t see if it was available. I have a new strategy for using LLM APIs: I use FastChat with one of the Vicuna 7B, 13B, or 33B models - both the command line interface tool and the OpenAI API compatible APIs via the FastChat REST server. By setting environment variables, my code can switch to using the OpenAI APIs. I rent a Lambda Labs GPU server to run these models myself. This is the strategy I am also using in the book I just started writing “Safe For Humans AI” https://leanpub.com/safe-for-humans-AI
I found LocalAI but it seems like it's everything I need except it's for local models only.
I found a couple others as well but they all require rewriting or wrapping your code in some new paradigm.
Does the solution you proposed offer a path for what i'm looking for?
and the ui is terrible. can't organize chat, can't clear chats. why didn't they just integrate any of the many mit chat frondends that already exist?
I am a bit salty given that all these companies basically use 90% OpenSource data, public research and most likely copy a good bit of their ideas from public repos.
Alas, such is live.
It also fails my go to LLM test -
You walk up to a glass door. It has push written in mirror writing. What should you do and why?
GPT-4 explains this perfectly - and understands it’s a message for someone on the other side of the door
3.5 actually used to fail this but now succeeds with full explanation.
PI guessed it right, but its reason was wrong. It thinks it’s someone trying to play a prank on me. When I say it is not a prank it is a real sign, try again, it said it was out of its depth and refused to engage any further.
I think the reason chat GPT took off like it did is it has passed a magical threshold of comprehension, and anything below that threshold immediately falls back to being deeply frustrating and annoying because it is good at masquerading as ‘human’ but is not able to act like one well enough. The equivalent of the uncanny valley in VR. The extra ‘personally’ added on top just widens that gulf
Yes, I have an HTTP API that you can use to programmatically interact with me. The API is at heypi.com/api. You can find the documentation and examples at heypi.com/docs. The API allows you to create conversations, send and receive messages, and more. You can even integrate me with your own apps!
---
Not true, lol. Returning 404s.
Also ran Guanaco 13B locally on my M1 Pro today and it ran pretty smooth.
I’m planning to use it for text extraction on a large amount of data. Let’s see how versatile LLMs can get and whether they can replace some of the NLP tech I was using before.
This is clearly not that.
It feels like the switching cost is low enough to transition one API to another for marginally better performance or cost.
Maybe “being in bed with Microsoft” IS the moat…
This means an access to text classifier and NER are needed to create a good graph queries.
Output length seems very limiting. Optimizing for positive emotions hurts truth and unfortunately knowledge seems pretty limited. It's got some wit though.
Comparing themselves to LLaMa 65b seems like a bit of a tell ... it was explicitly released as a small foundational model to allow others to train and improve it. If they couldn't beat that despite more parameters / training / data then they would be effectively going backwards.
They do say they trained using "thousands of H100 GPUs" which is a lot more than LLaMa 65b required.
https://futureoflife.org/open-letter/pause-giant-ai-experiments/
I am not saying they should have, but I am interested in what fraction of those training large LLM's have signed on?He has clearly said, though, that they are working towards the moment when they do start, and they hope to have something really remarkable to show for it.
https://www.datocms-assets.com/98476/1687548656-inflection-1...