781 karma · joined April 30, 2014
once you hide the reasoning, remove the knobs, and let the model choose its own effort, it gets much harder to tell whether the model got worse or just got harder to inspect.
that’s a real shift. less tool, more black box.
Exclude movies with very low number of rating or potentially very low scores too.
The long tail reduction would be significant
But it was on purpose not trained on the big “web crawled” datasets to not learn how to build bombs etc, or be naughty.
So it is the “smartest thinking” model in weight class or even comparable to higher param models, but it is not knowledgeable about the world and trivia as much.
This might change in the future but it is the current state.
It was available even before this, all they changed is that law abiding citizens can put apps in the App Store and charge money for it.
(More importantly law abiding companies can build on and fine tune it in hopes of profit).
BYOT - bring your own tests style.
Gives a better picture of real-world performance and more robust against contamination.
They collected over 6000 and 1500 votes for Mixtral-8x7B and Gemini Pro.
While ELO ratings are widely used to rank performance in Chess or among sports teams, here's a disclaimer by the makers of the leaderboard:
---
> Please note Arena is a "live eval" and pretty much a sampling process to estimate models capability.
> That's why we show the confidence intervals through bootstrapping. Statistically, these models (e.g., GPT-3.5, Mixtral, Gemini Pro) are very close and only looking at their ranking can be misleading.
---
So, we've been quietly building something new for running AI experiments and deploying models ...
Our Lightning AI Studios let you switch between different machines and GPUs flexibly in the same environment without any setup steps.
Everything can be accessed via your browser and supports
- VSCode
- Jupyter Notebook
- a regular terminal
- a control pane for multi-node jobs
- ... many, many collaborative and extra features
And there's no installation or setup step required at all.It's basically what I've been using internally as a productivity tool for the last few months to run AI experiments.
(*there's also a demo video in the linked tweet)
---
A persistent GPU cloud environment.
Code online. Code from your local IDE. Prototype. Train. Serve. Multi-node. All from the same place.
No credit card. 6 Free GPU hours/month.
Only Precise and Creative modes use GPT-4
https://twitter.com/emollick/status/1732495030143549541
Also see:
An Opinionated Guide to Which AI to Use: ChatGPT Anniversary Edition
https://www.oneusefulthing.org/p/an-opinionated-guide-to-whi...
But you’ll need it in less and less everyday scenarios and time goes on
Just like we need to write less and less assembly by hand
The whole universe might just be a stochastic swirl of milk in a shaken up mug of coffee.
Looking at something under a microscope might make you miss its big-picture emergent behaviors.
Some third party did these tests first (in article and spread on social) to which the makers of Claude are responding.
I knew it’s a weird test right when I first encountered it.
Interesting that the Claude team felt like it’s worth responding to.
But these LLMs were fine tuned on realistic human question and answer pairs to make them user friendly.
I’m pretty sure the average person wouldn’t prefer an LLM whose output is always playing grammar Nazi or semantics tai chi on every word you said.
There has to be a reasonable “error correction” on the receiving end for language to work as a communication channel.
Question is how good they are are retaining minute details from extremely long context, say 200k tokens.
That’s the frontier Claude and now GPT-4 Turbo are pushing
We tend to remember out of place things more often.
E.g. if there was a kid in a pink hat and blue mustache at a suit and tie business party, everybody is going to remember the outlier.
Now we have to tinker with it to learn instead of read Textbooks
GPT-4 Turbo is more watered down on the details with long context
But also it’s a newer feature for OpenAI, so they might catch up with next version