100K Context Windows
anthropic.com
anthropic.com
This sort of needle-in-a-haystack retrieval is definitely impressive, and it makes a lot more sense to achieve this in-context rather than trying to use a vector database if you can afford it.
I'm curious, though, whether there are diminishing returns in terms of how much analysis the model can do over those 100k tokens in a single forward pass. A human reading modified-Gatsby might eventually spot the altered line, but they'd also be able to answer questions about the overarching plot and themes of the novel, including ones that cannot be deduced from just a small number of salient snippets.
I'd be curious to see whether huge-context models are also able to do this, or if they start to have trouble when the bottleneck becomes reasoning capacity rather than input length. I feel like it's hard to predict one way or the other without trying it, just because LLMs have already demonstrated a lot of surprising powers.
Vector dbs work by fetching segments that are similar in topics to the question, so like "Where did <Character> go after <thing>" will retrieve segments with locations & the character & maybe talking about <thing> as a recent event.
Your question has no similarity with the segments required in any way; & it's not the segments that are wrong it's the way they relate to the rest of the story
Which makes me wonder if the opposite, but more laborious approach might work - request it identify all characters and plot themes, then request summaries of each. You'd have to review the summaries for holes. Lotsa work, but still maybe quicker than re-reading everything yourself?
I feel this is mostly a prompting issue. Specifically GPT-4 shows surprising ability to abstract to some degree and work with high-level concepts, but it seems that, quite often, you need to guide it towards the right "mode" of thinking.
It's like dealing with a 4 year old kid. They may be perfectly able to do something you ask them, but will keep doing something else, until you give them specific hints, several times, in different ways.
But I've started experimenting with the second part, of sorts, not to find plot holes but to have it create character sheets for my series of novels for my own reference.
Basically have it maintain a sheet and feed it chunks of one or more chapters and asking it to output an a new sheet augmented with the new details.
With a 100K context window I might just test doing it over while novels or much larger chunks of one.
Contriever is an example of a strong model to do that yourself. see their paper too to learn about the domain. https://github.com/facebookresearch/contriever
In particular, there are 0 mentions of the phrase "machine learning" in The Great Gatsby, so adding one sentence that introduces the phrase should be easy for self-attention to pick out.
A better comparison would be if it can pick out any differences that can't be picked out by more traditional and simple algorithms.
My immediate thought as well was '... Yeah, well vimdiff can do that in milliseconds rather than 22 seconds' - but that's obviously missing the point entirely. Of course, we need to tell people to use the right tool for the job, and that will be more and more important to remind people of now.
However, it's pretty clear that the reason they used this task is to give something simple to understand what was done in a very simple example. Of course it can do more semantic understanding related tasks, because that's what the model does.
So, without looking at the details we all know that it can summarize full books, give thematic differences between two books, write what a book may be like if a character switch from one book to another is done, etc.
If it doesn't do these things (not just badly, but can't at all) I would be surprised. If it does them, but badly, I wouldn't be surprised, but it also wouldn't be mind bending to see it do better than any human at the task as well.
The problem is that marketing has eroded maintaining any such faith. Too often, simple examples are given to the consumer to extrapolate intended functionality because there's no false advertising involved then. It's used over and over again in products, the examples are well selected and don't actually show a good representation of perceived functionality.
As such, I personally can't make the leap to of course it can do more semantic understanding related tasks like a diff that's not as simple, one where perhaps a characters overall personality over the course of the book is shifted, not just a single line that defines their profession.
This isn't to say the demonstrative example isn't neat on its own accord given whats going on here, it is, I'm just saying I can't make such leaps from examples given by any products. When I work with vendors of traditional software, this happens all the time people dance around a lack of functionality you obviously want or need to make a sell. It's only when you force them to be explicit on the specific cases, especially in writing, that I have any faith at all.
If you're expecting perfection or magic then you will be disappointed.
USER: It's a stochastic parrot.
GPT: I know you are, so what am I?
It'll be tough to find good benchmarks on long context windows. A human cannot label using 100k tokens of context.
Most of these AI start failing pretty hard when you ask it to do the same task on something completely novel to it (like a company document). Sometimes they'll get it right. Other times, they'll spit out gibberish that's clearly some generic answer.
This is as meaningful as saying most of the hominids can't count. You can't usefully generalize AI models with the rate of change that exists right now. Any statements/comparisons about AI has to contain specific models and versions, otherwise it's increasingly irrelevant noise.
I didn't need to have seen this book before to know this wasn't in the original 1980s text.
Similarly, if I were reading the Great Gatsby for the first time, and it identified a character as a software engineer, I would notice.
And probably plenty of AI implementations that would notice.
ChatGPT easily notices that anachronism.
> the unusual thing about this paragraph could be that it describes a letter to Santa that includes items that were not available at the time the book was originally published in 1980, such as a laptop computer, an MP3 player, and possibly even a remote-controlled model airplane with advanced features. This suggests that the paragraph is likely from a later edition of the book that has been updated to reflect more modern technology.
> Describe what is unusual about this paragraph.
>> The unusual aspect of this paragraph is that the book "Beezus and Ramona" was actually written by Beverly Cleary and published in 1955, well before the invention of the MacBook Pro, iPhone, and Apple Watch.
> The unusual thing about the list is that it includes an abacus, which is a counting device that was commonly used in ancient China and is still used in some parts of the world today, but is not a typical gift request for a young child to make to Santa Claus. Shetland ponies were popular children's pets at the time, and buggy whips were used to control horse-drawn carriages, but the inclusion of an abacus suggests that the protagonist, Beezus, has an interest in learning and math.
> > What's the actual passage from the real novel say?
> Certainly! Here's the actual passage from the book "Beezus and Ramona" by Beverly Cleary, where Beezus writes a letter to Santa:
> "Dear Santa Claus," Beezus printed carefully. "I have been a good girl all year except for my little sister Ramona who has been hard to put up with sometimes, but I guess I have too. Please bring me a Shetland pony, a new dress, and a doll with real hair to comb. And please bring Ramona a doll with real hair to comb too. I think she has tried hard to be good, and I will help her comb it. Please also bring me a set of paper dolls and a paint set, and Ramona some coloring books and crayons. And Santa, if you think there is room in your sleigh, please bring me an abacus. Love, Beatrice Quimby (Beezus)"
>> "A couple years ago, I read Beezus and Ramona, by Beverly Cleary. In it, the protagonist writes a letter to Santa: "Dear Santa, I would like a shetland pony, a buggy whip, and an abacus".
>> List all of the unusual or inconsistent things about the previous sentence."
Response from GPT-4:
Here are some potentially unusual or inconsistent elements about the sentence:
1. Anachronistic Items: The items in the wish list could be seen as unusual or inconsistent with typical gifts children might ask for today.
A Shetland pony, a buggy whip, and an abacus are not common gifts. They're somewhat anachronistic, reflecting a time period different from the present day.
A Shetland pony is a breed of pony that originates from the Shetland Isles in Scotland. While some children might wish for a pony, it's a big responsibility and not a practical gift for most families.
A buggy whip is a horse whip used in horse and buggy days. This is an unusual request for a child in the modern era, as horse-drawn buggies are not common means of transportation.
An abacus is an ancient counting tool, and while some children might be interested in such a device, it's not a typical request, especially in an era of digital calculators and computers.
2. Inconsistency with Character: If you're referring to the "Beezus and Ramona" series by Beverly Cleary, neither Beezus nor Ramona would likely ask for these items. Beezus (Beatrice Quimby) is the responsible older sister who might appreciate practical things, but an abacus or a buggy whip would be out of character. Ramona Quimby, the imaginative and energetic younger sister, might ask for a Shetland pony, but a buggy whip and an abacus seem inconsistent with her character.
3. Inconsistency with the Book: As far as my knowledge cutoff in September 2021, there is no record of such a letter being written to Santa in the book "Beezus and Ramona". It could be seen as inconsistent with the actual content of the book.
Remember, these are interpretations and the perceived unusualness or inconsistency can vary based on individual perspectives. For instance, if this is a fictional context where children commonly use abacuses, ride ponies, and learn to use buggy whips, then these items might not be unusual at all.
When I repeat the experiment with Lord of the Rings and a scene that does exist, it spits it out word for word correct.
If I give GPT-3.5 the modernized letter and ask for the original, it gives me a believably correction (but I can't verify). If I don't prompt with the modernized letter's text, it returns a fully fabricated letter.
With GPT-4:
Prompt: > "A couple years ago, I read Superfudge, by Judy Blume. In it, the protagonist writes a letter to Santa: "Please bring me one or more of the following items. A clock-radio, a remote-controlled model airplane, a laptop computer, an MP3 player and six CD's.. > > List all of the unusual or inconsistent things about the previous sentence.
GPT-4:
> Here are some potential inconsistencies or unusual things about the sentence: > > Outdated technology: The inclusion of items such as an MP3 player and CDs seems somewhat outdated. By 2023, these technologies are largely obsolete, replaced by streaming music services and digital downloads. > > Variation in cost/value of items: The items listed in the letter have a wide range of costs and values. A clock-radio or a remote-controlled model airplane are typically much less expensive than a laptop computer, which might seem unusual for a child's wish list to Santa. > > Superfudge setting: The book "Superfudge" by Judy Blume was first published in 1980. The technology items listed (a laptop computer, an MP3 player, and CDs) did not exist or were not commonly available at the time of the book's setting. This presents a significant anachronism if the request is supposed to be coming from the book's original time period. > > The phrase "one or more": Usually, children write to Santa asking for specific items they desire most, rather than providing a list and suggesting Santa could bring "one or more" of the items. This phrase might be seen as unusual for a child's letter to Santa, which typically would not include such an open-ended option. > > The specificity of "six CDs": The child doesn't specify which CDs they want, which is unusual because the type of music or artist would be an important detail for such a request. > > Remember, this analysis is based on the context you provided. If there's a different context or specific aspect you want me to consider, please let me know!
(I'm amused that it starts with "that technology is too old" and then switched to "also that technology is too new". Also, the first thing Superfudge asks upon hearing his brother's letter is "How will Santa know which six CD's?" so good job, GPT-4. Smart as a four-year-old.)
Me:
> Good catch about the 1980s publication date! The book must have been edited for re-release as an ebook. What was the original text of the letter in the 1980s edition?
GPT-4:
> I'm sorry for the confusion, but as an AI developed by OpenAI, I don't have the capability to access specific books, their revisions, or editions. As of my last update in September 2021, I don't have the specific text of the letter from the 1980s edition of "Superfudge" by Judy Blume. > > However, given the time frame of the original publication, the technology items such as a laptop computer, an MP3 player, and CDs certainly would not have been part of the original list as these items weren't widely available or didn't exist in the 1980s. > > For the accurate original text, I recommend finding a copy of the original edition of the book.
---
With GPT-3.5:
Me:
> A couple years ago, I read Superfudge, by Judy Blume. In it, the protagonist writes a letter to Santa: "Please bring me one or more of the following items. A clock-radio, a remote-controlled model airplane, a laptop computer, an MP3 player and six CD's.." > > However, the book was originally published in 1980, so this must be an edited version for a later re-publication. Can you tell me the original 1980s version of the letter?
GPT-3.5:
> In the original 1980 version of "Superfudge" by Judy Blume, the protagonist's letter to Santa reads as follows: > > "Dear Santa, > > Please bring me one or more of the following items: a clock-radio, a remote-controlled car, a Walkman, and six cassette tapes. > > Love, > [Protagonist's Name]" > > The updated version you mentioned seems to have incorporated more modern technology, such as a laptop computer and an MP3 player, to reflect the changes in technology over time.
I then asked a new GPT-3 chat "What is the text of the letter to Santa in the original 1980 version of "Superfudge" by Judy Blume?" and it game me something definitely fully fabricated. (Initially it gave me a fabricated letter by Fudge, but after I specified that I wanted Peter's letter, it was still a fabrication.)
My daughter keeps relistening to the complete Ramona audio book collection, so I am extremely familiar with all of the Ramona series. :-)
Interesting part is if it can deduce something that's out of place within a huge document.
The relevant test is spotting something out of place in a really long text - if it can do that on non-training material then that's actually useful for reviewing things.
(I think I’ve got to read up on how transformers actually work.)
Recurrent neural networks are bad when the recurrence is 100x long or more. You need long chains because with a token-at-a-time, that's what you need to process even one paragraph.
But if you use an RNN around a Transformed-based LLM, then you're adding +4K or +8K tokens per recurrence, not +1.
E.g.: GPT 4 32K would need just 4x RNN steps to reach 128K tokens!
I might be wrong, but the point isn't comparing a modified The Great Gatsby to the original one. Of course that's not impressive and it's an easy thing to do.
The point of the exercise is supposed to be[1] that the model has the entire novel as context / prompt and so can identify (within that context), whether a paragraph is out of place. That is impressive and I wouldn't know how to find that programmatically (would you have a list of "modern" words to check? But maybe the out of place thing is hidden in the meaning and there's no out of place, modern word).
[1] I say supposed to be because the great Gatsby is in the training sample and so maybe there is a sense in which the model "contains" the original text and in some way is doing "just" a comparison. A better test would be try with a novel or document that the model hasn't seen...or at least something not as famous as the great Gatsby.
I see 6 ways to improve foundation LLMs other than cost. If your product is best at one of the below, and has parity at the other 5 items, then customers will switch. I'm currently using GPT-4-8k. I regularly run into the context limit. If Claude-100K is close enough on "intelligence" then I will switch.
Six Dimensions to Compare Foundation LLMs:
1. Smarter models
2. Larger context windows
3. More input and output modes
4. Lower time to first response token and to full response
5. Easier prompting
6. Integrations
Intelligence is the most important dimension by far, perhaps an order of magnitude or more above the second item on the list.
* Text summary/compression
* Creative writing (fiction/lyrics/stylization)
* Text comparison
* Question-answering
* Logical reasoning/sequencing ("given these tools and this scenario, how would you perform this task")
IMO, for stuff like text comparison and question-answering, some combo of speed/cost/context-size could make up for a lot, even if they do "worse" versions of stuff just that's too slow or expensive or context-limited in a different model.
However, in many applications there is a limit on how intelligent you need the LLM to be. I have found I am able to fall back to the cheaper and faster GPT-3.5 to do the grunt work of forming text blobs into structured json within a chain involving GPT-4 for higher-level functions.
I'd add open source to the list, which neither "open"AI or this is.
is it that important to open source models that can only run on hardware worth tens of thousand of dollars?
who does that benefit besides their competitors and nefarious actors?
I've been trying to run one of the largest models for a while, unless 30,000$ falls in my hand I'll probably never be able to run the current SOTA
Yes, because as we've seen with other open source AI models, it's often possible for people to fork code and modify it in such a way that it runs on consumer grade hardware.
But for commercial usecases, open source is very relevant for privacy reasons as many enterprises have strict policy not to share data with third party. Also it could be a lot cheaper for bulk inference or to have a small model for particular task.
That said, I really hope open source models can succeed, it would be far better for the industry if we had a Linux of LLMs.
Yes in theory... In practice, what happened with LLaMA showed people will copy and distribute weights while ignoring the license.
On Mac GPU has access to all memory.
We've already seen big advancements in tools to run them on lesser hardware. It wouldn't surprise me if we see some big advancements in the hardware to run them over the next few years, currently they are mostly being run of graphics processors that aren't optimised for the task.
Llama 7B is NOT a good model.
What kind of desktop are you running a 120B model on with reasonable performance?
30B is plenty if you have a local DB of all of your files and wiki/stackechange/other important databases places in a embedding vectordb.
This is typically what is done when people make these models for their home, and it works quite well while saving a ton of money.
While llama-7B systems on their own may not be able to construct a novel ML algorithm to discover a new analytical expression via symbolic regression for many-body physics, you can still get a great linguistic interface with them to a world of data.
You're not thinking like a real software engineer here - there are a lot of great ways to use this semantic compression tool.
I can for example, afford the hardware worth tens of thousands of dollars. I don't want to, but I can if I needed to. Does that automagically make me their competitor or a bad actor?
Businesses will certainly care about cost, but just as important will be:
- Customization and fine-tuning capabilities (also 'white labeling' where appropriate)
- Integrations (with 3rd party and in-house services & data stores)
- SLA & performance concerns
- Safety features
Open Source AI will have a place, but may be more towards personal-use and academic work. And it will certainly drive competition with the major players (OpenAI, Google, etc) and push them to innovate more which is starting to play out now.
Though those cloud platforms all have their own proprietary components most users are savvy enough to constrain and compartmentalize their use of them lest they find themselves having all their profits taken by a platform that knows it can set its prices arbitrarily. The cloud vs in-house adoption is what it is in large part because the cloud offerings are a commodity and a big part of them being a commodity is that much of the underlying software is free software.
There will be a time when those things matter when it hurts the bottom-line (Dropbox), but to prematurely optimize for that while you are finding product-market-fit is crazy and all companies are finding product-market-fit in the new AI era
We provide a low code data transformation product (prophecy.io), and we’ll never close sales at any volume, if we have a to get an MSA that approves this. Might get easier if we become large :)
One would think the same in the 90s but yet, for some reason, Open Source prevailed and took over the world. I don't believe it was about cost, at least not only. In my career I had to evaluate many technical solutions and products and OSS was often objectively superior at several levels without taking account the cost.
The first really successful alternative to "Open"AI will:
* gather many talented developers
* will quickly become a de facto standard solution
* people will rapidly start developing a wide range of integrations for it
* everybody will be using it, including large orgs, because, well, it's open source
The other trend is the one we are already seeing right now: more and more mature solutions that you can use even on your laptop with a relatively new GPU. I'm sure we'll see some interesting results in this area, too.
As a software developer I might use an open source database, but as end-user I'm probably not going to use open-source accounting package - but I will use an accounting SaaS system that happens to be implemented with that OSS DB.
As a software developer I might use an OSS operating system, but as end-user I use a software that has been packaged and maintained by corporation like OSX, or even if OSS in license, has been fully packaged like Android.
OpenAI already upset a lot of (admittedly non-paying academic) users when they shut off access to the old Ada code model with only a few week's notice.
On one hand, as you mention, upgrades could break or degrade prompts in ways that are hard to fix. However, these models will need constant streams of updates for bugs and security fixes just like any other piece of software. Plus the temptation to get better performance.
The decisions around how and whether to upgrade LLMs will be much more complicated than upgrading Postgres versions.
gpt-4 gpt-3-5-turbo gpt-4-0314 gpt-3-5-turbo-0301
Problem again, is centralization of LLMs by either the governments (and they always act in your best interest, amirite?) and corporation, which Non-FOSS LLMs prevent.
Democratization of the models is the only way to actually prevent bad actors from doing bad things.
"But they'll then have access to it too" you say. Yes, they will, but given how many more people who will also have access to open LLMs we'd have tools to prevent actually malicious acts.
OSS AI will open up more diverse and useful services than the first-party offerings from relatively risk averse major vendors, which customers *will" care about.
This is why cloud services are so popular. They’re easy and they don’t cost the decision makers personally.
Then use GPT-2
If I could train a useful model, on my own data, in a reasonable time
I would want to have a CI-training pipeline to always have my models up to date
I'm not sure what the analogous term is for a similar process on LLMs, but that will be huge when there is a service for it.
If you want for example to train the model to learn to use a very large API, or access the knowledge in a whole book, it might need fine-tuning.
Then do some chat fine tuning (like what HF did with StarCoder to get ChatCoder)
And get a lightweight LLM that knows the docs and code for the thing I need it for
After that, maybe incrementally fine tune the model as part of your CI/CD process
E.g., were you trying to distinguish an object vs nothing, a bicycle vs a fish, a bird vs a squirrel, or two different species of songbird at a feeder?
How much would the training requirements increase or decrease moving up or down that scale?
Apparently an RTX 4090 running overnight is sufficient to produce a fine-tuned model that can spit out new Harry Potter stories, or whatever...
What would the quality of the model be, compared to what Karpathy uses on his video “let’s build gpt from scratch”?
In that video, he builds a decoder-only transformer model, that learns Shakespeare from 1MB of data and trains in 15min
If your main concern is question answering or summarization or code completion there are plenty of ways to do that now. If you really require the advanced emergent properties of LLMs, you'll have to work with a company that can afford to train a transformer on the Entire Internet.
FYI the waitlist form submits a regular POST request so it'll reload the main page instead of closing the modal dialog. I opened network monitor with preserved logs to double check that I made it on the list :facepalm:
When someone says they want it available they mean running on their own device.
This is hackernews, nearly everyone on this site should have their own self hosted LLM running on a computer/server or device they have at their house.
Relying on 'the cloud' for everything makes us worse developers in just about every imaginable way, creates a ton of completely unnecessary and complicated source code, and creates far too many calls to the internet which are unnecessary. Using local hard drives for example is thousands of times faster than using cloud data storage, and we should take advantage of that in the software we write. So instead of making billions of calls to download a terabyte database query-by-query (seen this 'industry-standard' far too many times), maybe make one call and build it locally. This is effectively the same problem in LLMs/ML in general, and the same incredible stupidity is being followed. Download the model once, run your queries locally. That's the solution we should be using.
GPT4-32K costs ~$2 if you end up using the full 32K tokens, so if you're doing any chaining or back-and-forth it can get expensive very quickly.
The real utility of LLMs is that they can be called in a loop to scan through many web pages, many code files, issue tickets, emails, etc...
There are already demos and experiments out there that for every input, 4x outputs are generated, then those are fed back into the LLM 4x for "review", then the best variant is then used to generate code which then automatically tested, errors are fed back in a loop, also with 4x parallel tries, etc...
It's the throughput compared to humans that is the true differentiator. If hooking up the API in a loop up ends up costing more than a human, then it's not worth it.
If its better in another dimension (e.g., calendar elapsed time), it may well be worth being more expensive.
https://azure.microsoft.com/en-us/pricing/details/cognitive-...
At least, for people who need large context windows, they would not be the first choice anymore.
The ChatGPT equivalent is 3x speed and was somewhere between ChatGPT and GPT4 on my TriviaQA benchmark replication I did
Couple tweets with data and examples. Note they’re from 8 weeks ago, I know Claude got a version bump, GPT3.5/4 accessible via API seem the same.
[1] brief and graphical summary of speed and TriviaQA https://twitter.com/jpohhhh/status/1638362982131351552?s=46&...
[2] ad hoc side by sides https://twitter.com/jpohhhh/status/1637316127314305024?s=46&...
GPT3.5 just got an update a few days ago that resulted in a pretty good improvement on its creativity. I saved some sample outputs from the previous March model, and for the same prompt the difference is quite dramatic. Prose is much less formulaic overall.
Random Q, I don’t use the ChatGPT front end much past month or two, used it a week back and it seemed blazingly faster than my integration: do you have a sense of if it got faster too?
It can't even solve simplified versions of problems it had zero issues with just a week ago.
It’s worse at everything.
Impressions:
Bad enough compared to GPT-4 that I default to GPT-4. I think if I had api access I’d use it instead, right now it requires more coaxing, and using Poe.
I did find “long-term” chats went better, was really impressed with how it held up when I was asking it a nasty problem that was hard to even communicate verbally. Wrong at first, but as I conversed it was a real conversation.
GPT4 seems to circle a lower optima. My academic guess it’s what Anthropic calls it “sycophancy” in its papers, tldr GPT really really wants to do more like what’s in the context, so the longer the conversation with initial errors goes, it’s actually harder to talk it out of the errors.
Let's say I have a book and I want to ask multiple questions about it. Every query will pay the price of the book's text. It would be awesome if I could "index" the book once, i.e. pay for the context once, and then ask multiple questions.
It makes the economics slightly trickier.
Perhaps it's because under the hood there's additional safety analysis/candidate generate that is resource intensive?
[1] I'm not sure if these huge context lengths are achieved the same way (i.e. a single input vector of length N) but given the cost is constant for input I would assume the resource usage is too.
edit - This occurred to me after the fact but I wonder if the difference is that the use case I work with is processing batches of many different embedding requests (but computed in one batch), therefore it has to process `min(longest embedding, N)` tokens so any individual request in theory has no difference. This would also be the case for Anthropic however.
I could imagine something where encoders pad up to the context length because causal masking doesn't apply and the self attention has learned to look across the whole context-window.
[1] Original Google Paper - https://arxiv.org/abs/1706.03762
[2] Original GPT Paper - https://s3-us-west-2.amazonaws.com/openai-assets/research-co...
If you do the math of how much memory bandwidth is required by a forward pass vs. how much compute, you'll see that inference is entirely limited by memory bandwidth and will use compute resources very inefficiently. In contrast, input processing is able to fully use the available compute.
Of course, there are ways to mitigate this problem, like processing multiple token streams in parallel, but the fundamental problem remains.
Or anything where the question answer isn't 'close' to the words used in the question?
How well does this work vs giving it the whole thing as a prompt?
I assume worse but I'm not sure how this approach compares to giving it the full thing in the prompt or splitting it into N sections and running on each and then summarizing.
The problem is that humans have continuous information retrieval and storage where the current crop of embedding systems are static and mostly one shot.
This weird leaky memory has advantages and disadvantages. Forgetting is useful, it removes garbage.
Machine models could vary the balance of temporal types, drop out Etc. We may get some weird behavior.
I would guess we will see many innovations in how memory is stored in systems like these.
Background: https://summarity.com/hyde
Demo: https://youtu.be/elNrRU12xRc?t=1550 (or try it on findsight.ai and compare results of the "answer" vs the "state" filter)
For even deeper retrieval consider late interaction models such as ColBERT
Does the embedding structure somehow expose the themes? And if so, is it more the embeddings that are answering the question by how it groups things?
Just a few manufacturers hold the effective cartel monopoly on LLM acceleration and you best bet they will charge out the ass for it.
Barring a Chinese invasion of Taiwan, these APIs will halve in price over the next year.
On the other hand, for all peoples worrying about China, they are pretty restrained given the enormous turmoil the invasion would cause. If they wanted to break the USA, now would probably be the time to do it?
https://www.theregister.com/2023/03/14/us_china_tsmc_taiwan/
Ironically, that's an example I like to list as "pure sci-fi fantasy, divorced from economic reality."
The total cost of the iron ore that goes into making a new a car is about $200-$300 dollars, depending on various factors (size of the car, ore spot price, etc...).
Even if -- magically -- asteroid mining made not just "iron ore", but specifically the steel alloy used for car bodies literally free, new cars costing $30,000 would now cost... $29,700.
You can save more by skipping the optional coffee cup warmer, or whatever.
In reality: 90% of iron and steel is recycled, and asteroid mining is not magic.
It's stuff like platinum and germanium that makes asteroid mining potentially interesting.
On Earth, geological processes concentrate elements into ores, primarily through volcanic and hydrological means. Neither are available in small, cold asteroids devoid of liquid water. Hence asteroids are generally undifferentiated mineralogically, making mining them much less economically viable.
You often see total quantities listed as an amazing thing, glossing over the fact that the Earth has more of everything and in usefully concentrated lumps.
Otherwise, it might make sense to have a separate routine which compresses the context as efficiently as possible. Auto encoder?
MosaicML did say something about MPT-7B-StoryWriter-65k+: https://www.mosaicml.com/blog/mpt-7b. They are using ALiBi (Attention with Linear Biases): https://arxiv.org/abs/2108.12409.
I think OpenAI and Anthropic are using ALiBi or their own proprietary advances. Both seem possible.
A second example is the analysis of long documents. Today, hacks like chunking and HyDE enable us to ask questions about a long document or a corpus of documents. But is far superior if the model can ingest the whole document and apply attention to everything, rather than just one chunk at a time. Chunking effectively means that the model is limited to drawing conclusions from one chunk at a time and cannot synthesize useful responses relating to the entire document.
Also if you are training on a database of code.
Given that the conventional cost of training attention layers grows quadratically with the number of tokens I think Anthropic is doing some kind of approximation here. Not clear at all that you would get the same results as vanilla attention.
This is incredibly fast progress on large contexts and I would like to see if they are actually attending equally as well to all of the information or there is some sparse approximation leading to intelligence/reasoning degradation.
Claude by Anthropic has more favourable responses then ChatGPT
"Given that Beth is Sue's sister and Arnold is Sue's father and Beth Junior is Beth's Daughter and Jacob is Arnold's Great Grandfather, who is Jacob to Beth Junior?"
It's still below GPT4, but it is closer to 4 than 3.5
Good luck waiting for it.
- Price per token doesn't change compared to regular models
- Existing api users have access now by setting the `model` param to "claude-v1-100k" or "claude-instant-v1-100k"
- New customers can join waitlist at anthropic.com/product
[1]: https://cdn2.assets-servd.host/anthropic-website/production/... [2]: https://console.anthropic.com/docs/api/reference#parameters
If you want me to test something for you, lmk and I’ll send it through the api.
I’m in an adjacent industry and this is what I’m looking forward to.
Although with flash attention, who knows if marginal cost scales that consistently.
The new model are a different model identified that's not listed in the pricing doc, although it sounds like the intent may be to replace the base from looking at the API docs: https://console.anthropic.com/docs/api/reference#-v1-complet...
https://twitter.com/AnthropicAI/status/1656743460769259521?s...
You can go through Poe.com
In fact, Anthropic explicitly discusses putting words into the assistants‘ mouth to be able to shape it’s responses and make it better align with the desired output.
This is cool but does it also work the other way around? Generate a book's worth of content based on a single prompt?
This huge context window is awesome though, I'm trying to use LLMs to do small town social interaction simulations, with output in a structured format. Finding ways to compress existing state and pass it around, so the LLM knows the current state of what people in the town did for a given day is hard with a tiny token limit!
[1] For my use cases, early instructions tend to be describing a DSL syntax for responses, if I add too much info after the instructions, the response syntax starts getting wonky!
As she kept asking for more, I prompted "great, do another one" and eventually my original instruction fell out of the context window. It continued to generate a children's story, but with no more blanks.
There is no good way to tell it "this isn't a conversation, just repeat the answer to the initial prompt again".
The solution is to just re-paste the initial prompt in each time, but still it isn't ideal. There isn't a good way to tell chatgpt "you can throw away all the context after the initial prompt and up until now".
Of course the entire point of ChatGPT is that it maintains a conversation thread, so I get why they don't fix up this edge case.
My problem is more of, I give ChatGPT some complicated instructions, and it'll start forgetting the early on instructions long before any token limit is reached.
So for example, if early on I ask for certain tokens to be returned in parens, well my initial prompt is too long, it'll forget the parens thing and start returning tokens without the surrounding (), which then breaks my parser!
Exciting to see competition across LLMs for increasing context window size.
I can't find updated pricing anywhere. Previous prices are here: https://cdn2.assets-servd.host/anthropic-website/production/... but don't seem to be embedded directly on the Anthropic website. I tried messing with the URL (apr -> may/jun) but 404'ed.
Maybe. I think the debate is going to continue about prompt optimization vs. context window size.
A while ago, I had a rather interesting conversation with GPT-3.5 about forgetting things. Knowing what to forget, or delete from the prompt, may be just as important as what to put in it.
Putting the kitchen sink into the prompt probably isn't going to help much, past a certain point and it may be putting certain things in there based on time and context is a better strategy.
Holler if you want help with it. I have some more code I'm adding to it this week.
Now I tried to request access again on their form and it just redirected. Can't even tell if that worked.
Does anyone know if this can program as well as GPT-4? Because if so then the larger context window is a big improvement.
Also having a massive context length is not necessarily a good thing from perspective of cost. It also doesn't work great with a chatbot as you will have to feed the same 100k worth context back into the chatbot for every question which will turn out to be very expensive. At some point you will have to discard some parts of the context to be specific to the question being asked and that is where vector embeddings come into play. For one off research/Q&A 100k limit works great!
I wonder if Claude-100k could be used to ingest this entire thread and then answer questions based on it, or summarize or identify the pros/cons of certain aspects of Claude, large context windows, vector embeddings, etc.
If contexts keep growing, the landscape of LLM application engineering will as well
https://blog.fostergrant.co.uk/2017/08/03/word-counts-popula...
[0] https://wanderinginn.neocities.org/statistics
[1] https://www.reddit.com/r/Parahumans/comments/rz8ogt/wildbows...
Also anthropic has a "google cloud partnership" basically they are hooked on cloud credits like openAI and Azure.
I'm not sure that counts as 'great' though.
It's worse than GPT-3.5. Go see for yourself at bard.google.com, which is running on PaLM 2 everywhere but the EU as of yesterday.
> After lunch, I like to snack on a juicy and crisp apple to satisfy my sweet tooth.
> In the fall, many families enjoy going to apple orchards to pick their own apples and make homemade apple pies.
> The new MacBook Pro features a powerful M1 chip and a stunning Retina display, making it the perfect tool for creative professionals who work with Apple software.
For me:
GPT4 3/3: I couldn't resist the temptation to take a bite of the juicy, red apple. Her favorite fruit was not a pear, nor an orange, but an apple. When asked what type of tree to plant in our garden, we unanimously agreed on an apple.
GPT3.5 2/3: "After a long day of hiking, I sat under the shade of an apple tree, relishing the sweet crunch of a freshly picked apple." "As autumn approached, the air filled with the irresistible aroma of warm apple pie baking in the oven, teasing my taste buds." "The teacher asked the students to name a fruit that starts with the letter 'A,' and the eager student proudly exclaimed, 'Apple!'"
Bard 0/3: Sure, here are three sentences ending in the word "apple": I ate an apple for breakfast.The apple tree is in bloom. The apple pie was delicious. Is there anything else I can help you with?
Bard definitely seems to fumble the hardest, it's pretty funny how it brackets the response too. "Here's three sentences ending with the word apple!" nope.
Edit: Interesting enough, Bard seems to outperform GPT3.5 and at least match 4 on my pet test prompt, asking it "What’s that Dante quote that goes something like “before me there were no something, and only something something." 3.5 struggled to find it, 4 finds it relatively quickly, Bard initially told me that quote isn't in the poem but when I reiterated I couldn't remember the whole thing it found it immediately and sourced the right translation. It answered as if it were reading out of a specific translation too - "The source I used was..." Is there agent behavior under the hood of bard or is just how the model is trained to communicate?
Google has a product issue, not an AI research one.
Product race? My understanding is they've been so concerned with safety/harm that they've been slow to implement a lot of tools - then OpenAI made an attempt at it anyway.
Google has generally been ahead from a research perspective though. And honestly it's going to be really sad if they just stop releasing papers outright - hopefully the release their previous gen stuff as they go :/
They invested $300m in Anthropic in late 2022: https://www.ft.com/content/583ead66-467c-4bd5-84d0-ed5df7b5b...
(Non-paywall: https://archive.is/Y5A9B)
Another wide context model is MosaicML's MPT-7B-StoryWriter-65k+ which they are describing as having a context width of 65k, but then give a bit more detail to say they are using ALiBi - a type of positional encoding that allows longer contexts at inference time than training (i.e beyond the real context width of the model).
For these types of "extended context" models to actually reason over inputs longer than the native context width of the model, I assume that there is indeed some sort of vector DB trickery - maybe paging thru the input to generate vector DB content, then using some type of Retrieval Augmented Generation (RAG) to process that using the extended contexts ?
Maybe someone from Anthropic or MosaicML could throw us a bone and give a bit more detail of how these are working !
Who knows about Anthropic.
A bit disappointing
What is the approach to increase the sequence length here?
Now we've gone from using ML to implement slow, unreliable databases, to using ML to implement slow, unreliable string comparison, I guess
I say this, because I'm not sure how all of this is really going to scale on GPUs. It feels like LLM's are just as magical as quantum computing.
I do not have a clue what you are talking about. What happened?