When ChatGPT summarises, it does nothing of the kind
ea.rna.nl
ea.rna.nl
The author states his conclusions but doesn't give the reader the information required to examine the problem.
- Whether the article to be summarized fits into the tested GPT model's context size
- The prompt
- The number of attempts
- He doesn't always state which information in the summary, specifically, is missing or wrong
For example: "I first tried to let ChatGPT one of my key posts (...). ChatGPT made a total mess of it. What it said had little to do with the original post, and where it did, it said the opposite of what the post said." He doesn't say which statements of the original article were reproduced falsely by ChatGPT.
My experience is that ChatGPT 4 is good when summarizing articles, and extremely helpful when I need to shorten my own writing. Recently I had to write a grant application with a strict size limit of 10 pages, and ChatGPT 4 helped me a lot by skillfully condensing my chapters into shorter texts. The model's understanding of the (rather niche) topic was very good. I never fed it more than about two pages of text at once. It also adopted my style of writing to a sufficient degree. A hypothetical human who'd have to help on short notice probably would have needed a whole stressful day to do comparable work.
It is basically a long winded way of saying in a bug report, “it doesn’t work”.
For example, looking at the ChatGPT link the author has, the model loaded 5 pages besides the one the author wanted. That clearly is going to cause some issues but the author didn't modify the prompt to prevent it. It was also a misspelled five (?) word prompt.
I don't see how you can draw conclusions from a model not reading your mind when you give it basically no instructions.
You need to treat models like an new hire you're delegating to and not an omniscient being that reads your intent on it's own.
And why all this talk about trying to engineer a prompt so that in the end the result is good? Should an actual usable system not just handle "Please summarise [url/PDF]"? That is, I suspect, what people expect to be able to do.
So I think that’s why you see so many reactions like this.
I’ve found chatGPT incredibly good at all sorts of things people say it is bad at, but you need patience and to really figure out the boundaries of the task and keep adding guidance to the prompt to keep it on track.
One example in the article is that if you have 35 sentences leading up to a 36th sentence conclusion, ChatGPT is very likely to shorten it to things in the earlier sentences and never actually summarize the important point.
Just the opposite: it calls into question if _we_ have thinking and understanding capabilities or if we are complicated stochastic parrots. [3] The best probing of these questions is done at the limits of comprehension and with unique and previously unseen information. I.e., how do you comprehend and process to previously unseen/unfelt/not-understood qualia? Not about how you deal with the mundanity of reactions between people (which are somewhat trivial to describe and model). [4]
[1] https://en.wikipedia.org/wiki/ELIZA [2] https://en.wikipedia.org/wiki/Anthropomorphism [3] https://www.newyorker.com/humor/sketchbook/is-my-toddler-a-s... [4] https://en.wikipedia.org/wiki/Games_People_Play_(book)
But other times I’ve persevered and once it’s ‘got’ it, it can then repeat it as many times as I need. That’s the knack really. Get it to the point of understanding and then reuse that infinitely and save yourself a lot of time.
I was impressed at those high-level summaries. If I had assigned this task to several humans, I'm not sure how many would have been able to achieve similar results.
> What the colleague used, I can ask, but I suspect standard ChatGPT based on GPT-4. But my test was with GPT-4 (current standard), so that would mean about 8000 tokens (or roughly 4000 words, I think?). That may have influenced the result.
A summary for which you must always read the un-summarized text is useless as a summary, this should be obvious to literally everyone, yet AI developers stick their heads in the sand about it because RAG lets them pretend AI is more useful than it actually is.
RAG is useless, just fucking let it go and have AI stay in it's lane.
If you repeat these steps over and over for enough iterations eventually you will run out of money and the problem will be moot anyway
If you have everything, you have anything
If you have infinite capital, then you have AGI
> The solution is to feed both the text and the summary back to ChatGPT and ask it to identify any inconsistencies.
Hmm... Sounds interesting, ...
> If you repeat these steps over and over for enough iterations
Oookay, continue...
> eventually you will run out of money and the problem will be moot anyway
Oh. :D
If you ask an intern to summarize some text, you trust them to do a half decent job. You're not going to re-read the original text. The hiring process is meant to filter out bad interns.
You can't go through this process with an AI, every time it's just a shot in the dark.
Yes.
And if I were to propose that we filter all information we consume through completel unqualified interns summarizing it, I'd be laughed out of the room.
Yet that is the future all these AI firms seek to build.
I regularly work with a wide variety of project managers, product owners, secretaries, etc…
I swear that most of them willfully misunderstand everything they’re told or sent in writing, invariably refusing to simply forward emails and instead insisting on rephrasing everything in terms they understand, also known as gibberish that only vaguely resembles English.
All of them are still “gainfully” employed.
> A summary for which you must always read the un-summarized text is useless as a summary, this should be obvious to literally everyone
Nah, it's still useful if the summary is usually right or mostly right. At the limit, it's not even clear that something can be summarized perfectly.
Consider that the alternative to reading the summary often isn't reading the entire text yourself. It's reading nothing.
Also, in my experience, these tools often fail when it comes to questions with a definitive answer. E.g. if you pass them a lot of text and ask a detailed question with a very clear answer, they often get it wrong. But when your question is vague like "summarize the text," they're very useful.
And you'd confirm that by having to read the unsummarized content. Thus useless.
Consider that the alternative to reading the summary often isn't reading the entire text yourself. It's reading nothing.
Reading an inaccurate summary is actually less useful than reading nothing. It's not like the utility of the summary is to exercise a reading muscle.
> Reading an inaccurate summary is actually less useful than reading nothing. It's not like the utility of the summary is to exercise a reading muscle.
No it's not. You're acting like these tools generately wildly inaccurate text. Even when they're wrong, they're mostly accurate. Almost everything people read is a mostly accurate summary, whether it's from some guy's article or wikipedia. We rarely go to the primary source.
edit - as a concrete example, I did my taxes this year with help from ChatGPT. It was a big improvement over using Google or reading through the instructions myself. And if it was wrong, well, maybe I'll get a bill or a check in the mail, but that was always a possibility and making a mistake on your taxes isn't illegal.
It is when it amounts to tax evasion.
If this is your approach to your taxes then you are probably wasting your time by using either of Google or ChatGPT. Just punch some numbers in to TurboTax based on a quick skim of the buttons, say “good enough for government work,” and wait for a check or a bill in the mail baby.
Just because you can amortize the cost of ChatGPT over more operations doesn't mean you aren't paying for it. You said you didn't want to pay someone for software to do your taxes, not that you wanted to pay only a little.
Asking it to do a summary on a paper or meeting notes I’m mot familiar with, yes the utility is very limited exactly the way you’ve pointed out.
However, asking it to summarise or shorten a paper or meeting notes I’ve been directly involved in - that has enormous utility for me.
It provides me a quick starting point to create summaries, and I get to see if the salient points are addressed, and if not, I get to add them in.
I’m no a fast writer, and often experience writers block, so having a fast way to start is enourmously useful to me.
But that could be actually good alternative. Sometimes its simply not worth it. You can loose more time with the tool that will give you uncertain summary.
Having no information is always preferable to having wrong information.
There are situations where AI is good enough, other cases where you need more accuracy, and others still where you should be reading the reference directly.
AI is improving quickly though, and context windows will allow for summaries to be tailored to each end user.
imho a side effect of promoting RAG is that the vector search by itself (on chunks of documents) might be a good-enough thing for most people. if we create a system without the LLM-summarization part, it might be the best of both worlds. Alas, people actually don't care about that stuff.
I would suggest a less strong but more plausible claim that GPT4o has trouble summarizing for longer form content outside the bounds of it's context window or something like a lossier attention mechanism is being used as a compromise for resource usage.
Summary:
https://chatgpt.com/share/21d81811-db45-4ac5-b3c7-b25a79b2ba...
This extends to your human readers.
One of the more useful AI tools I’ve built for myself is a little thingy that looks at a piece of writing and answers ”What point is this article making?”. If the AI gets it wrong, I know my readers will also misunderstand what I’m saying. Back to the drawing board.
The problem with making a subtle point that hinges on 1 detail in a vast sea of text is that 80% of humans will also miss that 1 detail.
I think that one of the author's key premises is false:
> To summarise, you need to understand what the paper is saying.
A summary is not about the author or about the summarizer, it's about the reader. It's about picking up on the portions of the original work that will matter to the reader's estimation of whether it's important to read the full work. And that actually depends much more on context and how the work relates to other works than it does about the specific details contained in the work. For example, Betteridge's Law of Headlines [1] basically provides a summary of any article whose title ends in a question, that summary is just "No", and it does so by making an observation about the authors rather than the content of the articles (about which it's completely agnostic).
It reminds me of the problems that plague AI sentiment analysis. Machine learning is actually very good at that task, but you top out at about 70% precision because humans top out at 70% precision on sentiment analysis. At best, people only agree with each other ~70% of the time when judging the sentiment of a piece.
[1] https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...
- I would say Betteridge's Law of Headlines does not provide "No" as a summary to all of these articles. The accurate summary would be "[Title]? No.", since it seems fairly obvious that the word "No" is very light in information conveyed.
On a personal note, I have to say I can't see the point in this current craze about summarizing everything. I never saw the point of those subscription programs which promise you'll "read" a book a week/a month because they'll send you a 5-min audio about the book. I think you're better off choosing one really awesome book a year and actually reading that.
So I can't see how a (flawed or not) ChatGPT summary will provide any epiphany on the level you'd get from consuming fewer works, thoroughly.
- The author is pushing a specific proposal (a "council of stakeholders") in the pension space.
- To support that, he wrote a long paper with much supporting information.
- The ChatGPT summary system didn't pick up on his proposal being the important thing.
- The author doesn't like that.
But evidence of what, precisely? How do we measure failure and what failure rate is sufficient?
As is, would you be comfortable with doctors applying it to all of your medical records?
No, I might not want AI summarising my medical records right now, but I might be quite happy for it to summarise a blog post for me.
Medical records are not my bar though, nor would I be comfortable giving my medical records to 99% of the human race. The bar these tools need to meet for me is far more innocuous (code and apparently summarizing fluff articles) and two things can happen at the same time:
1. Models, and not to be understated, the engineering around them can improve until they pass the point we are comfortable using them in high stakes situations like medical records.
2. People's expectations come down until we settle on the agreed upon tasks where LLMs provide real value.
My perspective is 1 will happen in the long term and 2 is what we should be focusing on right now to provide real immediate value with an eye on 1 so we reap the continued improvements.
For 2 calibrating people's expectations are going to be messy and people rarely estimate the utility of something correctly in the early days, either over or under which is the reason for the mess.
LLMs have been in the works for almost a decade (building on research that's been around since at least the 70s), but their utility has only been apparent for about two years.
We're still super early days for that settling period.
We probably need a fully rigorous study to conclude anything. Which is not what the author or myself did.
I think here's the key why some people like llm summaries while others don't.
Two people can read an article and take away different points from it. If the llm summary contains the points you would have taken, you like it. If it didn't, you don't.
(e.g. if I look at your kindle highlights for a book and compare them to mine, they'll be very different - this is why I find it hard to use a service like Blinkist - but I think a good llm like gpt-4o or claude 3.5 sonnet does as good a job as a wikipedia article would... Not sure what else people expect)
I work for the client side and this bothers me a lot. It's very hard to get a true honest value analysis done with all the sales influence and office politics going on.
That being said, these initial results are not reassuring.
1. Take the documents, chunk them up on paragraph then sentence then word boundaries using spaCy.
2. Generate embeddings for each chunk and cluster them using the silhouette score to estimate the number of clusters.
3. Take the top 3 documents closest to the centroid of each cluster, expand the context before and after so it's 9 chunks in 3 groups.
4. For each cluster ask the LLM to extract the key points as direct quotes from the document.
5. Take those quotes and match them up to the real document to make sure it didn't just make stuff up.
6. Then put all the quotes together and ask the LLM not to summarize, but to write the information presented in paragraph form.
7. Then because LLMs just can't seem to shut up about their answer make them return JSON {"summary": "", "commentary": ""} and discard the commentary.
The LLM performs much better (to human reviewers) at keyphrase extraction than TextRank so I think there's genuinely some value there and obviously nothing else can really compose english like these models but I think we perhaps expect too much out of the "raw" model.
It's also all long context models getting pages of data, which even for these flagship ones is certainly just RoPE or similar which is a cheap hack but isn't super accurate [0]. 4o is best and still showing haystack benchmark accuracies below 80% and Gemini is just completely blind. That certainly needs fixing up to 100% before we can say for sure that nothing will ever get skipped.
[0] https://preview.redd.it/rlgauej7ve4d1.jpeg?width=1086&format...
For my uses, I find AI to be a much better search/answer engine than Google ever was. It can produce answers with hyperlinks for further reading much better and more efficiently than any other option. I no longer have to read through a bunch of seemingly random Google search results, hoping that my specific question is addressed.
Nonethless, try looking more left and right. There are really good opensource solutions which can surfe the web for you like stormai, Or you can use anythingllm and give it all your local files.
You can write your tips and tricks, start commands, upgrade procedures etc. in Markdown and reference it through your local LLM.
I like it for coding specifically for languages or things i write seldomly (i'm not coding every day but did for 15 years).
Nonetheless, googles internal code review tool is already suggesting things which are getting accepted by more than 50%. Thats a lot and will only get better. GitHub with Copilot will also just get better every day too. They probably struggled (as the whole industry) with actually getting used to having ML stuff in our ecosystem. Its still relativly new.
That being said, I often come in a situation that its not the code that is bad but the solution. Where I need to tell the LLM that it's not logical to do it in a certain way leveraging on my own knowledge.
As it happens, I've just been using it this weekend to write code for me. Two things:
(1) it's not business critical code, it's a side project that I want to get done but otherwise wouldn't have energy for. Especially not in this heatwave in a century old building that has no air-con.
(2) my experiments are a weird mix of ChatGPT wildly messing things up and it managing to get basically everything done to an acceptable result (not quality of code, quality of output). Sometimes I have the same experience as you, that it's just aggravating in its non-comprehension, sometimes it's magical.
I don't know if it can be magical more often if I was "better at prompting". But I do know that I also get aggravated (less often) by other humans not understanding me, and those are much harder to roll back to a previous point in the conversation, edit the prompt, and have them try again :P
As for code quality… well, sure. Stuff I'm asking ChatGPT for is python and JavaScript, and I'm an iOS dev. I can't tell when it's doing something non-idiomatic, or using an obsolete library or archaic pattern in those languages.
Most non-tech people think AI is no different to “algorithms” and is just another IT buzzword that means “computer people click a few buttons and it does all the work for them”.
AI is basically ml at this point. And it did already A LOT.
Whisper, great jump in quality for speech to text, segment anything, AlphaFold 2, all the research paper Nvidia publishes regarding character movement, AI Raytracing, Nerfs, all the medical research regarding radio imagin, advances in fusion reactors...
We have never been so close to a basic AI/AGI / modern robots. we have instructGPT which allows for understanding 'steps' easier and more stable than anything we developed before in multip languages.
ChatGPT and LLM advances are great and helpful.
Image generation is already poping up in normal life.
AI is not wildly overhyped at this point. We are in the middle of implementation after the first LLM breakthrough and a LOT more money is funnelt into AI/ML research now as it was 10 years ago.
The future is, at least for now, really interesting and there has not been any sign of a wall we are hitting.
Even the missing GPT-5 might feel like a slight wall, but we just got GPT4-o mini which makes all of the LLM greatness a LOT cheaper and a lot easier to use.
We switched from text parsing and avg bad results to just using llama3 (with a little bit of saveguarding) and its a lot better.
Really? where ?
In my company we even have a LoRa for a specific company style of images (icons and similiar)
really curious about this. do you have an example link by anychance
Its a german news site for it people.
But i have seen ai image already on the street, unfortunate i was in public transport and not able to take a picture fast enough when i saw it and i currently work most of the time from home.
Your account is 3 days old, you haven’t read anything over and over again.
I just create a new account to get away from fomo and add a little bit more effort to commenting.
But as you can see, it doesn't work very well
The problem with LLMs is the lack of reliability and coherence. If it gives you a wrong answer and you ask it if it's sure, in most cases it will show another wrong answer and you need to go through multiple of these hops to get something fine.
I think the go to is everyone said the same thing about the internet. Look where we are now.
AI is also different in the sense that it has already gone through several hype cycles, each followed by an "AI winter" of broken dreams. Clearly large DL models are a breakthrough, but the amount of hype and hot air is entirely out of proportion to the actual results.
I think the tide is turning for sure. Metaverse scam unraveled pretty fast after building for months.
It's both, the level of criticism is also overwhelming and delusional.
I hear things like, AI will never be as smart as me!, It will never take over my job, Look it got this and this wrong, No chance it can ever be as smart as me! It's always a comparison to their own abilities so I think a lot of it is just an attempt to stay relevant in a world where technology is about to replace all of us.
I think by now everyone knows the limitations of current LLMs. The result of this "in depth" analysis is not only expected but very obvious. Nobody is surprised and none of this is suppressed because it's so obvious.
What's delusional is the amount of criticism around the performance. The denial that AI is more and more matching human intelligence. Of course it's not there yet, but LLMs made a giant leap and bridged a huge gap.
AI is no match for humanity now, that much is obvious. I think the delusion lies in the fact that many people are trying to deny the trendline... the obvious future and trajectory of what progress has been pointing too. Milestones are getting surpassed at a frightening pace. AI is now running circles around the turing test and I can now ride a car with no driver and it's normal in SF.
We are here bitching about the fact that AI shortens a text rather then summarizing it without remarking at the fact that you can actually bitch to the AI directly about this fact and demand the AI to stop shortening the text and start summarizing it.
It is true that LLMs appear to be an exciting technology, but it's also delusional to assume they're following a positive trendline. Performance between GPT-3.5 and GPT-4 was like a 25% improvement that took 2500% more resources to train. It's clear that there's diminishing returns to bigger and bigger models, and the trend we've been seeing in the industry is actually smaller models trained for longer periods in an effort to bring down inference costs while maintaining current performance.
More intelligent models may require new techniques and technologies that we don't have yet. I'm sure it will get better but the path to improving isn't as obvious as you're making it sound. Making comparisons to Moore's law is also disengenuous because we're actually running into physical limits on how dense we can make chips due to the size of atoms themselves, so past trends for technological development may not continue to bear out.
You are doing exactly as I said. Focusing on obvious criticisms on the current state of the art. We know the obvious pitfalls of LLMs. It's completely obvious nowadays.
I also never pointed to an obvious path forward. I pointed to the an obvious trend that indicates whether you like it or not, we will move forward.
>It is true that LLMs appear to be an exciting technology, but it's also delusional to assume they're following a positive trendline. Performance between GPT-3.5 and GPT-4 was like a 25% improvement that took 2500% more resources to train. It's clear that there's diminishing returns to bigger and bigger models, and the trend we've been seeing in the industry is actually smaller models trained for longer periods in an effort to bring down inference costs while maintaining current performance.
AI is following a trendline. I never said specifically LLMs are exactly on this trendline. LLMs are only a part of this trendline along with other technologies and part of the overall progress deep learning is making. It is extremely likely that there will be an AI that will solve all of current issues with LLMs in the coming decades. Whether that AI is some version of an LLM remains to be seen.
>Making comparisons to Moore's law is also disengenuous because we're actually running into physical limits on how dense we can make chips due to the size of atoms themselves, so past trends for technological development may not continue to bear out.
I never made a comparison to moores law... are you replying to me or someone else? AI with the performance of a human brain is highly realizable despite physical limits because intelligence to the level of the human brain already EXISTS. The existence of humans themselves is testament to the possibility it can be done and is not a fundamental limit.
Perhaps everywhere else. On HN mostly what I see are these fairly shallow dismissals TBH. It is a natural reaction when your own livelihood is affected, old as tech itself to be sure [0]. Still, it's getting tiresome. A new technological revolution is unfolding, and the people best positioned to lead and keep it in check are largely balking.
It is still unclear to me whether this is a deficiency in all ChatGPT models or this is just one of the many.
https://chatgpt.com/share/d5709aeb-d24c-488b-985c-c13eba0c01...
"4. IORP Directive: The IORP (Institutions for Occupational Retirement Provision) Directive is analyzed, highlighting its scope and its impact on pension funds across the EU. The paper suggests that the directive's complex regulations create inconsistencies and may need clarification or adjustment to better align with national policies." "5. Regulatory Framework and Proposals: A significant portion of the paper is devoted to discussing potential reforms to the regulatory framework governing pensions in the EU. It proposes a dual approach: a "soft law" code for non-economic pension services and a "hard law" legislative framework for economic activities. This proposal aims to clarify and streamline EU and national regulations on pensions."
^^ these corresponds to the author's self-selected two main points.
I have been working on summarizing new papers using Gemini for the same purpose. I don't ask for summary though, i ask for the story the paper is trying to tell (with different sections) and get great output. Not sharing the links here, because it would be self promotion.
It's not a problem when you are aware of it and with some follow up input you can get it mitigated, but often I see that people tend to take the first output of these systems at face value. People should be a bit more critical in that regards.
How do you benchmark something or someone understanding text?
I'm asking because the magic of LLM is the meta level which basically creates a mathematical representation of meaning and most of the time, when i write with an LLM, it feels very understanding to me.
Missing details is shitty and annoying but i have talked to humans and plenty of them do the same thing but actually worse.
I guess at best you can say these models have an ‘understanding’ of language, but their ability to waffle endlessly and eruditely about any well-known topic you can throw at it is just further evidence of this — not that it understands the content.
You can also just quiz it on certain basic definitions. Ask it for examples of objects that don’t exist (graphs or categories with certain properties, etc.). Sometimes it’ll be adamant that its stated example works, but usually it will quickly apologise and admit to being wrong only to give almost exactly the same (broken) argument again.
Another thing you can try is concocting some question that isn’t even syntactically well-formed (i.e. fails even a type check) like ‘is it true that cyclic integer lattices are uniformly bounded below in the Riemann topology?’. I imagine that one is too far out to work, but when I’ve played around I’ve found many such absurd questions ChatGPT was only too happy to answer — with utter nonsense, of course. It’s interesting (and, I think, quite telling) that such systems are seemingly almost completely unable to decline to answer a question. And the reason is that there’s no difference between hallucination and non-hallucination. Internally, it’s exactly the same process. It either knows or doesn’t know — but it doesn’t know that it knows (or doesn’t).
LLMs basically only work on questions that are very similar to, or identical to, questions that have already been widely asked and answered online or in books… hence their lack of utility in mathematical research, or even in calculating one’s taxes, or whatever.
I could provide some more literal examples, but I’d have to go and try some and pick the ones that work, and even then they might not work on your end because of the pseudorandomness and the fact that the model keeps getting updated and patched. It’s better to just play around on your own based on the ideas I’ve given.
The moral is to use LLMs as a powerful way of finding information, but don’t trust anything it says. Use it to find better sources more quickly than you’d be able to via a search engine.
If you for instance threatened to shoot a human if it refused a request or admitted it didn't know something, they might answer in a very similar fashion.
Interestingly enough, tx to a talk from a brain researcher i understood that there are two major brain modes we run: Either the i just observe and do things i observed (were it doesnt' matter that you are gay or black but still vote for trump) and the logical mind were you see a conflict in stuff like this.
LLMs are interestingly enough somewere inbetween.
Is a system prompt "provide a summary of this text" the best possible prompt? Do different models respond differently to prompts like that? At what point should you attempt more advanced tricks, like having one prompt extract key ideas and a second prompt summarize those?
Products like ChatGPT are rewarded in their fine tuning for doing happy sounding cheerleading of whatever bland unsophisticated corpo docs you throw at it. Consumer products like this simply aren't designed for novelty, although there's plenty of AIs that are. For example, AlphaFold is something that's designed to search through an information space and discover novel stuff that's good.
ChatGPT is something that's designed to ingratiate itself with emotional individuals using a flawed language that precludes rational thinking. That's the problem with the English language. It's the lowest common demonstrator. Any specialized field like programming, the natural sciences, etc. that wants to make genuine progress, has always done so historically by inventing a new language, e.g. jargon, programming languages.
The only time normal language is able to communicate divergent ideas is when the speaker has high social status. When someone who doesn't have high social status communicates something novel, we call it crazy. LLMs, being robots, have very low social status. That's why they're trained to act the way they do.
However, it still managed to pick up several clickbait headlines about NASA’s asteroid wargame and write a scare news summary:
Truth: https://www.space.com/dangerous-asteroid-international-coope...
> The participants — nearly 100 people from various U.S. federal agencies and international institutions — considered the following hypothetical scenario: Scientists just discovered a relatively large asteroid that appears to be on an Earth-impacting trajectory. There's a 72% chance it will hit our planet on July 12, 2038, along a lengthy corridor that includes major cities such as Dallas, Memphis, Madrid and Algiers.
Glancias Summary:
> NASA has identified a potential asteroid threat to Earth in 2038, revealing gaps in global preparedness despite technological advancements in asteroid trajectory redirection and the upcoming launch of the NEO Surveyor space telescope.
What prompt am I missing? Find the edge-case details and other similar "what's only mention once" I can't get it to highlight.
It's still a bit experimental. Hacker News comment sections are very large (100k+) characters so it probably won't find everything. It also adds a summary section at the end which is annoying but not too bad. And it will literally say 'these are niche points'. I find it a bit funny but there's no reason to fix it. I'm using Gemini Flash.
Here is an example for something less subjective: https://arxiv.org/abs/2309.04269
I am very skeptical of the author's claims. Perhaps the parts of the articles being summarized are not actually important so the LLMs did not include them. Or perhaps the article does an exceptionally bad job of explaining why the argument is important. Also there's a difference between the API and free web interface. I think the web version has a bunch more system prompting to be helpful which may make a summary harder to do.
Giving a LLM too much context causes the same effect, as the sliding window moves on from the earliest tokens.
It's also why summarization is bad.
It's not exactly linear though according to the text from start to finish, bit's of context get lost from throughout the input text at random and will be different each time the same input is run.
A good way to mitigate this is to break up the text and execute in smaller chunks, even with models boasting large context, results drop off significantly with large inputs so using several smaller prompts is better.
Likely it added the content to the prompt, but the content didn't stay in the prompt for the next prompt. The next prompt likely only had general web results as context.
In this case I would take a similar approach, split the document into multiple smaller (and overlapping) fragments, let a LLM summarize each one of those into key findings, and in a next step merge those key findings to a summary.
I have not a lot experience though, if this would provide better results.
[1] It's annoying to see Google initially market their Gemini models about their 100K to 1M tokens context size, and even OpenAI has been doing a lot of their model making and marketing around it too recently.
I've been having a surprisingly good time in my 'discussions' with the free online chatgpt, which has a cutoff date of 2022. What really impresses me is the results of persistence on my part when the replies are too canned, which can be astonishing.
In one discussion, I asked it to generate random sequences of 4 digits, 0000-9999, until a specific given number occurred. It would, as if pleased with its work, give the number within 10 tries. I suppose this is due to computational limitations that I don't understand. However, when with great effort, I criticized its method and results enough, I got it to double the efforts before it lazily 'found' an occurrence of the given number. It claimed it was doing what i asked. It surely wasn't. But it seemed oblivious. I'm interested to understand this.
I'm sure I'll get some contempt fo my ignorance here, but I asked to analyze pi to some unremembered placeholder until it found a Fibonacci sequence. It couldn't. Maybe one doesn't exist. As obvious as this might be to smarter primates here, I don't understand. I was mostly entertaining myself with various off the hat things.
What I did realize, is what by my standards, is fierce potential. This has me wanting to, if even possible, acquire my version with, perhaps, the possibility of real time/internet interaction.
Is this possible without advanced coding ability? Is it possible at all? What would be a starting point and some helpful pointers along the way.
Anyway, it reminded me of my youth, when I had access to a special person or two and would make them dizzy with my torrential questions. Kindof a magic pocket Randall Monroe, with spontaneous schizophrenia. Fun.
Edit note: those were but a couple examples of a lot more that I cannot remember. I'm hooked now, though, and need to come out my cave for this, and learn more. I have some obsolete python experience if that might be relevant.
Of course, it obvious in hindsight that to create a useful summary requires reasoning over the content, and given that reasoning is one of the major weaknesses of LLMs, it should have been obvious that their efforts to summarize would be surface level "shortening" rather than something deeper that grokked the full text and summarized the key points.
Many people just use GPT 3.5 because it's free, not realizing how much it sucks in comparison to newer models.
People use 3.5 for testing the capabilities of LLMs, and then conclude that LLMs are inherently bad at that task, when in reality there are better models.
Wasn't that kind of the previous classical AI method of doing summaries? Something something, rank sentences by the number of nouns that appear in other sentences to get the ones with the most information density and output the top N?
They all fail quite spectacularly at this, at least for my use case (cel shading, outlines, height sensitive fog)
I don’t need an “AI” to help me with something widely discussed, I can simply read the docs. On the other hand I’d love a tool that opens me niche topics in a reliable way.
Or rather, as I honestly don’t know what a GLSL shader is, and barely know what a canvas is, it would be a bit like the scene from blackadder - I would love for Baldrick to read this book, but that will mean teaching him to read, which will take about ten years.
I have no idea if there's an expectation of interdisciplinary research but LLMs are well suited to breaking down barriers between different levels of understanding, so it follows. I'm part of a patient research group and we use it all the time to digest academic papers, test hypotheses, and bring forward more informed questions to researchers.
``` This article examines the governance of Dutch pension funds in light of the Future of Pensions Act (Wtp). The new legislation shifts towards more complete pension contracts and emphasizes operational execution, necessitating changes in pension fund governance. The authors propose strengthening pension funds' internal organization, improving accountability to participants, and enhancing the powers of participant representation bodies. Key recommendations include establishing a recognizable governance structure with clear responsibilities, creating a College of Stakeholders (CvB) to replace existing accountability bodies, and granting the CvB more authority, including appointment and dismissal powers. The proposals aim to balance the interests of social partners, pension funds, and participants while ensuring transparency and effective oversight. The article emphasizes principles such as transparency, trust, loyalty, and prudence in shaping governance reforms. It also discusses the impact of digitalization (DORA), the need for pension funds to demonstrate value, and the potential for further consolidation in the sector. International perspectives, including insights from the World Bank, inform the proposed governance improvements. These changes are designed to help pension funds adapt to the new system, manage risks effectively, and maintain their "license to operate" in a changing landscape. ```
Similarly, the second article's summary also captures the key points that the author points out (emphasis mine).
``` The article "Regulating pensions: Why the European Union matters" explores the growing influence of EU law on pension regulation. While Member States retain primary responsibility for pension provision, the authors argue that EU law significantly impacts national pension systems through both direct and indirect means. The paper begins by examining the EU's institutional framework regarding pensions, focusing on the principles of subsidiarity and the division of powers between the EU and Member States. It emphasizes that the EU can regulate pension matters when the Internal Market's functioning is at stake, despite lacking specific regulatory competencies for pensions. The authors note that the subsidiarity principle has not proven to be an obstacle for EU action in this area. The article then delves into EU substantive law and its impact on pensions, concentrating on the concept of Services of General Economic Interest (SGEI) and its role in classifying pension fund activities as economic or non-economic. The authors discuss the case law of the Court of Justice of the European Union (CJEU), highlighting its importance in determining when pension schemes fall within the scope of EU competition law. They emphasize that the CJEU's approach is based on the degree of solidarity in the scheme and the extent of state control. ** The paper examines the IORP Directive, outlining its current scope and limitations. The authors argue that the directive is unclear and leads to distortions in the internal market, particularly regarding the treatment of pay-as-you-go schemes and book reserves. They propose a new regulatory framework that distinguishes between economic and non-economic pension activities. For non-economic activities, the authors suggest a soft law approach using a non-binding code or communication from the European Commission. This would outline the basic features of pension schemes based on solidarity and the conditions for exemption from EU competition rules. For economic activities, they propose a hard law approach following the Lamfalussy technique, which would provide detailed regulations similar to the Solvency II regime but tailored to the specifics of IORPs (Institutions for Occupational Retirement Provision). ** The authors conclude that it's impossible to categorically state whether pensions are a national or EU competence, as decisions must be made on a case-by-case basis. They emphasize the importance of considering EU law when drafting national pension legislation and highlight the need for clarity in the division of powers between the EU and Member States regarding pensions. Overall, the paper underscores the complex interplay between EU law and national pension systems, calling for a more nuanced understanding of the EU's role in pension regulation and a clearer regulatory framework that respects both EU and national competencies. ```
I'd bet that the author used GPT 3.5-turbo (aka the free version of ChatGPT) and did not give any particular prompting help. To create these, I asked Claude to create a prompt for summarization with chain of thought revision, used that prompt, and returned the result. Better models with a little bit more inference time compute go a long way.
Don't get me started with Sonnet 3.5
AI winter is going to be brutal
ChatGPT Isn't 'Hallucinating'–It's Bullshitting