Graph of Thoughts: Solving Elaborate Problems with Large Language Models
arxiv.org
arxiv.org
I was focusing more on an engineering perspective; modeling a complex LLM-and-code process as a dependency graph makes it easy to:
- add tracing to continuously measure and monitor even post-deployment
- perform reproducible experiments, a la time-rewinding debugging
- speed up iteration on prompts by caching the parts of the program you aren't working on right now
My test case was using GPT4 to implement the operators in a genetic algorithm, which tbh is a fascinating concept of its own. I drifted away after a while (curse that ADHD) but had a great time with the project in the meantime.
Instead, I think composing LLMs needs to be done in a way that degrades gracefully, with resilience to failure being a fundamental consideration. Biology has similar properties; complex biological systems (ecosystems, cells, etc) have feedback loops, redundancy, and most of all diversity. If we take a similar approach to building LLM apps, we'll end up with things like:
- multiple different prompts used in parallel, with results joined e.g. with voting. A change in how one prompt behaves can thus only have a bounded effect on the system as a whole.
- some way for an LLM to productively express 'this thing you're asking me to do is nonsense', with monitoring and continuous evaluation hooked up to that signal, and maybe runtime retry behavior as well. This can help with when you get into situations where prompt A gets an "I'm afraid I can't do that" response, and then you give that to prompt B as if it is a valid thing, and that cascades through the rest of the application.
llmtaskgraph as a library is designed to make building, operating and maintaining systems with these sorts of features easier - without good observability, it's impossible to know if some feedback loop is doing its job, or which prompts in a pool are behaving well vs poorly, much less what effect they are having on the rest of the system.
Sorry for the wall of text, I got a bit nerd-sniped. :)
Some kind of prompt like “does paper P contain idea A and does it suggest that A is true.” Then you could automatically categorise citations by whether they agree/disagree with the cited paper.
Sometimes I see papers with 2,000 citations and I wonder: how many of those are dis/agreeing with the paper.
I don’t think you need to trust the LLMs for this kind of thing to be very useful. The LLM could generate the KG with every node labelled as “autogenerated.” When you use the graph for research, you are still going to read the papers you are interested in so you can then update the relevant citation node with the label “human checked.”
If a research group uses the same graph over time, the nodes will gradually become “trustworthy” (ie verified by humans). Maybe even get reviewers to update a papers graph during review and publish that for other groups to add to their graphs.
One example of an author that is very influential, despite causing a lot of disagreement (even in more than one discipline) is Noam Chomsky, who is also the most cited person alive, and the second most cited person in recorded history after Aristotle. His views about generative grammar are in part revolutionary, in part plain wrong; your assessment of his views about the Palestine conflict and U.S. foreign politics will largely depend on your political leanings; and his contribution to formal language theory is fundamental regardless of your leanings (Chomsky hierarchy; Chomsky Normal Form).
> This has already been studied. Negative citations are vanishingly rare. So virtually all of them will be either neutral or positive. Might be a difference between science/engineering (where true) and humanities (where a larger amount is negative).
I guess most of conditions is because when biologists debate, they only have one opponent.
Out of 762,355 citations from 15,731 articles in the Journal of Immunology (1998–2007), we identified 18,304 as negative (about 2.4% of the total).
Though with LLM and sufficient context length you could probably just use that prompt directly on the academic paper without ever generating a knowledge graph
They could make it more efficient by implementing a kind of "hard attention". Each token should have access to a sparse subset of the whole input, so it would be like a node in a graph only having access to a few neighbours. Could solve the very large context issue. This can also be parallelised, running all thought nodes in parallel, of course each with a sparse view of the whole input making it much faster.
For example when reading a long book, the model would spawn nodes for each person, location or event of interest, and they would track the source text as the action develops. A mental map of the book. That would surely help a model deal with many moving pieces of information.
Or when solving a problem, the model could spawn a node to work on a subproblem, parametrised by the parent node with the right inputs. Then the node would report back with the answer and the parent continues. This would work with recursive calls.
The new cpu is the LLM and the clock tick is 1 token.
https://en.wikipedia.org/wiki/Graph_rewriting
probably not was GP meant, but something along those lines.
It seems pretty intuitive that you'd get a "task / subtask" split for example, with feedback from the latter, but semantic content largely flowing from the former to the latter.
- Complex generalization with a simple unstated justification: last 'paper' like this was ToT, and a tree is a graph with constraints.
- Framework is discussed cognitively, units of "thoughts" "scored". (AutoGPT redux, having the LLM eat its own output repetitively improves things but isn't a panacea)
- Only sorting demonstrated "due to space constraints" -- unclear what that means, it seems much more likely it was self-enforced time constraints
- Error rate is consistently 14%.
- ~10x the cost for ~15% error rate in sorting instead of ~30%
- GPT3.5
Still not sold that it'll fly in finance without that sort of observable, intermediate representation.
Number sorting is faster using code.
A) difficult to impossible for an LLM to do in a single pass B) easy to verify the correctness.
In general, the point isn't finding things that only an LLM can do, but find things that LLMs can do with decent results at lower cost than getting a human to do it.
[1] https://jbconsulting.substack.com/p/its-not-just-statistics-...