Btw it’s not just those two, add iGaming into the mix which is cell phone gambling glammed up in Candy Crush form. Leeches, all of them.
565 karma · joined January 29, 2018
Btw it’s not just those two, add iGaming into the mix which is cell phone gambling glammed up in Candy Crush form. Leeches, all of them.
- memory systems are a specific type of knowledge base where you generate all the documents. You might as well generate them to be less than your embedding token limit to obviate the need for chunking.
- embedding models are getting better and are no longer just semantic averaging.
- small models are getting dirt cheap, making parallel reads cost manageable
What they describe is sort of the simplest architecture that takes advantage of these observations. I believe them when they say it works well.
I do suspect though that things like keyword lookup will completely fail if every memory is just a vector. Hence why something like Typesense hybrid search can still be useful.
Plausible scenario. Individuals predisposed to Alzheimer’s experience different mental sharpness from birth and this makes them enjoy taxi driving less (more intense than long-haul trucking), and so they pursue taxi driving as a career at lower rates. Under this scenario, driving has no effect, it just induces a selection bias.
More practically, I think you could give a board of directors access to an LLM that sees all levels of company operations and let them talk to that chatbot as a stand-in for the CEO on a 24/7 basis.
To be clear, I’m not talking about subjective style issues, I mean conforming to their own spec and avoiding careless bugs.
All remaining work fell on the backs of the physics referees. I’m not sure what value Springer provided from an editorial standpoint. It was disappointing to say the least after all that hard work.
That said, I’m kind of having a blast using CC in corporate with all the connectors available at our disposal, and I baffled how little some of my coworkers know about what’s available and what the capabilities are. So it’s clear that perhaps some encouragement is prudent for those who are slower to embrace new technologies, but I’m not sure tokencounting and tokenmaxing are the answer.
It’s also a good ledger of exactly how much I socialize and with whom. Often I find that I think I hangout with certain friends more than I do, and this helps me confront the reality that I need to put in more effort.
To be clear the catchup frequency is not meant to promote “hangout every X days” but rather to encourage me to be more mindful with who I hangout with and how often (roughly).
I think the solution will be small (1-5 person) teams where product and engineering sit next to each other and have clear authorization to launch directly to prod at their discretion. The gripes about performative work tracking mechanisms and the realization that non-tech considerations are now the bottle neck are not mutually incompatible.
We originally had RAG as a form of search to discover potentially relevant information for the context. Then with MCP we moved away from that and instead dumped all the tool descriptions into the context and let the LLM decide, and it turned out this was way better and more accurate.
Now it seems like the basic MCP approach leads to the LLM context running out of memory due to being flooded with too many tool descriptions. And so now we are back to calling search (not RAG but something else) to determine what’s potentially relevant.
Seems like we traded scalability for accuracy, then accuracy for scalability… but I guess maybe we’ve come out on top because whatever they are using for tool search is better than RAG?