If you've never been totally recommend. K Road in general is always so amazing (well, like >7 years ago).
1,294 karma · joined June 29, 2021
If you've never been totally recommend. K Road in general is always so amazing (well, like >7 years ago).
I had a stopover in Bangkok and _everyone_ was wearing masks. Not just Asian people, _everyone_. And that's when I was like "oh...this covid thing in the news is gonna be a bit bigger than we thought".
Then I land in the UK after that and all my friends & colleagues either haven't heard of covid or aren't really bothered by it ("it won't come here"). Was back at work for 2 weeks, then the first lockdown hit.
And we had lockdown, after lockdown, after lockdown. I would walk to the supermarket (the one thing we were allowed to do as per the rules) and be walking past parks filled with moron families socialising in the middle of lockdown.
The UK had so many long lockdowns because people here are quite selfish and cannot follow rules - and before people take offence - this is in contrast with NZ culture. There's much more "me me me" and "mine mine mine" culture here in the UK.
I'm not the type, myself. But damn, no wonder downtrodden men end up snapping - because they most often have nobody to turn to and are forced to exist within a bubble that analyses their weakness and victimhood as much as women are for beauty standards, etc.
If it's neuro (llm/transformer similar) symbolic (symbolic with hand crafted rules) then, imo, that's no different from current tool calling harness implementations and I'd love for someone to explain it further to me if I've misunderstood.
If it's neuro (llm/transformer similar) symbolic (symbolic rules that have also been learned via training) then yeah I can understand how it's a distinct concept.
But every definition I've seen tells me that a standard harness of:
query->llm->[tool call symbol + tool name + tool params] aka symbols governed by logic, including param/arg validation->llm->etc
Already meets the requirements for being "neurosymbolic"...
I've found with traditional embeddings that obviously you're getting an average of the content of the chunk even with the semantic awareness magic. And our (or I guess the) core problem seems to be a lack of enough contrast between chunks with makes one-shot pure embedding based RAG extremely difficult and low quality.
I am wondering if models will end up being some sort of evolution of MoE where it has something internally like the model the author refers to that gets surfaced when it needs to search in some way. I guess it makes sense; our own brains have so many distinct task-specific regions.
I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?
It seems like a lot of the problems I have been running into with RAG on large/complex documents with generally low contrast in the information is not one that has been perfectly solved yet - here I am thinking I'd been a bit behind.
It's just unfortunate that none of the cloud providers are flexible enough to deal with the pace of change. Probably going to have to shove one of those 8b~ models into an instance to use when needed.
It's interesting you mention late interaction (retrieval), I had recently been using ChatGPT as a mirror to throw ideas back at me on this issue and had been musing about how nice it would be to have some sort of hierarchical embeddings that capture a whole chunk, then sentences and then sentence fragments or individual word and it seems that that fits the bill!
Basic embeddings give decent-ish results (in the top say 20 chunks). Basic agentic retrieval gives slightly better results so long as the agent part of it doesn't go down the wrong track.
I like the idea of what's discussed in the link, however atm we are on Bedrock KBs and so locked in to a very basic implementation of RAG, because Amazon doesn't have the foresight to make things flexible enough - including making it an absolute pita to use their hybrid search. But, I guess they "work" reliably.
One of our core issues centers around a 1300 page document all about the same overall topic but with minor various for specific procedures/situations. Typical embeddings waters this down so that each chunk really just represents the common theme and therefore lacks a lot of contrast.
But now that luna's (and others) price has been cut, perhaps I'll start experimenting with giving it free rein to explore the data a little in the same way that I do a web search.
One thing that definitely helped was providing a separate index of each section where I had another model summarise the primary unique topics in each section to act as a guide for the agent. I think either we should be chucking the entire doc at a model (400k tokens...so not really ideal at this time) or improving RAG accuracy. For the latter I think even with embeddings, meaning of words and semantic connections are not enough at all - attention is KV so it is 2 dimensional and once I started getting into it I've kind of realised that 2 dimensions aren't really enough to represent the logic that exists between tokens (i.e. sections of documents that refer to a sequence of actions dependent on some logic that references "variables" from another section, i.e. "if x, y has happened then refer to z sequence). There's much deeper meaning to human language than I think basic embeddings covers.
I think it's becoming clear to me that in the same way that embeddings encode the web of semantic meaning of a chunk of text, I need something similar to a hybrid of the author's model + reranker + super-embeddings that encodes as much of the entire meaning of a text as possible and not just semantic.
But when it actually happens to you, as I have been violently mugged several times. Oh, would I push that button - pull that trigger and I'd rewind time so I can do it a few times over. Cathartic.
These little fuckers hide out in alleyways late at night and they've made a career of it. They do it so often they embellish their muggings with other bits like choking and mock stabbing or slamming people's heads against pavements. So fuck 'em.
But I would rather it do that than my phone still be unlocked after being mugged or scooter thieves'd'd.
I've also ensured I switch to esim so that they can't just eject it. Unfortunately they will still dump it into faraday bags.
Who goes when companies need to downsize? Rarely executives; they'll always find a way to be retained...hell, they're the ones with the power to decide who goes.
However thinking about it...if someone asked me to manually create an SVG, or hell even draw a quick doodle on a bit of paper of the same, I'd still probably be much slower than an LLM and potentially end up sketching less accurate anatomy than the machine.
I think the general "organic task" stuff has been mostly sorted out, but in personal and professional experiences using AI to try to _do_ something, I've found less so recently problems with hallucinations and moreso problems with attention.
For example GPT5.6 still has issues where if I provide it with a list of documents and then ask it to raise questions from that information. Then provide it with additional documents that answer some of those questions and ask it to summarise which outstanding questions there are again, it still asks questions that have become irrelevant with the additional documents - but when this is pointed out it knows exactly what to do and produces the correct list of outstanding questions.
I'm sure frontier models are doing all sorts of crazy stuff with attention already, but it seems to me like we almost need some hierarchical attention mechanism like KVL (with Level added) so that it's aware not only of semantic connections between tokens in the context but also of where there are gaps, missing links to assist the model in becoming aware of its own attention span (I guess).
Let alone a decent sized model on a couple 4090s or similar, pretty gud t/s to churn out tts/stt/response/control model actions and emotes.
Interesting how Genghis Khan got away with it, to most he's now just a "badass" historical figure, I don't think most people could tell of all the terrible things he was responsible for.
Just seems like the ideal way to spend the last period of your life; quietly making the small mechanical pieces and hopefully finally assembling something to be left behind.
Though I don't imagine I'd ever be able to produce something as small, accurate or intricate as these students are able to.
Many modern watch parts are CNC machined, often the finishing is done by hand such as zaratsu polishing - but even that is a repetitive motion that can be mechanised.
I would not be surprised if given enough time even what we have today - a decent VLA model + some other specialised models, 6-axis CNC machine, an SMD pick and place etc would be capable of designing, manufacturing and assembling a mechanical watch.
On a Samsung S24U I held down the "circle to search" homescreen button which brings up the AI tools interface (I don't know what it's officially called), held down on the text and copied the whole thing in one shot.
It took like 2 seconds.
Humans aren't like that lmao. We're reactionary, tribal animals.
I can totally get the "don't tell people to kill themselves" aspect of these models, but certain parts of the Internet have been telling people that for decades.
Certainly the models should be trained/tuned to avoid conversations like that wherever possible and redirect people to get the help they need...but that's exactly the problem; doing that is MORE than what the state and the strangers surrounding a person would do. That's the problem, a mental health crisis that is ignored, particularly in men.
I wouldn't think that the copy of some movie Netflix is streaming to me will be 60-100GB over the duration of the movie. Not to mention when their services have issues and you're watching 5-10 minutes of low quality content until it settles and snaps up to full (streaming) quality.