> when I can have it do something like write entire working kernel module fixes for old MacBooks on a whim
They’re not saying it can’t do that, and that’s not proof it doesn’t hallucinate. In fact, having used 6-8 agents at a time for a year plus while writing AI tooling for an AI startup, I can definitely surely tell you that they’re almost inversely correlated as in models that hallucinate a lot sometimes also put out the best most impressive solutions.
I’m definitely not anti AI and I definitely have found a way to make it work very well and I’m content with the work I get out of it (again maxing out several max 20x subs), but I have had sol definitely hallucinate this week and I’m a bit shocked you’re trying to say otherwise.
Listen I know it’s going to be I’m holding it wrong too, but I’ve been reading white papers and research on LLMs for a long time and was definitely at the cutting edge of context engineering, implementing features in our tooling harness a year before they were in codex or Claude.
maybe I am holding it wrong still but but like at some point if I’m holding it wrong who else will be holding it right? Dozens of people? At some point, the technology has to be approachable enough for everyone to have your point of view automatically.
I'm no expert on the inner workings/harnesses/etc beyond a basic understanding of the architecture. Maybe I've just developed a good sense for effective prompts? I could share some recent sessions.