123 karma · joined January 28, 2026
Like if I take his example: "Front-line staff may be skipping mandatory fields because the process adds fifteen minutes of friction to every customer interaction." First it's unclear whats the policy for those fields are when they are mandatory and also can be skipped. Then why are those mandatory if skipping them seems only lower friction with no other consequences? How should a model decide if it should enforce the policy for those fields, code an automation or just make them voluntary?
Next i'll do is to implement tavilly, exa, tinyfish etc. as search engines for searxng. No agents, no mcp, just their search api endpoint.
Sometimes I try other models but they always feel less concise and tend to not as strictly respect the instructions. Maybe the the top frontier models from openai / anthropic would not feel under this category but they are way to expensive on openrouter compared to my coding plan.
Same with mcp. I want them to use the jsdelivr cdn instead of them scraping github against the rate limit. etc. But if I dont explicitly state to strictly use the $%!$@@! mcp for searching in repositories they simply ignore the mcp and even if clearly instructed, they still often fall back to gh.
Putting every detailed instruction in the AGENTS.md would just unnecessarily bloat the context and it works well enough to just instruct them in the AGENTS.md when to use which skill. Yet I agree that Skills are not some voodoo magic to provide your model super capabilities.
This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...
I wont advertise any commercial mcp I use but to give an example for a well designed and useful mcp server I could name the nixos mcp. Its useful because it bundles all the nix resources to one endpoint which is more efficient than web search and gives you better control over the sources.
https://github.com/utensils/mcp-nixos
Another one would be this filesystem mcp which is in my opinion to prefer over direct cli access. Of course this depends also on your general sandbox strategy but if you just use a generic docker image there are still many potentially dangerous binaries available and such an mcp can restrict the models capabilities.
https://github.com/modelcontextprotocol/servers/tree/main/sr...
And of course there are many service provider offering their mcp with its own llm / agent behind e.g. most web search provider. In this case you most likely already use an mcp without noticing it.
Also, why is he talking about "ethernet"? Its the IP layer, not the ethernet layer...
> Try inserting a few jokes, puns etc into a conversation and you'll see that it responds kn kind.
A couple days ago I was setting up new SSH keys encrypted with Ubikey but because I feared losing them and lock me out I evaluated some backup plan with 2FA. Turned out in in my homelab arent that much alternative options, a fingerprint without Linux drivers, an old Galaxy S9 on pmOS without working camera etc. After some ruling out many solutions I proposed a butthole recognition because your butt is in your pants where a face can be recorded by security cams. It recognized the joke and honestly it wasnt the worst answer. It's reply felt like your CEO makes a really bad joke and you have to answer something in order to avoid awkward silence.
I for myself use the trick to ask the model, after I explained what it has to do, what it thinks about my proposal. Of course I have no scientific evidence but to me it feels as it prevents some misunderstandings and the model follows more correctly my intention.
The fact that there are currently hardly any established methods, and that the combination of all the LLMs, harnesses, MCP servers, etc., results in extremely different experiences, and that nobody really has a comprehensive overview, only exacerbates the situation.
A tool can only be as good as the person who use it.
> Also, it seems that many don’t want to learn but instead expect to have all understanding outsourced to LLMs. Many seniors have noticed this and have stopped teaching juniors as the seniors don’t like the feeling of having their time wasted by teaching people who don’t want to learn.
Or may be it is because now a junior dev is expected to deliver to output of a senior?
> Also stop saying “please” to an LLM. It does not have any feelings.
Yes, but I still prefer a nice tone. Like why should I change my manners just because it has no feelings? If anything the statistic predicts a friendlier answer when I say "please".
> Again, I recommend people try running small LLMs locally where temperature and other settings are fully exposed and configurable to see this themselves. It is a good antidote to falling for the illusion that LLMs would actually be intelligent.
It's like recommending someone to buy the cheapest Lenovo Thinkpad to prove that Lenovo sucks.
> However, the best models still have a pass rate of only about 50% on the Humanity’s Last Exam.
Haha, as if he (or really most people) even would understand 50% of those questions. I find it rather mind blowing that it's possible to put such diverse knowledge into a couple of TB. Or may be I'm just an idiot and it's common knowledge, "how many paired tendons are supported by the sesamoid bone of hummingbirds within Apodiformes".
> LLMs are not a scam, but a useful tool and technology that has its uses. But the idea that AI has or will surpass humans any time soon in either capabilities or efficiency is simply not true
He's not wrong that LLMs are just useful tools but isn't it the purpose of a tool to surpass human capabilities and efficiency? Like even a bicycle makes traveling more efficient than just walking and a car has the capability to transport more items than any human. And its know more than two decades since computer surpassed human capabilities in chess. If a tool is neither more capable nor efficient, it's just a useless tool.
I mean I get his point that GenAI is to some degree over hyped but his arguments just dont hold in my opinion.