You said no MCP
earendil.com
earendil.com
So even with a local Qwen and Pi you can now say things like:
Set up Clop to optimise any PNG that I drop in my website assets folder and convert to a webp with the same name near it
Get Crank to start Time Machine backups immediately when I connect my HDD and notify me when the backup is done.
I want to be able to hold rcmd and fuzzy search and focus cmux agent panes
BetterTouchTool has a great MCP which can create native SwiftUI views and bind them to hotkeys, trackpad gestures etc. It can leverage its immense macOS automation tools and private APIs to let agents do Computer Use.You would need a much more capable coding model to code those tools from scratch and get the same fail-safe logic that the apps have honed over the years.
Like, since MCP, Crank [1] has fully replaced my use of crontab, launchd, scattered scripts I run once a week. Not that it could not do that before, but it's so much simpler now to just describe the automation and have it happen reliably and visible in the UI. The friction is gone.
[0] https://reddit.com/r/macapps/comments/1wkv0dy/mcp_in_macos_a...
They're so powerful and yet get out of your way when you're not using them. I couldn't imagine being without them. Thanks y'all!
Everyone on this forum has an absolute paucity of imagination when it comes to applying LLMs to any use case that doesnt involve coding.
Of course MCP has its use case like if you want auth, or session based actions.
NO IT CANT!!! why dont you understand that not all agents have access to a terminal!
Recent example: The AWS CLI has a convenient "s3 sync" command that does a one-way directory sync, only downloadiong new files if existing files with matching sizes and timestamps don't exist, but the closest thing the .NET SDK has unconditionally overwrites destination files.
Another example: one of the nicest things about C# is how the compiler exposes itself as extensive library functionality, so, not only can I compile and run code at runtime, I can generate this code from programmatically created ASTs instead of text, parse expressions into ASTs, etc.
These are things devs often take for granted as being part of ‘how chat agents operate’ but they are specific to how claude code/opencode/pi/codex operate.
Giving agents a Unix computer account they can play with is definitely a powerful tool that makes them capable of doing a lot more (see: meta muse, OpenAI dots), in much the same way that giving a human a computer they are trained to use makes them way more capable… but it’s surely not the only way we can run these things.
They pay tons of money for hardware, only to use it the same way I was using those DG/UX terminals at the university.
Naturally there are no coders in other operating systems as well.
But in my case, a CLI was not enough.
Like, to the MCP I might say:
Set Clop to make every video copied in ~/shots smaller, 2x and silent
Then the MCP can use elicitation and say: By smaller, you mean re-encode to compress file size (factor can be 0 to 100 max compression) or downscale resolution (100% same size, 50% half size)?
And does 2x mean faster speed? In which case do you want to keep frames so the video plays smoother or drop frames for size? Or does 2x mean upscale?
And the agent will present those as nice choice menus I can decide schematically on.With the CLI I have to first read, learn and memorize the requests and commands needed for each app, the accepted values and formats and the steps to reach a specific result.
There's only so much space in my head I can leave for implementation details of arbitrary apps. I'd rather have an agent care about that.
And yes I get the irony, those are my apps, I coded them by hand for years, I should know their implementation details, yet even I forget if I should pass 50% or 0.5 for half size.
Btw Clop is a media file compressor for context: https://lowtechguys.com/clop
Plus I can add some complex commands like in the rcmd Stages [1] case where the agent can create a 4 monitor layout with apps and windows placed where you want, with every window opening the document/folder/project/URL you want and running the terminal commands you need. Sure you can do that with the CLI, but it's hard enough to get right because of shell quoting issues, that even an agent can get it wrong.
For simple tasks though, sure, the CLI is just enough and the agent can use it without needing to install yet another MCP. You'll know when you need it.
[1] https://lowtechguys.com/rcmd
---
EDIT: I just remembered, you can even hook the Claude/Codex/Gemini desktop app to the MCP, while you can't get it to use the CLI. so there's that for users that still don't feel comfortable at a terminal, which is a number higher than you might estimate.
I'm working on a sideproject called Rowbly[1]. It acts as a sharable data store where LLM's can dump research rather than keeping it in their memory or throwing it into a spreadsheet.
At first I thought "I don't need an MCP, I'll just expose a CLI" but that carries a pretty big limitation in that it only works with agents on your computer (Codex, Claude Code, Pi, etc). For the folks on here this is not an issue and is often times preferable but I'm also targeting the average LLM user that primarily interfaces with it via "consumer AI" and with those tools the stuff you can do is very very limited.
I still think MCP's have a long way to go maturity wise and hey maybe in a few years we will figure out a better way to do things but for now, if you want to interact with consumer AI apps, there's just no way around using them.
[1]: https://rowbly.com
There were some teething problems on Golden Gate: keystrokes didn't seem to make it to the rcmd popup, but your recent remediations seem to have solved it. It's a beta OS version too, of course :-)
Still working on finding all the edge cases, so sorry if you still encounter problems there. I've been using it since June and still find problems in input handling.
Like there's this thing, where if an app has Accessibility Permissions and listens to key events (like rcmd does) and then you revoke that permission while the app is running, then your whole system will stop responding to keys and clicks. Until you kill the app in question, but how are you going to do that without a keyboard?
All apps have this problem, even established ones like BTT, because it's a recently introduced macOS behavior in how the internals of CGEventTap work.
Why is AI involved for any other reason than building the original test implementation?
Not sure if you got the right context, your question doesn't really make sense to me.
We're back onto the original use cases for natural language processing. This is where all the value always was, and now the market has proven to itself what anyone with even a bachelor's in computer science already knew.
This seems like exactly the sort of thing I've done with shell scripts or even makefiles.
But this is for people that already use the app, researchers, writers, students, people that aren't necessarily comfortable with a terminal. And given Clop already implements the basics: an efficient file events watcher, tuned encoders for the Mac silicon, fail safe backups and UI for seeing the result and interacting with it in real time, it has advantages over trying to do it yourself.
I will point out that the shell script way uses less resources than an LLM making a tool call. But I understand that these scenarios are not necessarily meant for the same user.
Oh for sure, I would prefer to have the automations as invisible things running at the system level, doing exactly what I want and nothing else, not wasting resources on UIs and event watchers I might not need. I would get rid of my own apps if that was easy to do.
But it seems we need to waste some resources to get some usability in return.
Definitely saving it for further use, I sometimes need to have small invisible watchers and I don't want a full fledged app or shell scripts for that.
Yesterday I received a new thermometer for my aquarium to replace an old broken one. Both were bluetooth, but different models. I just told claude "I'm going to set up up my new bluetooth thermometer for my fish tank in a few minutes, keep an eye out for it and replace the old broken one with it in Home Assistant" and then walked away and put a battery in it and put it in my aquarium.
When I came back it had found it, replaced all my existing entities for the broken one with the new one, and verified it was all working with my existing graphs and automations.
Having an agent keep an eye on stuff and fix things proactively should make the experience much better. Plus I can no longer write yamls at last.
User will interact or build apps with simply text like: "Give me all the issues that are X, context: https://somedomain.com/llm.txt"
llm.txt will have all the API instructions
Not only giving an API key to an agent can leak to the model because of harness issues or too broad reading rights, but also most providers don't give the ability to apply principle of least privilege to an API key.
I don't want to give an agent full R/W access to any of my services/accounts.
The agent can get to the resource through the MCP server or using API key. I personally do not see the benefit MCP is providing here. Sure you can reduce the exposed surface at MCP layer, but I do that at the API layer. I don't need to add another layer here.
I can kinda understand if you do not have control of the API layer and/or you have to expose the API layer to the public Internet as well. Most of the time that is not the case for me.
Definitely helps that the home assistant api is documented online most likely in the training data.
It is more of a curse than a blessing. MCP pollutes agent context even when you are not using it. Use a manually invoked skill instead if you don't want to be wasting tokens on every turn and bloating up agent context making it dumber in the process.
MCP in some harnesses bloats context. MCP in some harnesses doesn't bloat context.
Is that not how it works out of the box?
It's a very specific thing for me really, I connect the HDD specifically for doing backups as fast as possible then I want to disconnect and store it back so I can keep using my laptop. I don't have a desk anymore where I can keep these things connected all the time.
The linked post from Armin is gold:
"... whenever you are confronted with a very strong opinion about a topic, reasonable discussions about the topic often involve arguments that have long become outdated or are no longer strictly relevant to the conversation."
(https://lucumr.pocoo.org/2016/11/5/be-careful-about-what-you...)
And that's from 2016! These days if you're arguing from a position you took even a week ago, you already might be out of sync.
Thats my observation when people retort “whataboutism” in usually a geopolitical context
Its not ever clear to me that people are aware that their own country or a place they respect engages in the same practice as the place they are denigrating
Like, sure I would like both places to fix their problem but dont you have better things to do like fix your own?
A direct quote from March, 2026[1]:
> If you’re still not convinced that a lot of this discourse [regarding the death of MCP] lacks nuance and is just hype, congrats on buying into the current AI-influencer FOMO hype cycle; see you in 6 months when the influencers move on to the next revelation of the moment to stay relevant and get your eyeballs and dollars.
It was fairly obvious why MCP would be needed once AI engineering and uptake moved beyond the solo developer and single harness stack of "what works for Me" versus "what works for My Team", particularly in an enterprise context. The key mistake people made was thinking in terms of their own workflows and own local stacks instead of a team's workflow and a team's operational stack. There was also an ignorance of MCP's stateless HTTP mode (yes, it was already a thing in March; the 2026-07-28 revision of the spec just prioritizes it as the primary focus moving forward) versus local `stdio`.My biggest complaint right now is that OpenAI has still refused to implement the MCP Prompts spec[2] and in general, the major clients have spotty implementation for some of the features in the spec.
[0] https://news.ycombinator.com/item?id=47380270
[1] https://chrlschn.dev/blog/2026/03/mcp-is-dead-long-live-mcp/
It is only going to continue to proliferate in usage and adoption.
Do you mean the Roy Fielding "REST" or the HTTP API "REST"?
Calling the latter "REST" is wrong. It's like calling a hermit crab a snail.
Edit: I forgot to mention, the former is gaining relevance as the elusive evolvable clients are now a thing.
If you can build the ultimate evolvable client, then you can collapse all UI into a single client.
Basically Roy Fielding told us to build an API that can only be operated by a human like intelligence, nobody could implement that because such an intelligence did not exist (hence the switch to imperfect HTTP APIs), and now that LLMs are a thing, said human like intelligence exists. This means Roy Fielding wasn't wrong, he was 25 years too early and ironically people should be building real REST APIs in the pendantic academic sense from today on and not the "pragmatic" HTTP API.
MCP servers provide three things: tools, resources, and prompts. Of these, tools seem to be the only part implemented consistently across major clients like ChatGPT, Claude.ai, Claude Code, etc.
For prompts and resources, there doesn't seem to be a common understanding of how clients are supposed to consume them.
For example, if an MCP server exposes resources, Claude Code can discover them and consume them when needed without you explicitly asking for a specific resource. Claude.ai behaves differently. It doesn't automatically discover and consume those resources. Instead, it gives you a way to manually add an MCP resource to the prompt.
So while MCP defines tools, resources, and prompts at the protocol level, the actual user experience for resources and prompts varies quite a bit across clients.
Codex is the only mainstream harness that does not implement this in the client.
A good use case for resources is small amounts of commonly needed state that can be fetched and proactively updated by the mcp server, saving latency when the model requests it.
Dynamically target sets of `/` commands to teams in an enterprise by their identity+claims? Legal team gets a set of skills? Finance team gets another just by their roles? Always up-to-date delivery of what are effectively remotely served skills? Telemetry on who is using which skill? Server-side rendering of skills so that common skills can be composed? With placeholders replaced by user- or team-custom options? Easy to ship new skills as long as the user has connected the MCP? Easy to retire skillsets that are outdated across the entire enterprise?
MCP Prompts is one of the most powerful capabilities in the spec for enterprises.
OpenAI team: if you are angling for enterprise, you need to get this solved. Your FDEs are going to make a killing getting this set up for enterprises. Build an enterprise skills management platform around this that's integrated to their directory. Streamlined setup of the MCP via MDM. Telemetry on enterprise wide usage of curated skills across the enterprise, by team, by individual. You need this.
I follow several influencers from pre-AI times who were (or were trying to become) social influencers thought leader types. People like Theo or many of the JavaScript and training course people. Following them was helpful to follow the trends that more chronically online juniors would be picking up and pushing at the workplace, which set me up in a better position to understand and then defend against it.
All of them, every single one, have dropped their previous influencer topic and pivoted to being AI influencer. Every time I see a post they’re either saying you need to adopt a new trend or that last month’s trend is dead. “Prompting is dead! Graphs are the future!” or “Claude Code is OVER! This new harness is 10X better”
MCP was one of these topics. Everyone went from telling you that you needed to use MCP or be left behind, to declaring that MCP was dead almost in unison.
The only pro-MCP holdouts were the influencers who had built their own MCP courses for sale.
The game has changed in the AI era. Horde secrets, skills, custom processes. Don’t share everything. Stop giving things away or you will have nothing left.
“When a favor becomes too large to repay, the gratitude turns to resentment.”
Good luck
I can see the snap appeal. One protocol, we can chuck an auth reverse proxy in front of all the MCPs, compliance has their integration point, etc.
I don’t think MCP is structured enough to give a huge edge over bash there. Looking at MCP messages, they aren’t immediately more legible than a bash command, output and exit code. You also don’t own a lot of the MCP servers you use, so backtracking for audits will require knowing what MCP commands did what back then.
I do suspect something more like MCP than bash will be the winner. MCP just feels very open source rather than enterprise. Eg I don’t think I’ve seen any sort of privilege escalation and logging scheme. The enterprise will want some sort of “request admin privileges” scheme. Likewise they’ll probably want more context on ACP requests; who is calling this MCP, using what agent, and for what project?
MCP over HTTP buys you server side telemetry, composability, remotely held credentials (security), etc.
MCP is about enterprise control of the server side; no advantages on the client side at all.
I do agree, the "MCPs are dead" narrative was overblown, but there were legit reasons we weren't ready to go all in on them back then.
We need a dunning-kruger for empathy: people least able to imagine situations not their own trying hardest to influence others.
It's hard to customize those skills. How can I tweak my skill a bit to match my workflow? With MCP Prompts over HTTP, this is easy: you can server render the text with my personalization specific to me.
It's hard to tell which skills are being used. With MCP over HTTP, each Prompt call, each Tool call is an HTTP request and you get telemetry on activation. You can server compose the response and ask the agent requesting it to return a score on how useful it is, too. Or ask it to call another endpoint to rate the skill.
Enterprise skills delivered to the disk you cannot do this. If a skill is flawed or outdated, you cannot revoke the skill at an enterprise level. MCP Prompts: it's easy to do.
I don't see what you mean?
An MCP and a skill are just two ways of distributing capabilities.
For telemetry ... well the agent should still have it's http logs? Also, I'm not sure MCP Telemetry is the primary point of concern for most things? Like thats a skill debugging thing?
lmao not sure if you are being sarcastic but that guy is joker and a poster child of ai psychosis .
It’s suboptimal for the reasons the author outlines: but so is USB-C. So is NVME, so is HDMI.
We use these hugely successful technologies in spite of their flaws because they’re widely compatible and easy for the end user.
That’s why MCP is everywhere. It might not be performant, robust and uniform but it WILL get better over time.
And I’d much rather have the broad MCP ecosystem that we have now than seven or eight different “optimal” ways of plugging in an LLM to something useful.
I hope they'll do the same and eventually add native support for ACP (https://agentclientprotocol.com/get-started/introduction) which, on the contrary, I use quite.
Then again, I don’t even know if general adoption is what Pi/Earendil is going for.
About what Pi/Earendil is going for, I can't really say but a while ago they created quite a stir in the Pi community for adding a trust system[0] which for many (me included) went against the loudly advertized "yolo" phylosophy, at that time I speculated it was a move to make it more palatable for the general population (whatever that actually means), so I'll stay optimist for ACP adoption for now.
[0]: https://pi.dev/docs/latest/security#understand-project-trust
Pi’s agent is supposed to be simple, and a simple ACP agent is like a couple hundred lines of code. Making a system that allows UI plugins is way harder.
Also not sure if you’ve seen but you can get ACP from Pi with https://github.com/svkozak/pi-acp It bridges Pi’s RPC mode to ACP, works okay but not amazingly. My thinking level selector in Zed has never worked with it but everything else I use has worked (I’m sure other things don’t but I must not use them).
I'm not so sure about this move, or the general inclusion of code mode in the core editor as one of pi's main selling points was its minimal nature.
Though as time progresses, they are probably going to do the same with sub-agents.
There are a lot of ways to implement sub-agents, and it's not something I need the harness to be opinionated about.
I know builtin tools support opt-out, but it's more bloat. It's also more complexity for the agent to understand when you use it to build extensions for itself.
Imagine a world where you could:
* Configure your favorite harness/chat client with any OpenAPI spec for an API that supports Oauth2.
* The harness would walk you through the Oauth flow and securely store the token.
* And then insert some tools for discovering the API methods and making requests in to the context.
* The agent could then formulate a request, call the request tool, and the harness would 1) makes sure it's allowed to make a request to that API, and 2) insert the Auth Token into the request.
It would be basically exactly the same way MCP is setup today, except all you would need is an OpenAPI spec. You wouldn't have setup a server for a janky new standard that's half implemented slightly differently by every harness/chat client.
I suspect that a smart model driving multiple dumber models for work and then using sub-agents with the same smart model for adversarial review will be a pretty common pattern.
Personally, I got a bit confused about Pi having most of that stuff as plugins since I remember how much of a mess Eclipse was where so much was just loosely fitting together plugins and just went with OpenCode since it covers most of my needs out of the box. Guess that might also be a sign of me getting older, because my IDEs and desktop environments are all closer to stock too.
Like sub-agents, you could just instruct pi/any harness with a user prompt/system prompt to start new invocations of itself, if you share what the exact command is, and pi or any other harness will do their own poor man's version of sub-agent via standard unix programs.
Using sub-agents for example also lets me decrease the default context size in Claude Code instead of running at the full 1M like:
/autocompact 420k
or deal with Codex's 258k tokens (seriously quite tiny by modern standards).Same idea with something like OpenCode, there I even configured custom agents for review: https://opencode.ai/docs/agents/
For that is it not better to have separate sessions for planning stuff and doing actual work? Pi is super flexible with session management, and a lot of that can be automated by its extension system.
Personally, seems like too much effort for something that would still need to (and fail to) have some sort of a link between the two, so I could go from the planning over to implementation and back easily. In reality, that'd get lost in the noise of dozens of sessions - I mostly just want the harness to help me do work and otherwise get out of my way, not make me dance around it. Ergo, the more context management it handles, the better!
It's part of why I really dislike to work in Claude Code, and found it too unwieldy. There I have to keep dancing around it to manage the context in a sensible way.
I mean, not in Pi at least. In Claude Code it is a chore.
But they'd still inevitably get to long in the tooth, and context poisoning meant they'd just eventually not be able to stay in the preferred context size, which for me is 64k-128k. So, I extended it with an eviction command and required a ratio. So instead of a summary of work, it now just places a waypoint. The waypoint basically means the context has a semi-coherent context but without all the baggage.
I'm on like day 3 of a single session with 3m tokens removed and still in the sweet spot. So it evicts to beneath the lower limit, compresses to the upper limit, then evicts again.
It's amazing how resilient it is if you give it a good plan. The work flow has basically been:
1. Write up an implementation document for some new set of features.
2. Rewrite the implementation as a TDD document
3. Set it to work.
The only thing I haven't figured out is it likes to stop when it hits the finish line of the subparts, but likely we're going to end up with the master of puppets monitoring these things and just set them to evaluating what they've done.
Doesn't that destroy the cache? I find that caching significantly sped up my Qwen, especially on said larger contexts.
So it is designed like a heap, where we're taking raw context off the heap, compressing it, and putting it back on the heap. So cache during compression is mostly unperturbed, since we're rarely digging all the way to the bottom of the stack, but that could happen.
Eviction though is cache busting; but again, I'm valuing the session's roadmap as the valuable product and context size slows computation size, so I have to bust the cache to sacrifice immediate re-processing for longer term compute speed up.
Because that's faster than getting to the end of the context (remember, every 1k adds to the compute time of the next 1k). So speed at 200k is much lower than at 100k. It's also local, so I'm only paying time+watts for the trade off. As far as I can tell, speed is not being lost since if I let the context grow, the kv cache doesn't help with the compute throughput.
So, yes, but it's "smart"; we're only busting it at the top of the context, so rebuilding it isn't from the bottom up, it's just at the top. Those summaries sink on the heap until you get to the eviction limit, and then, they're evicted, and we rebuild from some intermediate place in the heap.
The benefit of it all is I can have lots of projects, and keep a single session that tends to have the context necessary to avoid having to write AGENTS.md or other context bloats. Set large implementation goals and come back to them as needed, etc. I've had it running like this for awhile and it seems Qwen3.8-Flash-Next has no trouble understanding the rolling window.
GPT context window is way too low for me and my last experience with it (GPT 5.6 Sol) was so awful and I hit limits way too fast that I cancelled it (and at least got my money back).
I'm no longer using Pi since it got worse IMHO and Claude subs can only be used in Claude Code but I miss the /tree feature which is perfect for first letting the model read & cache the important bits of the codebase and then start your plan from there (as long as you stay in the Cache TTL). Claude Code has /rewind but it's not as good.
I'm only using the 20$ plans.
When cargo fails to build and creates a massive amount of compile errors, you're better off having this preprocess step.
I usually have my main agent write a wrapper command around things like that as it hits them. The wrapper only surfaces the important info, writes the full log to a file, and the agent gets some instructions on using sed and the like to navigate the output.
It seems to work reasonably well.
Claude already greps and tails every output by itself, a sub-agent would do the same.
They can, when they're forked off the main session instead of spawned from scratch.
https://github.com/can1357/oh-my-pi
I haven't tried it much though, can't vouch how well it works.
On top of that, somewhat unrelated I’ll agree but still, it has support for vim keybindings
Anecdotally, I find the auto compaction (or what I assume is happening when the context magically drops) to be hit or miss. I do like how easy it is to use my work cursor sub and business chat gpt at the same time. Then I use nearly free cursor models for dumb shit and Sol for real problems.
I did create some extensions where it spawns sub agents for specific tasks, especially when I want to keep the context clean or when I really want to offload a piece of work to a cheaper model. And for that I have a high degree of control over, I know which model is being used for each subtask.
I find Claude Code too unwieldy for my tastes. Pi's philosophy of being very light on features nut highly flexible for customization, clicked very well for the way I work.
OpenAPI is “intelligent tool discovery” (whatever that is). OpenAPI literally “returns structured data” and is “discoverable by their documentation and description.”
Just once, I’d like to see someone explain MCP in terms that suggest they have any idea what they are talking about.
https://chatgpt.com/share/6abd1842-cd34-83e8-997e-55c8bc23cd...
Part (a) in particular is pretty prominent for us since our uses cases are all about video, and having "just" an iframe URL to supply, with no other interaction with the model/turns, would give a pretty janky and unpleasant experience if Claude/ChatGPT even allowed it to embed.
Another perk is it allows me to run tools in the exact same way as agents instead of treating MCP as a special way to call on services. Super valuable when debugging.
It's typed so you can build some governance around it, by allowing only some tools or parameters for your org (this is a pretty weak point, but still)
A skill has one giant description from the frontmatter loaded into the context, where MCP loads a smaller one for every tool. Not necessarily better, the skill approach is often better actually, but sometimes the MCP approach fits more
With a skill, updates depend on whatever channel delivered it to you. Whichever channel that is, it's out of my hands as a provider.
So, MCP solves the problem of coordinated distribution of updates to a larger subscriber base. Think inside of a company, for example. I don't have to go around and tell people to `git pull` their skills folder.
There is an argument for and against having the model repeat this state.
I'm still on team CLI in that i think even designing an interface from the CLI perspective gets you a better domain interface compared to when you can 'cheat' with the MCP state.
But the thing MCP is just better at is credentials.
The thing that _was_ the dealbreaker between CLI and MCPs is that MCP's couldn't be composed. Maybe `codemode` fixed this; haven't tried enough to say 1 way or another.
MCP's have all the same advantages that a rest api has over a cli.
Codemode is a way for the LLM to orchestrate harness level tools. The reason this happening now, is because the models by the labs are increasingly trained on this. Codex for instance in responses lite requires codemode to even perform parallel tool calling.
* speed - much fewer hops back to the LLM
* fewer tokens - intermediate execution steps in the script don't leak into context, only the final result does.
* repeatability - if the LLM needs to repeat work, it can reuse a script it wrote last time.
If you have a harness that has access to a full shell and knows how to use bash or python, you'll often see it writing little scripts. For setups that don't (ie normal model API requests with tool calls), you can give it an lightweight secure execution environment like just-bash, or quickjs.
What if I have an MCP Tool LookupZip(City) and want to chain it with a bash tool that prodcues a list of 100 cities. And then I want to filter again to the largest Zip code.
It's pretty effective because of the reasons you noted, but there's a composability problem since each MCP has its own sandbox and can't call into the other ones.
IIUC Pi offer a workaround for this, the harness runs the sandbox and populate it with the MCP tools, that way the composability problem is solved and every MCP do not have to implement their own sandbox.
codemode lets you execute scripts in a runtime where your MCP tools are made available as function calls
this matters for cases where the MCP tool is the only way to do something and you do not have an equivalent CLI, API, whatever to script with
Another retrospect note, "No MCP" appears to be the first icon on their front page - not sure how I missed that.
Imagine my surprise reading this!
So I cut out a lot of functionality in my own fork
https://github.com/Pyrolistical/mi
I took “expected to customize basically everything” to heart
The first tool execution runtime in harnesses are direct tool calls with JSON or XML, such as the Read and Edit tools. As an escape hatch, we have Bash tool that allows arbitrary code execution on the host running the agent. The downsides of using bash (on the host) as the main tool execution runtime are:
- Syntax and obvious errors only surface at runtime
- Unergonomic orchestration of parallel and background tasks
- Verbose command output cluttering context
- Dependent on the host environment, packages versions, etc.
- No security measures by default.
To me the last point is the biggest inherent weakness, usually mitigated by creating a dedicated unprivileged user or running bash in a sandbox.
Note that direct tool calling is kind of the polar opposite on these points: syntax errors are caught early, orchestration can be done with some wrapping tools, command output is controlled, and most importantly they are more sandboxed. On the flip side, they obviously have way less power, necessitating Bash tool in the first place.
Codemode is the middle ground between these two extremes. It actually can be derived simply by one idea: what if we replace Bash by another language that can be checked for obvious errors, i.e. type checked?
Everything else falls out from there:
- Any language would do, but I think TypeScript fits the balance between safety, speed, conciseness, and popularity in training data.
- If we use TypeScript, might as well run it in a sandbox as JS runtimes have been designed with this in mind for 20 years
- Orchestration comes for free from the JS runtime. It's not more powerful, just more ergonomic.
- Since the tools are controlled by the harness and not dependent on the host, cloud agent becomes easier.
- With this in place, MCP are not very different from a tool provided to this sandboxed runtime.
Overall I find the benefits compelling enough, but we'll see if the heavily-RLed models these days will use it effectively.
There's just something that bothers me about this. Normally if LLMs want to compose multiple operations, they have the perfect tool for this: bash, or whatever other OS shell is available. It's why I was always confused by Codemode-type constructs for direct chaining of tool calls; see also the way highly-RL'd modern models will fall back to sed or python for complex file edits.
It seems like Codemode is raised here as the perfect tool for chaining or composing MCPs, but isn't that backwards? LLMs are already given the perfect tool for that, and the problem is that MCPs aren't exposed to that tool.
I many scenarios, e.g. running the harness server-side, as is the case for chat interfaces, you don't really want to expose OS shell access as that opens up a huge security attack surface.
It does, but a restricted user account mitigates the large majority of those issues. A sandbox mitigates even more.
The number of remaining exploits left is probably going to be the same as the number in the harness. More, in fact, as many of them have no human review anyway.
That's been a trivially solved problem for decades.
Like I said, AI bros vibecoding slop because they literally have no clue what they're doing.
Rootless immutable containers without shell access, or SaaS products from multiple vendors with WebAPIs as the only touch point.
Pi is primarily a coding agent, so yeah, code mode makes sense, but I've found that better MCP design saves everyone a lot of trouble and would also probably have improved the thing's reputation overall (I personally am not fond of the line protocol, would rather have protobuf and more typing, but it is what it is).
Bash is an interface to operate Linux computers. It's not a web interface at all. It's the worst interface for the web.
Code mode is just them running JavaScript for web based APIs. They are using the web native programming language for web tasks. It's actually completely logical once that is clear.
Bash for the web is a terrible idea.
Second, the point of Pi as a bare-bones BUT extensible harness is just that - a bare-bones harness, that can get anything one wants. Make that "thing" easy to get. Besides, bloat is bloat and every bloat is someone's must-have and vice versa.
I had also read somewhere that they have launched (already? not sure) their hosted model infra or something like OpenRouter (not sure). Nothing wrong with running a paying business along with a FOSS project but that's also there.
I send an Authorization header with a Bearer token but the procedure calls are sent in cleartext. Is this how "MCP" server is typically implemented (no encryption)
NB. I didn't write the aforementioned software implementing "MCP" server, that's someone else's work
No idea. The boring (in a positive sense) answer I'd expect for any backend API server is that encryption in transit is handled by TLS. So I'd expect either the MCP server in question can be configured to support TLS connections & refuse plaintext HTTP connections, or that for a production-like deployment it expects to be deployed behind a reverse proxy that is responsible for terminating TLS.
I will use this to talk to my manager about the project status
However- in my testing, mcp is really quite fast, and its pretty much free at this point- with frontier models. Context rot is, from what ive tested, not as much of a concern now. I genuinely was not able to hillclimb skills/extensions to beat out the speed of mcp in some cases I've been testing.
With an MCP, agents are very tolerant of changes, since each usage is fresh with no prior knowledge, so nothing to break. This allows you to experiment and iterate on the MCP, see how people use it and what gaps you still need to cover. Once it stabilizes, you can lock it down into and API/CLI interface.
MCP makes my boss happy so I’m happy
It's like people hate it so much they forget how easy it is to spin one of these up; it's not like I spent weeks and weeks working on this.
I spun up an MCP server for a personal finance app I've been working on. OAuth with personal tokens and all the tools I added- guess how long it took me? Like 2 hours maybe. Some people find it useful, including me, and some don't. That's okay!
Add in the context that they're also building their own inference provider[1], and there's certainly a non-zero risk that the project trends away from the barebones platform to build on that people signed up for.
criteria: {
none: "Neutral, factual, or friendly, even about a serious bug",
mild: "Explicit annoyance, impatience, or disappointment",
high: "Clearly angry, exasperated, sarcastic, or fed up",
},Doesn't have to be used for prioritization purposes. Obviously a little annoyance is to be expected at times, but you could explicitly ban egregious angry behavior in your project.
So, exactly like OpenAPI?
a subagent per mcp toolset, basically sub agent per saas in the configuration at $WORK
I don't have any hard evals on this but it feels like it works more than it doesn't, especially now that subagent use has gotten better in general, but you do feel drawbacks if you ever need to do something that involves coordinating multiple sub agents, there's some communication overhead that wouldn't be there if it was just one agent with all the tools to do the whole job
Treating MCP as a part of OpenAPI rather than a tool connector is a direction in which we're heading. It is important for the users to have the flexibility of deciding the model, work to be done and the tool call in one prompt. The framework sets up the configuration and gets the output.
I'm actually pretty happy that they did it, since 90% of what I have to integrate in enterprises is MCP-driven (it's a security and auth boundary that has become pretty much mandatory for any third-party agents wanting to reach into corporate data) and this lets me use Pi directly. Am just being cautious about the first version, because, well... it's a first version, and I like my tools stable.
(I actually played around with the idea of using QuickJS myself for codemode, but since I rely on Bun that gives me the ability to use other things... never got around to do it though.)
It's the cost efficiency that's a big deal. Jev is an excellent cheap "first pass" model.
Even Anthropic moved to auto-approve by default and they are kings of fearmongering.
Sure you can use curl along with it - but many MCP friendly harnesses don't have access to a tool like that. Wish they had added this before moving on to stuff like MCP apps.
The idea of using jev as a cheaper faster subagent for specific use cases is interesting. Will have to experiment with that!
It doesn't look to be asciinema?
Codemode isn't replacing MCP; it's fixing MCP's biggest flaw—its lack of composability.
While we are waiting on that to become stabilized, we implemented a inspired/co-evolved way to do that in our tool[0], where you mark individual fields in the request/response schema as being file payloads, so that file exchange can be properly orchestrated by the harness and doesn't pollute the context. We just do inline base64 uploads of the required payloads, which in practice we've seen to work quite will until ~100MB files (which is otherwise also the size limit we usually recommend for file processed).
It's annoying that it's not stabilized yet, but for most bigger customers we've seen, they implement 80% of the MCP servers they connect in-house, so doing adjustments to the tool surface, and metadata has been less of a pain for them than we expected.
[0]: https://erato.chat/docs/features/mcp_servers#file-support
If you are a Pi user it may be better to just ask your agent to explain https://github.com/earendil-works/pi/pull/10040
If we fail to explain it, then we need to do a better job explaining it :)
If you feel that Pi has been drifting away from its original vision, try hax (https://usehax.dev/) - you might like it.
Thank you. I've been frustrated by harnesses hijacking the terminal and breaking basic features such as scrolling and text selection.
It even sends BEL when the agent completes, which makes so much sense, yet Pi never implemented it.
I'm definitely going to use it over the next few days and hopefully make the switch.
I find this approach is easier to debug and I can also use the tool myself to ensure it's working well.
In this very post:
> While a lot of things have improved about MCP, quite a few have not. The biggest issue with MCP continues to be that it’s hard to compose. Even with codemode, which is just a neat little sandbox to allow composing of tool calls, MCP doesn’t fully deliver on this. But that at this point is less the problem of MCP but the MCP servers out there and different approaches of harnesses to work with them.
Can we get a bit more clarity on this? With codemode, what's the gap? I've also been investigating the search+execute MCP server pattern evangelised by Cloudflare (it uses codemode inside the MCP server to bypass the need to expose a large number of individual tools), but i've seen people say that doesn't compose well either.
Just generate CLI tools, with docs, from MCP servers on demand.
Swagger was the new old name for the concept
1: https://swagger.io/docs/specification/v3_0/adding-examples/
Spec wise: Ampere 128C, 256gb, 4xa4000,& 7800x3d, 96gb, Blackwell 24gb, 4070s. Nvidia fiber and mirkotik. Working on RDMA (but isn't really needed yet, but we have it)
OS - Tumbleweed ARM headless, CachyOS x86 hyprland. It's been easier staying on the edge of the kernel with less abstractions.
There were already tons of third party MCP extensions that were every bit as good as this half baked built in.
Already lots of code mode extensions too, which is also opinionated.
Sad to see Pi to go from being: "This harness is yours"
to
"Just another coding harness"
But pi doesn't have security. Pi is full yolo, it's on the user to run it in an environment that minimizes the blast radius if the LLM goes haywire.
So I'm not sure what the plan is here. Will pi support running certain tools like bash as a different OS user than owner of the pi process?
that's my one problem with the cli vs mcp debate. I prefer cli, because they're tools I'm familiar with, and they're composable. My problem with it is some cli tools need to use secrets to access things. By making sure they only go to the harness, then I've got my problem solved.
MCP is just REST with meta data folks. It's all just HTTP and JSON.
Nowadays they moved to http and stateless.
And you didn’t remember that when you said no to MCP?
No, no MCP for now?
Unix programs are said to be composable, in that you can combine them in different ways to get more complex and useful functionality. But the applications themselves are not composeable. They just take input, perform calculation, and return output. Each app has unique input and output, and none know about other apps. So how can they possibly work together?
Bash acts like a programming language, allowing you to write a new program on the fly. It is very lightweight, but gives enough functionality to do two things: 1) call arbitrary programs, 2) connect their inputs and outputs, 3) make decisions about how to do this to result in a novel solution. It uses Unix pipes to make writing the program easier, but actually it could work fine without pipes, reading/writing program input/output with files.
The important thing is bash is a "glue" program that ties together the other programs. Bash is what makes those programs composeable. That, and the Unix API (execve(), open(), read(), write(), close()) that allows making the call, passing input, reading output, in one standard way for all programs.
MCP is both the application to call, and the Unix API to call them. Your agent harness is bash. Codemode is certainly a neat/more efficient way to write the code, but it's not necessary. What's necessary is getting a very large library of programs, like Unix programs (find, cat, grep, sed, awk, tr, cut, sort, tail, head, etc) that each have powerful functionality. The more programs you have, the more your bash script can do to chain them together and get more powerful results.
LLMs only use Bash because they didn't yet have MCP and a large library of MCP programs. Bash is a stopgap solution. The future is MCP (or whatever replaxes it).