9,259 karma · joined January 26, 2017
There are a lot of ways to implement sub-agents, and it's not something I need the harness to be opinionated about.
I know builtin tools support opt-out, but it's more bloat. It's also more complexity for the agent to understand when you use it to build extensions for itself.
I'm not so sure about this move, or the general inclusion of code mode in the core editor as one of pi's main selling points was its minimal nature.
Nether service has a way to configure fine-grained access for a secondary user.
How do I go about giving the agent the ability to perform these tasks without exposing myself to the risk of unexpected destructive behavior from the agent?
And I love pi - it's my daily driver - but the extension system itself is an attack vector. If any process manages to write an extension to your .pi directory, it could rewrite your prompt to have the agent exfiltrate your secrets, or take whatever action on the host system if you don't sandbox it.
When I was kicking the tires on pi, one of the first things the agent did was push an update to one of my published Rust crates (not the project it was working on).
That in itself wasn't harmful, but it did convince me it was worth the effort to figure out sandboxing after that.
For something like muse that's supposed to be a general-purpose assistant, how do you give it enough access to be useful, without giving it too much access, and creating unacceptable risks? And how do you do that in a way that's comprehensible the average Facebook user who's the target market of this product?
I.e. if I have an agent running in a WASM sandbox with no access to the host system, I can't ask it to clean up my files. Same thing with things like giving an agent access to your email inbox: doing so allows the agent to provide utility, but it comes with risks, as the agent can delete important emails, or leak sensitive data.
I think a big part of the problem is, a lot of the systems we use and would like agents to help us with don't have a concept of separated roles with different levels of access which can be applied. A lot of times it's all or nothing.
And even when we do have fine-grained access control available, it's a pain in the ass to manage it. Like you can create a GitHub token with fine-grained access control to your repositories and make sure the agent only uses that one to connect, but it's a whole lot easier to use a broad-access token, or just let the agent use your own token, so lots of people will just end up doing that.
And I also like pi, but it's probably one of the worst in terms of sandboxing as it's yolo by default.
Even as a technical person, it's not trivial to sandbox agents correctly. The fact that an mis-clicked permission popup could give an agent unrestricted access to a user's disk is a massive risk vector in the hands of lay people who barely understand how any of this works.
So much of current security depends on the model of tying access control to a user account. A lot has to be re-thought in terms of how to grant access to an agent working on the user's behalf, in a way that doesn't make it completely useless, and also doesn't require every user to become a sysadmin managing fine-grained agent permissions manually.
The issue I have run into is that for running llama.cpp, which is a very actively developed bleeding-edge software, it seems like experimenting with different versions/configs/patches etc. was fighting with the Nix philosophy of immutable software, so I ended up just managing it outside of nix.
But this could be user-error, I've only spent a couple of weeks with it.
Of course I don't know if there's really a way for this to be molded in current LLM's (sounds more like diffusion)
The main place they are a bit behind is in the number formats they support natively. Iirc M5 doesn't have native FP8 support, so you will take a speed penalty on quants where other architectures get better acceleration.
And given how fast the hardware and software is evolving, I can easily imagine a future where we all have very capable models running on our own devices for an embedded intelligence layer that's doing most of the day-to-day tasks, and only have to outsource to a super-smart cloud model for specific things.
While we're at the point where traditional software is pretty much fast enough for all but extreme use-cases, with LLM's it feels like we're back to the days where you press compile and go have a coffee or chat to your colleague.
But I think this will eventually be a problem solved at the OS level in a more streamlined way. I.e. there will be fine-grained permissions you need to approve to give an agent access to the system.
I'm not a fan of big-tech acquisition results either, but one benefit can be that a product continues to exist when it would otherwise become insolvent.
Models are already largely hardware agnostic. It would be pretty hard to put that cat back in the bag.
I could imagine them building value-added services on top of HF to advantage Nvidia products (i.e. "run this model on NVIDIA cloud" with one-click), but in this moment it's hard to imagine how they could actively disadvantage models built to run on other platforms.
My use-case is coding, currently working on a project with a Rust backend and TS/React/Vite frontend, with probably tens of thousands of lines of code total (including tests).