HNHacker News
TopNewBestAskShowJobs

skohan

9,259 karma · joined January 26, 2017

spencerkohan (at) gmail
submissionscomments
skohan··on OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network
If they take Neural Rendering far enough, they can get rid of those useless raster and RT cores completely and ship compute and tensor cores only on all their chips
skohan··on You Said No MCP
I think MTP is kind of ok, since it's a common protocol, but I think incorporating sub-agents could be a mis-step.

There are a lot of ways to implement sub-agents, and it's not something I need the harness to be opinionated about.

I know builtin tools support opt-out, but it's more bloat. It's also more complexity for the agent to understand when you use it to build extensions for itself.

skohan··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Do you have any plans to support ROCm?
skohan··on Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Do you have anything published on the quality benchmarking using your caching strategy?
skohan··on You said no MCP
You could already use MCP perfectly well on pi via extensions.

I'm not so sure about this move, or the general inclusion of code mode in the core editor as one of pi's main selling points was its minimal nature.

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
Let's say I have an aligned agent, and I want to give it access to my bank account to track my budget, and my email so it can automate responses to certain messages, and automatically unsubscribe from junk emails.

Nether service has a way to configure fine-grained access for a secondary user.

How do I go about giving the agent the ability to perform these tasks without exposing myself to the risk of unexpected destructive behavior from the agent?

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
So how would you go about sandboxing Muse in a way that it still functions as a general purpose assistant?
skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
Yeah I think people are probably widely underestimating the risks. Like if you gave another programmer unlimited access to your system and ssh keys, and they never slept and could code and run terminal commands at dozens or hundreds of words per minute, you would have to trust them a lot to give them that.

And I love pi - it's my daily driver - but the extension system itself is an attack vector. If any process manages to write an extension to your .pi directory, it could rewrite your prompt to have the agent exfiltrate your secrets, or take whatever action on the host system if you don't sandbox it.

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
That helps you if the blast radius is on your machine, but it doesn't really help if the agent is using your credentials to cause some damage with some remote system you have access to.

When I was kicking the tires on pi, one of the first things the agent did was push an update to one of my published Rust crates (not the project it was working on).

That in itself wasn't harmful, but it did convince me it was worth the effort to figure out sandboxing after that.

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
I sandbox my coding agents using bubblewrap, but I don't think it's as trivial a problem as you make it sound.

For something like muse that's supposed to be a general-purpose assistant, how do you give it enough access to be useful, without giving it too much access, and creating unacceptable risks? And how do you do that in a way that's comprehensible the average Facebook user who's the target market of this product?

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
Sandboxing the agent application is easy. The tough part is sandboxing in such a way that it's still useful.

I.e. if I have an agent running in a WASM sandbox with no access to the host system, I can't ask it to clean up my files. Same thing with things like giving an agent access to your email inbox: doing so allows the agent to provide utility, but it comes with risks, as the agent can delete important emails, or leak sensitive data.

I think a big part of the problem is, a lot of the systems we use and would like agents to help us with don't have a concept of separated roles with different levels of access which can be applied. A lot of times it's all or nothing.

And even when we do have fine-grained access control available, it's a pain in the ass to manage it. Like you can create a GitHub token with fine-grained access control to your repositories and make sure the agent only uses that one to connect, but it's a whole lot easier to use a broad-access token, or just let the agent use your own token, so lots of people will just end up doing that.

And I also like pi, but it's probably one of the worst in terms of sandboxing as it's yolo by default.

skohan··on Unsurprisingly, Meta's new Muse AI agent blatantly ignores users permissions
Agent sandboxing/access control is one of the biggest problems to be solved before this technology really should go mainstream.

Even as a technical person, it's not trivial to sandbox agents correctly. The fact that an mis-clicked permission popup could give an agent unrestricted access to a user's disk is a massive risk vector in the hands of lay people who barely understand how any of this works.

So much of current security depends on the model of tying access control to a user account. A lot has to be re-thought in terms of how to grant access to an agent working on the user's behalf, in a way that doesn't make it completely useless, and also doesn't require every user to become a sysadmin managing fine-grained agent permissions manually.

skohan··on Claude partial outage
A more cynical version of me could imagine scheduling an outage to make people think about how dependent they are on your product in advance of the IPO
skohan··on Dutch governments builds alternative for Microsoft based on NixOS
I recently got my feet wet with Nix for the purpose of agent sandboxing, and it's a really great concept!

The issue I have run into is that for running llama.cpp, which is a very actively developed bleeding-edge software, it seems like experimenting with different versions/configs/patches etc. was fighting with the Nix philosophy of immutable software, so I ended up just managing it outside of nix.

But this could be user-error, I've only spent a couple of weeks with it.

skohan··on M5 Ultra Mac Studio Review
1.2 T/s is not that modest is it? That's very close to an RTX pro 5000
skohan··on Claude Fable 5.1 and Claude Mythos 5.1
I don't think smart people generally solve problems by talking through reasoning steps at a mile a minute. They clear their mind and let the solution come.

Of course I don't know if there's really a way for this to be molded in current LLM's (sounds more like diffusion)

skohan··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
16GB VRAM is probably not quite enough. With 32GB, or maybe even 24GB, you can do serious coding work with local models.
skohan··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
For decode, memory bandwidth is the main bottleneck, so these machines will likely perform well even without a ton of GPU horsepower. Not as well as Blackwell, but I expect they will be a reasonable choice in terms of price/performance if you want to run large models with a lot of context.

The main place they are a bit behind is in the number formats they support natively. Iirc M5 doesn't have native FP8 support, so you will take a speed penalty on quants where other architectures get better acceleration.

skohan··on Asahi Linux Progress Report: Linux 7.2
That's what I mean though. The comment I was responding to was talking about "forever laptops" - my point is there's plenty of room for new capabilities which will make current hardware obsolete. Just like how GPU's didn't exist at all, and became a standard part of computing.

And given how fast the hardware and software is evolving, I can easily imagine a future where we all have very capable models running on our own devices for an embedded intelligence layer that's doing most of the day-to-day tasks, and only have to outsource to a super-smart cloud model for specific things.

skohan··on Asahi Linux Progress Report: Linux 7.2
As someone who does a lot of work with local LLM's, today's systems feel woefully under-powered. I'm looking forward to a future where my laptop has 10x the memory, 100x the memory bandwidth, and optimized cores to make inference workflows that currently take minutes or hours go down to seconds or milliseconds.

While we're at the point where traditional software is pretty much fast enough for all but extreme use-cases, with LLM's it feels like we're back to the days where you press compile and go have a coffee or chat to your colleague.

skohan··on The Harness Is the Thing
Ok that's fair if it's targeting a non-technical audience.

But I think this will eventually be a problem solved at the OS level in a more streamlined way. I.e. there will be fine-grained permissions you need to approve to give an agent access to the system.

skohan··on The Harness Is the Thing
That's the most basic version of a harness. But a harness is really about automating context delivery to the LLM based on your use-case.
skohan··on The Harness Is the Thing
Is that the harness' job? It seems to me the best place for sandboxing is at the OS level (i.e. running the harness inside a container with correct access configured).
skohan··on The Hugging Face incident and the road ahead
And despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost. Sometimes constraints are healthy for inducing creative solutions.
skohan··on Nvidia agrees to acquire Hugging Face for $13B
Is huggingface profitable?

I'm not a fan of big-tech acquisition results either, but one benefit can be that a product continues to exist when it would otherwise become insolvent.

skohan··on Nvidia agrees to acquire Hugging Face for $13B
That could be a factor, but the optimistic interpretation would be that they want to support the open ecosystem because it sells more chips.

Models are already largely hardware agnostic. It would be pretty hard to put that cat back in the bag.

I could imagine them building value-added services on top of HF to advantage Nvidia products (i.e. "run this model on NVIDIA cloud" with one-click), but in this moment it's hard to imagine how they could actively disadvantage models built to run on other platforms.

skohan··on Nvidia agrees to acquire Hugging Face for $13B
They didn't, but their price increases haven't kept pace with general memory price increases (yet) so currently their pricing seems reasonable.
skohan··on Nvidia agrees to acquire Hugging Face for $13B
Yeah I think this is probably one of the best outcomes you could hope for if you want to use open models.
skohan··on GLM-5.3-Flash
I'm using unsloth dynamic Q4_K_XL.

My use-case is coding, currently working on a project with a Rust backend and TS/React/Vite frontend, with probably tens of thousands of lines of code total (including tests).

skohan··on GLM-5.3-Flash
I'm using the unsloth dynamic Q4 and getting good results. I was running Q5, but Q4 gives more context headroom so I can run two agents in parallel with ~100k context each with 32GB vram.
Page 1 of 34Next →