HNHacker News
TopNewBestAskShowJobs

nordic_lion

1 karma · joined August 14, 2024

submissionscomments
nordic_lion··on Ask HN: How do you shut down misbehaving AI in production?
That makes sense binding to the smallest viable control surface, and the sampling strategy for hot paths sounds like a pragmatic balance between latency and coverage. Thanks for the additional feedback here.
nordic_lion··on Ask HN: How do you shut down misbehaving AI in production?
One thing I’m still unclear on: what runtime signal is the soft-rule evaluator actually binding to when it decides “semantic drift”?

In other words, what is the enforcement unit the policy is attached to in practice... a step, a plan node, a tool invocation, or the agent instance as a whole?

nordic_lion··on Ask HN: How do you shut down misbehaving AI in production?
Where/how do you define the policy boundary line that triggers course correction?
nordic_lion··on Ask HN: How do you shut down misbehaving AI in production?
Solid breakdown. Yeah, the agent-level observability gap you call out is real.

One direction been considering is to treat intent and plan state as first-class runtime signals (for example, intent spans and plan-step annotations alongside traces), so when agent goes offrail can at least see what it believed it was doing, not just which calls it made.

nordic_lion··on Ask HN: For vertical AI does MCP break need for 3rd party API libraries?
Great points, and thanks for sharing your experience (Skeet looks sweet)—this is exactly the kind of real-world insight that’s missing from the MCP hype.

1. api design vs. natural language - totally agree, porting specs directly to MCP is little like fitting a square peg into a round hole. APIs weren’t built with conversational interfaces in mind, and the weird outputs shows. These quirks make it clear that MCP needs a lot of thoughtful abstraction to “just work” for end users.

2. Definitely early days still… the gap between the MCP promise and its current state is real. OAuth setup, inconsistent auth models, and half-baked community servers are huge speed bumps. It’s easy to talk about how MCP simplifies things, but the devil’s in the details—and those details are still messy.

3. SSE scalability - yea, this is a big one. Getting SSE to work reliably at scale is a nightmare, and it’s not something you can handwave away. Until MCP proves it can handle real-world, production-grade workloads, it’s hard to fully buy into the vision.

MCP feels like it’s addressing the right problems, but it’s clear we’re still in the “rough draft” phase. Curious to see how much of this friction gets smoothed out over the next year or so