226 karma · joined August 5, 2021
I build AI agent infrastructure. The post came from a real debugging session. An agent modified 47 files, the build failed, and I spent twenty minutes scrolling terminal output before giving up and starting over.
The core argument: we solved observability for microservices over the last decade (OpenTelemetry, Datadog, Honeycomb, Grafana). AI agents are also distributed systems. Multiple LLM calls, tool invocations, file operations, decision points. But there is no structured trace, no cost attribution per task, no permission audit trail, and no session replay.
Four questions you cannot answer today:
1. What did the agent do? (no structured trace) 2. Why did it do it? (context is ephemeral) 3. What did it cost? (no per-task attribution) 4. What was it allowed to do? (no permission audit trail)
The patterns exist in distributed systems observability. They need to be adapted, not invented. OpenTelemetry's data model (trace IDs, spans, parent-child relationships) maps directly to agent execution.
Happy to discuss the technical details. Particularly interested in hearing from teams that have built ad-hoc agent logging and what they learned.
Earlier this year. The DMs from that post, hundreds of engineers describing the same problems - made it clear there was no single reference covering the full agent stack for engineering teams.
33 chapters, 10 parts. Early version, open source.
There are rough edges and likely mistakes. PRs welcome:
https://github.com/Siddhant-K-code/agentic-engineering-guide
Happy to answer questions about any of the chapters.
The result is a 33-chapter guide, free to read online: https://agents.siddhantkhare.com
Individual patterns for using agents well are one piece. The infrastructure, security, and team practices for running them safely at scale are another. This covers the second part.
The result is a 33-chapter guide, free to read online: https://agents.siddhantkhare.com
Early version — open source, CC BY-NC-SA 4.0.
There are rough edges and likely mistakes. Corrections welcome: https://github.com/Siddhant-K-code/agentic-engineering-guide
Individual patterns for using agents well are one piece. The infrastructure, security, and team practices for running them safely at scale are another. This covers the second part.
The Check Point disclosure this week (CVE-2025-59536, CVE-2026-21852) showed that malicious repo configs could execute shell commands and steal API keys before the trust prompt even appeared. Anthropic patched the specific bugs. But the underlying problem is architectural.
Claude Code gives you two options: approve every mkdir & npm test individually, or pass "--dangerously-skip-permissions" & give the agent unrestricted access to your filesystem, network, and shell. Most devs end up on the second option within a week.
We solved this for CI/CD and service accounts decades ago. Declarative policies, scoped permissions, audit trails. None of that exists for AI agents yet.
The post lays out what a real permission model would look like: declarative policy files per project, relationship-based scoping (so a feature branch agent gets different access than a production hotfix agent), and structured audit logs by default.
Happy to answer ques. about the auth patterns or the Check Point findings.
The short version: AI made code generation fast, but nobody invested in making code verification fast. The human became the bottleneck. That's what causes the fatigue.
The fix is backpressure, a systems engineering concept. Automated feedback (types, tests, linters, architectural rules) that catches agent mistakes before they reach you.
A few things I learned from talking to teams:
- One team cut their test suite from 15 min to 90 seconds specifically for agent iteration speed. Paid for itself in a week.
- Pre-commit hooks went from "annoying" to essential. Agents don't complain. Turn everything on.
- BoundaryML calls this "agentic backpressure" - they did a whole podcast on it. The Ralph Wiggum loop community builds workflows around the same idea.
- The post has a hierarchy (types > tests > linters > architectural rules > human review last) and a Monday-morning checklist.
Backpressure won't catch everything, an agent can pass every test and still be the wrong approach. But it reduces the noise so you can focus on the signal.
Curious what feedback loops you've built around your agents.
“My machine” is identity and control. That worked when the human was the execution engine. With agents, undocumented setup turns from annoyance into hard failure.
But that agent workflows expose environment debt the way CI exposed testing debt.
If an environment can’t be provisioned programmatically, it’s not infrastructure, it’s folklore.
The companies moving fastest right now have all done the boring work first: to standardize their dev environments. The agent harness is a thin layer on top.
Let us know what you think!
I've started doing it now, still needs to work on it. Thanks for the tip though, i hope it is working well for you!!
Most frameworks show token latency. But why are some tokens slow? What’s stalling the GPU? Is it poor SM occupancy? Kernel launch delay? Cache stalls?
I built *LLMTraceFX*, a token-level GPU profiler for LLM inference workloads.
*What it does:* - Parses GPU execution traces (like vLLM outputs) - Analyzes performance at token granularity - Detects kernel-level bottlenecks: stall %, cache latency, launch overhead, etc. - Uses Claude API to explain why a token was slow and how to optimize it (e.g. "fuse kernels", "fix memory access pattern") - Generates flame graphs + bottleneck dashboards
*Output*: JSON reports, HTML dashboards, CLI summaries, Claude suggestions.
*Stack*: Python, FastAPI, Plotly, Modal.com for GPU runtime, Claude API (no infra required)
*GitHub repo*: https://github.com/Siddhant-K-code/LLMTraceFX
---
Would love feedback on: - Other formats to support (HuggingFace, llama.cpp, ONNX?) - What to show beyond kernel breakdowns - Ideas for integrating with compilers, optimizers
Cheers!
It handles: • Synonyms & fuzzy queries (e.g., Paracetamol ≈ Acetaminophen) • Multilingual data • Contextual search beyond keywords
No Credits given :( => https://github.com/conwnet/github1s/issues/346