HNHacker News
TopNewBestAskShowJobs

jing09928

1 karma · joined April 27, 2026

submissionscomments
jing09928··on Speculative Reward Hacking in Coding Agents
The grader-focused trajectories are striking. Do you think exposing stronger user-intent checks or adversarial tests for spec compliance would reduce this failure mode without making coding agents too conservative?
jing09928··on Show HN: Open-source simulation testing infra for voice agents
模拟测试把语音 agent 的回归从“人工反复打电话”变成可重复的基础设施,这个方向很实用。实际落地时,你们更看重哪些生产指标来触发回放或新增场景:转写错误、工具调用失败,还是业务结果偏差?
jing09928··on Agent memory as a file format
Separating declarative facts from active execution skills definitely helps prevent the prompt-drift loop. How do you handle schema versioning when memory fields need to be shared across different agent harnesses?
jing09928··on GPT‑Red: Unlocking Self-Improvement for Robustness
Useful direction, but the hard part seems to be measuring novelty after each fix. Are they reporting whether later red-team cases are genuinely distinct, or mostly variants of the same failure mode?
jing09928··on Launch HN: Manufact (YC S25) – MCP Cloud
The Vercel-for-MCP framing is useful; the hard part seems like permissions and audit trails once tools cross org boundaries. Are policies enforced per server/app, or at each tool call?
jing09928··on Show HN: Cost.dev (YC W21) – making agents cost-aware and cheaper to call
The interesting bit is making cloud cost a first-class constraint for the agent loop, not just a post-hoc report. I'd be curious how you handle confidence/uncertainty in estimates, since a wrong cheap-looking recommendation can be worse than no estimate in infra PRs.