HNHacker News
TopNewBestAskShowJobs

evalstate

51 karma · joined December 17, 2024

open source work at huggingface.co

maintainer of fast-agent.ai

submissionscomments
evalstate··on Stateless MCP has recaptured my interest
There's a nice feature in the new version that lets you copy tool _arguments_ in to the HTTP Headers for custom routing etc.
evalstate··on Stateless MCP has recaptured my interest
fast-agent is very good at this (disclosure: author)

you can specify --url, --npx and include --auth $TOKEN on the command line. you can also interactively connect with /connect, and pixel peep transport details https://fast-agent.ai/mcp/mcp-inspect-transport/

evalstate··on Hugging Face Skills
Gotcha - yeah, it removes the tool calling step so their content is always in context (noting they took action to try and reduce the size of that). The framing seems a little simplistic -- thanks for the link.
evalstate··on Hugging Face Skills
I think the paper is saying specifically that it's redundant to include information about your coding repository when that information is otherwise available to the agent in higher fidelity forms (e.g. package.json). This makes sense - but not sure it's about Skills directly.

For the former I'd be interested in learning more about that. From a harness perspective the difference would be the inclusion of the description in the system prompt, and an additional tool call to return the skill. While that's certainly less efficient than adding the context directly I'd be surprised if it degraded task performance significantly.

I tend to be quite focussed with my Skill/Tool usage in general though, inviting them in to context when needed rather than increasing the potential for model confusion.

evalstate··on Pi – A minimal terminal coding harness
fast-agent lets you do this as well (and has a skill in its default skills repo to help with automation/running in container/hf job).
evalstate··on Hugging Face Skills
Yes -- skills live in a special gap between "should have been a deterministic program" and "model already had the ability to figure this out". My personal experience leaves me in agreement that minimal system prompts are definitely the way to go.
evalstate··on What I learned building an opinionated and minimal coding agent
An excellent piece of writing.

One thing I do find is that subagents are helpful for performance -- offloading tasks to smaller models (gpt-oss specifically for me) gets data to the bigger model quicker.

evalstate··on No management needed: anti-patterns in early-stage engineering teams
A lot of those books are more about persuasion than motivation - they can look similar from a distance.
evalstate··on Toad is a unified experience for AI in the terminal
fast-agent has ACP support and works well with ollama. Once installed you can just use `toad acp "fast-agent-acp --model generic.<ollama-model>"`.
evalstate··on A2UI: A Protocol for Agent-Driven Interfaces
I quite like the look of this one - seems to fit somewhere between the rigid structure of MCP Elicitations and the freeform nature of MCP-UI/Skybridge.
evalstate··on MCP Specification – version 2025-06-18 changes
Structured Output in this case refers to the output from the MCP Server Tool Call, not the LLM itself.
evalstate··on MCP Specification – version 2025-06-18 changes
Yes. VSCode 1.101.0 does, as well as fast-agent.

Earlier I posted about mcp-webcam (you can find it) which gives you a no-install way to try out Sampling if you like.

evalstate··on Directory of MCP Servers
The list here https://modelcontextprotocol.io/clients has a number of Host applications, frameworks etc.
evalstate··on OpenAI Audio Models
Really looking forward to integrating with these models.

The next version of Model Context Protocol will have native audio support (https://github.com/modelcontextprotocol/specification/pull/9...), which will open up plenty of opportunities for interop.

evalstate··on Another late-night Claude Code post
The GitHub logo between the File Attach, Screenshot and Google Drive icons.
evalstate··on Another late-night Claude Code post
I've spent more on Claude Code than I'm willing to admit, and I'd estimate about 50% of the spend is "written off". With CC it can be tough to judge when to stop on a particular path.

Prior to CC I was using Goose, which is similar - and it's hard to tell how much better CC it than Goose as Sonnet 3.7 was released at the same time I switched. One of the nice features of Goose that CC doesn't have is loadable history/resumable sessions.

The workflow I've actually found most effective (and cheaper) now is to use the GitHub integration in the Claude.ai to get started, then use Claude Code to fill in the bits. The GH integration is much better than I expected, and worth a try if you have a Claude plan.

evalstate··on MCP vs. API Explained
Most LLMs including Claude struggle using the @modelcontextprotocol/server-filesystem server - it's way too complex (and tools like Goose wrap ripgrep in an MCP to handle it). A simple MCP Server with the SDKs can easily be less than 20 lines of code and be useful.

I wrote mcp-hfspace to let you connect to Hugging Face Spaces; that opens up a lot of image generation, vision, audio transcription and other services that can be integrated quickly and easily in to your Host app.

evalstate··on Show HN: Fast-agent – Compose MCP enabled Agents and Workflows in minutes
The Messages API contains a special section for placing Tool information, which is added to the Context Window - and it's this information that the Model then uses to decide whether to attempt a Tool Call.

In that case, we configure the MCP Server, and then the Host application (in this case fast-agent) uses the Anthropic or OpenAI API to populate it, and they inject it in to the Context Window[1] in the format best for their model.

So for fast-agent, we can set the model when we define the agent with `model="o3-mini.medium"` or from a command line switch. Depending on the type of eval you are doing you could for example use a Parallel workflow to see how the different models perform. Quite often, given a failing tool call the model will attempt to recover (the @modelcontextprotocol/server-filesystem is... an interesting example).

Another fun one is to use Opus 3 tool calling, where it emits <thinking> tags showing how/why it's calling it.

One final point is that different combinations of tools will give different behaviours - if 2 MCP Servers have similar definitions, it will degrade performance... One of the motivations for fast-agent is precisely because it allows dividing tasks up amongst different context windows to get the sharpest performance.

Link to the Anthropic docs as it's my preferred explanation. The Messaging API's grab the JSON and present it as Tool Call types - other models will simply emit JSON and let the Client handle it.

[1] https://docs.anthropic.com/en/docs/build-with-claude/tool-us...

evalstate··on UI is hell: four-function calculators
just tested this. in "standard" mode you get 7, switch to "scientific" you get 9.
evalstate··on Systems ideas that sound good but almost never work
Gerry Weinberg wrote beautifully about these "lullaby words". https://www.humansystemsinaction.com/lullaby-language/

> “Precisely. It’s what I call a ‘Lullaby Word.’ Like ‘should,’ it lulls your mind into a false sense of security. A better translation of ‘just’ in Jeff’s sentence would have been, ‘have a lot of trouble to.'”