2,796 karma · joined December 13, 2012
thomasalwyndavis at gmail dot com @ajaxdavis on x
Not overly confident in my position, but I believe agents prefer the extra information albeit noise to some.
Maybe users reporting otherwise are just looking at their client reports which wouldn't be able to tell the difference.
Somewhere should definitely make this for missing persons.
I have done interviews with companies that I generally thought were wholesome enough, but you can't control how individuals feel on certain days, they could be going through some dark days at home etc.
I'm not sold on AI interviews, but it could actually end up letting you fully share your experience more than a human could on average.
---
If you have codex I just typed "can you find in my home directory the last claude session I had"
And go from there if you got the fu
Don't hate me aha and no, there is no reason other than I can
---
Also, I'd love to use these sound effects, but I am an rts player and love aoe and wc franchise, these noises just trigger me to want to play too much.
---
Also, also, if you haven't seen AgentCraft, you are missing out -> https://x.com/idosal1/status/2021661861163544818 (worked in one npx command for me using my claude, a+ for creativity and smoothness)
The idea of representing UI as state goes back forever, I’m not that old but at least in the advent of the web, plenty of JSON -> UI specs or libraries have came into existence. If the specification is solid and a large portion of people agree upon it, I don’t doubt it will take over what we think of UI. (current contenders are json_render, a2ui etc)
The first benefit being that if I can describe my entire UI and the actions each component wants as json, theoretically, I can pass that file to any client to render be it mobile, react, a java swing app etc The responsibility of rendering becomes of anyone who wants to do it.
UI JSON -> UI Framework <- Design Tokens
Above is a simple way of describing how it would generally work. Where the UI framework can be whatever it wants to be as long as it knows how to connect up the UI JSON in a meaningful way.
Now for existing apps and their respective UI’s it’s never made all that much sense to describe how your components behavior in state, useful for some, and many have done it, but a hard pitch for others.
In the agentic era, the pitch is a lot more appealing.
- LLM’s are great-enough at writing JSON
- A lot of people are sharing the sentiment that they can just vibe code small apps for themselves. Hinting at they love the actual ability for full personalization.
Though having the user generate HTML and the rest all the time by LLM’s is more error prone, slow and costly.
The user can just ask an LLM to compose a composition of components in JSON laid out how they want and connected to the API’s they care about. (that can be rendered anywhere)
Personally, if I had a catalogue of 100 distinct services/API's, and I could ask an LLM to generate a UI in JSON that I can copy and paste anywhere to render it, I would be in heaven.
If I had subscriptions to services that; (fake services)
- EMAILR: Sent Emails
- BOOKLAND: Explore books
- DEEP_RESEARCHER: Researches Topic
I could ask an LLM to "With my services, EMAILR, BOOKLAND and DEEP_RESEARCHER and their attached tools.
Can you generate me a dashboard that lists out the top 20 BOOKLAND books, below each one added a button that posts the book title to DEEP_RESEARCHER when I click it. Also add a button below each book that uses EMAILR to email me them"
It would then return something like;
{
"view": "dashboard",
"title": "Book Research Hub",
"children": [
{
"type": "grid",
"columns": 4,
"source": { "service": "BOOKLAND", "action": "list", "params": { "limit": 20, "sort": "top" } },
"each": {
"type": "card",
"children": [
{ "type": "text", "bind": "$.title", "variant": "heading" },
{ "type": "text", "bind": "$.author", "variant": "muted" },
{ "type": "image", "bind": "$.cover_url" },
{
"type": "button",
"label": "Deep Research",
"variant": "primary",
"action": { "service": "DEEP_RESEARCHER", "method": "research", "params": { "topic": "$.title" } }
},
{
"type": "button",
"label": "Email me this",
"variant": "secondary",
"action": { "service": "EMAILR", "method": "send", "params": { "to": "$user.email", "subject": "Book: $.title", "body": "Check out $.title by $.author" } }
}
]
}
}
]
}Users could share their layouts and what they like and you could end up with a market place or sane defaults for those who don't want to bother with describing what they want. No longer do you have to rely on the UX team of the service for it to be laid out how you want.
There is a metric tonne of work that has to be done to make a specification that can handle more complex things. But I'd bet a lot of users will learn to love and appreciate that the 5% of features they care about they can finally just actually place how they want it to across all their disparate apps.
which lets users create collections of any publicly available tool which let's us do awesome things
We can auto generate a skills.md for collections https://tpmjs.com/ajax/collections/unsandbox/skills.md (https://tpmjs.com/ajax/collections/unsandbox to see the collection)
We analyze all the tools source code to write the skills.md using LLM's and then we auto inject three different ways agents can iteract with them
- hosted mcp servers for collections automatically e.g. https://tpmjs.com/api/mcp/ajax/unsandbox/sse
- a cli tool that can invoke each tool in the collection e.g. tpm run --collection ajax/unsandbox --tool [tool-name] --args '{"key": "value"}'
- restful endpoints for direct execution of tools e.g. curl -X POST https://tpmjs.com/api/mcp/ajax/unsandbox/http
It's very much alpha, but after building this, I don't see why you can't let agents choose any method they want of interacting with your service.
It means your agents can use CLI, MCP or REST, only in one style or any style all at once.
Side note: You gotta hustle people!
tpmjs is a registry of ai sdk npm packages (i created them for sprites), which you can add to personal collections, we automatically server your collections as mcp servers if you want.
Creating sprites in an example chat bot -> https://imgur.com/a/ETNxR1o
Creating sprites in claude desktop -> https://imgur.com/myC0U28
Listing out my sprites in claude desktop -> https://imgur.com/rgBU0jm
---
You can view the collection of tools here -> https://tpmjs.com/ajax/collections/sprites (fork it to use it yourself)
I'm looking into exe.dev and sprites.dev to build out extra features into tpmjs, agent sandboxes make a lot of sense.
---
> which ironically I think Scott would have liked
Agreed, RIP.
I like your thought process, and agree with it all.
Everything other than what you described towards the end seems easy to build useful abstracts around.
I'm going to tackle the problem this weekend, probably just use Lightroom mcp as the example, I don't have any good ideas to begin with;
- These applications should probably adapt their codebase to the evolving landscape (that might take a while so in the interim...)
- Another easy idea, is to boot up a sandbox and runs the software, maybe even shares projects across mcp users or something, a service orientated model but pretty much sucks too
- Best but kind of worst idea I have so far is to just make a software service that users download and run that orchestrates software and processes etc (kind of like anti cheating software or something, with far too elevated permissions)
A bit stressed for time so couldn't distill what I think properly just yet, will edit later)
Useful for service providers who want to expose themselves to technical consumers without having to write custom sdk's that consume their restful/graphql endpoints.
The best implementation of MCP is when you won't even hear about it.
I definitely agree that it is currently pretty shit and unnecessary for agentic coding, cli's or some other solutions will come along. (the premise being the same though, searchable/discoverable and executable tools in your agentic harness is likely going to be a very good thing instead of having to document in claude.md which os and cli specific commands it should run (even though this seems far more powerful and sensible at this point in time))
Often these days I vibe code a feedback loop for each project, a way to validate itself as OP said. This adds time to how long Claude takes to complete giving me time to switch context for another active project.
I also use light mode which might help others... jks
I understand we are all in different camps for a multitude of reasons;
- The jouissance of rote coding and abstraction
- The tree of knowledge specifically in programming, and which branches and nodes we each currently sit at in our understanding
- Technical paradigms that humans may have argued about have now shifted to obvious answers for agentic harnesses (think something like TDD, I for one barely used that as a style because I've mostly worked in startups building apps and found the cost of my labour not worth it, but agentic harnesse loops absolutely excel at it)
- The geography and size of the markets we work in
- The complexity of the subject matter / domain expertise
- The cost prohibitive nature of token based programming (not everyone can afford it, and the big fish seemingly have quite the advantage going fourth)
- Agentic coding has proven it can build UI's very easily, and depending on experience, it can build a very very many things easily. it excels in having feedback loops such as linting or simple javascript errors, which are observability problems in my opinion. Once it can do full stack observability (APM, system, network), it's ability to reason and correct problems on the fly for any complex system seems overly easy from my purvue.
- At the human nature level, some individuals prefer to think in 0's and 1's, some in words, some inbetween, and so on, what type of communication do agentic setups prefer?
With some of that above intuition that is easily up for debate, I've decided to lean 100% into agentic coding, I think it will be absolutely everywhere and obviously with humans in the loop but I don't think humans will need to review the pull requests. I am personally treating it as an existential threat to my career after having seen enough of what it's capable of. (with some imagination and a bit of a gambling spirit, as us mere mortals surely can't predict the future)
With my gambit, I'm not choosing to exit the tech scene and instead optimistically investing my mental prowess into figuring out where "humans in the loop" will be positioned. Currently I'm looking into CI level tooling, the known being code quality, and all the various forms of software testing paradigms. The emerging evals in my mind will keep evolving and beyond testing our ideas of model intelligence and chat bot responses will do a lot more.
---
A more practical rant: If you are building a recommendation engine for A and B, the engine could have X amount of modules that return a score which when all combined make up the final decision between A and B. Forgive me but let's just use dating as an example. A product manager would say we need a new module to calculate relevance between A and B based off their food preferences. An agentic harness can easily code that module and create the tests for it. The product manager could ask an LLM to make a list of 1000 reasons why two people might be suitable for dating. The agent could easily go away and code and test all those modules and probably maintain technical consistency but drift from the companies philosophical business model. I am looking into building "semantic linting" for codebases, how can the agent maintain the code so it aligns with the company's business model. And if for whatever reason those 1000 modules need to be refactored, how can the agent maintain the code so it aligns with the company's business model. Essentially trying to make a feedback loop between the companies needs and the code itself. To stop the agent and the business from drifting in either directions, and allowing for automatic feedback loops for the agent to fix them. In short, I think there will be new tools invented that us human's will be mastering as to Karpathy's point.
I currently use ai sdk by the vercel team heavily, thought it would be nice to have a registry for all tools as a matter of convenience with a few added goodies. Currently, it is simply an npm mirror where package authors can set a top level property in their package.json called "tpmjs" which flags it to be picked up by tpmjs. Beyond just simply being a directory, I've also added an executor playground which is currently free to use (and open source) which makes it very simple to just add the tpmjs search and execute tool to your agent's which allows them to take advantage of all listed tools without installing any of them locally in your own codebase. (sounds horrible I know but if you play around with the idea enough it could actually make sense if you hosted everything yourself)
if you play around with https://playground.tpmjs.com the thing to note is, is that app does not have any of those tools compiled in, it is using the external executor at runtime. It fetches via esm.sh and runs it inside a sandboxed deno which returns the results to your agent.
currently will be reaching out to tool authors to get their feedback on adding the "tpmjs" property to their package.json
all feedback is welcome