676 karma · joined September 20, 2023
Though I agree with OP's point of view that MCP was overhyped for too many use cases. Complex agent-system integration problems are usually better solved with just AI agent + curl + SKILL.md. It's just way more flexible.
It's another variant of the 'fat client/thin server vs thin client/fat server' debate. Some people want rigid, thin (e.g. web-based) frontends with the LLM doing work in secret behind the scenes. Others want fat, versatile frontends through which the LLM can interact with the user's own environment.
I've always been a fat client guy and this time is no exception. I doubt the constrained approach is going to lead ground-breaking innovation. I also wish companies would treat SKILL.md + curl as the main mechanism for agent tool calling as opposed to MCP. MCP is niche.
The people who prided themselves on raw puzzle-solving ability got hit hard. Those who are idea-driven and architecture-oriented feel like they got handed a superpower.
It's quite a big shock at the industry level because the entire software engineering job interview process at essentially all large companies was heavily biased towards well-defined, raw problem-solving under time constraints which is precisely the skill which AI has replaced.
Now AI has surpassed me in fine-grained problem-solving ability. It can write more reliable code than I can, faster than I can, provided it is given the right guidance.
But one thing that I'm still much better at is identifying technical opportunities and choosing the right tradeoffs.
My feeling is that frontier models are incredibly smart in some ways, but incredibly dumb in other ways
When I chat with Claude using deep technical language about distributed systems issues, the arguments it presents are mind-blowingly good. I worked with many skilled engineers on complex projects but the kinds of arguments Claude makes are on another level. It feels like it can read my mind because it already identified all of the relevant aspects to the current topic and it's incredibly persuasive. I would say it's even better at deep, nuanced technical discussions than I am.
When I ask it to implement a feature using technical language, it does a really good job and the solution usually works out of the box. No bugs at all 95% of the time even after adding 1000 lines!
But the problem it still has is that when it implements solutions, it misses so many low-hanging fruits/opportunities. Especially in terms of performance, maintainability, scalability and UX. It's like it sees and recalls all the relevant parts perfectly but yet somehow misses opportunities which seem extremely obvious to me.
It feels like AI has 0 creativity. It only sees the opportunity once I mention it... And once it sees the opportunity, it demonstrates deep understanding of the technical implications. I think what's surprising is that it understands the suggested solution so well, with such nuance, that I can't understand how it didn't see the opportunity and why it never seems to see it until I mention it.
This is very unhuman-like. There is no way that a human being with that degree of understanding of a topic would be presented with such highly relevant context and not make the connection.
That said, I also think everyone who contributed content online deserves to be rewarded for their contribution, especially since OpenAI (which made the largest contribution) was a non-profit in the beginning. So there should be some kind of tax. Otherwise I agree, it's not ethical. It's not 'fair use' of copyrights.
It's going to destroy a lot of startups which were monetizing this exact idea. But clearly it's a low-hanging fruit so it makes sense that OpenAI would do it.
Seems we're on the same wavelength because I had the exact same thought about this.
But actually, I'm probably not so fussy. I wouldn't mind cleaning up the mess if I'm paid hourly... But now my concern is: What if they want me to clean up the mess but demand that I do it in a particular way which makes it impossible?
I'm already seeing signs of this. This was already kind of the case in my last job; I had to fix things and implement new features but we had to keep the same clunky, over-engineered architecture. My current job doesn't have the same degree of architectural legacy baggage but there is a lot of bureaucracy to deal with instead, which is itself restrictive.
Just watch the software engineers who say "Let's be careful" lose their jobs and get replaced by non-technical vibe-coders churning out 10K lines of dirty insecure, unmaintainable code per day... Ticking time bomb.
I swear, good engineers are going to move to North Korea to monetize because of the amount and size of 'opportunities' this will create.
It seems so far-fetched but that's the direction it seems to be heading.
There are project I've built from scratch that I would feel confident to hand off to a bunch of non-technical vibe coders and I know they would be productive and the product would likely be secure; because the existing codebase already exhibits all the patterns and principles that are required for that kind of project.
It would probably slowly degrade over time if a lot of vibe-coded logic is added on top but I think they could get very far feature-wise whilst keeping the software reliable.
But even though the value of such codebase has increased, people haven't adapted to this new reality. People are generally not good at telling what is good code. Because we don't actually have consensus on a definition. My definition is that good code is code that is easy to extend and maintain.
If implementing a feature requires a huge amount of tokens, then there's a good chance the codebase is not great.
I've worked on a codebase where a small feature requires might require 3k tokens, but on a different codebase, a feature of similar complexity would require 30k tokens minimum... And it's not about the size of the project; it's more about how the logic is divided and the architecture. And importantly; it's not a one-off; it's a clear observable, repeatable pattern.
Anyway databases nowadays are a commodity. A sticky commodity but nonetheless they are replaceable; increasingly so in the age of AI where data migrations are easier than ever.
Even if each user gets their own sandbox, they will still want to configure different access rules for different kinds of data which they host.
That said the idea that each user could control and host their own data is interesting and could work. I imagine you could have apps which link data from many different user sandboxes via remote foreign keys.
You could have a centralized data schema controlled by the application owner but the data itself would be held/scattered across a large number of sandboxes.
I got interested in this specific idea back in the early days of Firebase I saw someone built a realtime HTML component with PolymerJS called 'collection' and I became consumed by the idea of fully generic realtime self-updating components. My approach is a bit different than OP or that of HTMX though; it's JSON over the wire, not HTML.
I've built a full implementation in Node.js with a set of declarative frontend components.
https://github.com/Saasufy/saasufy-components?tab=readme-ov-...
I'm thinking to make open source.
It's nice to see major frameworks coming to a similar conclusion.
I think partly it's because the training set contains a lot of JS, but also because complex software written in JavaScript must have impeccable architecture in order to exist at all.
It's rare to encounter a complex, functioning JavaScript application with bad architecture. I've never met any engineer smart enough to maintain a large spaghetti-code JavaScript project.
On the other hand, I've seen horrible TypeScript projects. If it wasn't for the helpful type annotations, no human being would have been able to maintain it.
But we can't drop all abstractions. It's not physically possible. The AI agent is going to break logic up into files... The AI will create/impose an abstraction for each piece of logic whether we like it or not.
So saying "drop abstractions" is actually saying "let the AI decide what the abstractions should be" and unfortunately because LLMs are trained on average code, those abstractions are often pretty poor.
Coming up with the right abstractions is actually one of the main skills which AI hasn't automated and where the human brings the most value. It affects maintainability directly; including when using AI.
I've worked on projects with poor abstractions and ones with good abstractions using AI. The ones with poor abstractions require 10x to 100x more tokens and time to solve problems and implement new features. You're constantly fighting it to prevent it from coming up with nasty hacks and workarounds. Poor abstractions create the need for hacks and workarounds.
A good abstraction is a single edged sword which simplifies the task. A mediocre abstraction is a double-edged sword. A bad abstraction is like a single edged sword with a restrictive handle and the sharp edge is facing towards you.
For vibe coders.
With a focus on:
- Security
- Analytics
- Advanced data sharing scenarios
- Realtime updates
Claude can be given full access to your control panel and can impersonate any role to exhaustively test all your data models and views. So even a fool can exhaustively prove the security of their application for all data in their system.
The article stops exactly where I did and doesn't go further into SPHINCS+ which builds on top of the same primitives to provide statelessness.
And the reason I stopped at that was probably the same. There is already a fair amount of complexity involved. I wanted a signature scheme which would be relatively simple to understand and implement. Also, there were no good SPHINCS implementations at the time for my engine/language and I didn't feel confident to implement from scratch.
Also, no project at the time (except I think IOTA) had stateful signatures and I liked the idea of being able to change passphrases as it could potentially allow people to sell their wallets (along with associated DEX memberships or delegate spots) to other people.
At the time, I was forging as a delegate on a different DPoS blockchain and I was thinking that it would be nice if I could sell my wallet (which would have been worth 5x to 10x my annual block earnings)... Unfortunately, I couldn't do that (no mechanism for it) and I ended up losing that income stream, never having the option to cash out.
Firstly, people don't agree on what good code is (developers who get paid by the hour tend to favor complexity so they think good code is complex code), secondly, AI companies have an incentive to produce more tokens; this works against good architecture since good architecture would cost fewer tokens to maintain. Thirdly, many decision makers tend to conflate complexity with intelligence, hence they are more likely to prefer complex solutions over optimal ones.
If I'm knowledgeable in a niche over which there is no mainstream consensus or it's disorganized, and I want to reason more about this topic, I need the AI in the rabbit hole with me. I don't need it to question the existence of the rabbit hole I'm in. Otherwise all of its arguments sound like gaslighting; asking me to disbelieve my eyes and decades of experience about basic ideas which I've already researched many times and proven not to work (at least with high confidence).
For me this happens with some niches of software engineering and architecture and I have lots of physical proof of working, secure, maintainable software to back up my position.
Also, when I say that "most of my contrarian ideas are correct" it's with the caveat that I don't tend to adopt or hold onto contrarian ideas that are easily proven to be incorrect. I'm referring to the kinds of niche ideas which most non-experts would not feel confident to make strong arguments about but AI will.
All this doesn't necessarily translate to sales though. Because, to evaluate those things requires attention, trust, time and effort.
This is a significant barrier because a lot of software will appear to meet all of these superficially.
After 1 day of usage, a piece of software may appear to be intuitive, reliable, secure, performant and useful... But then after some time (sometimes a whole week or longer) you run into a critical scenario and discover that it cannot be solved with that software... Or performance drops off sharply after you created the 1000th record in the software... Or a hacker takes months to find that one endpoint which allows full remote code execution.
Even in an optimistic scenario, 1 week is a long time to evaluate a piece of software. I've encountered software which took 6 months and large teams of people to realize that it wasn't suitable. That's how long it took to hit the critical limits. It's a very long evaluation loop.
So you cannot judge new software efficiently by just looking at it. Even industry consensus is problematic if the software is very new and complex... Plenty of trends have fallen off a cliff in the past. You need to understand who is behind the software... And even that's not so easy; social proof can be misleading when it comes to deep technical ideas. The people who are good at social networking aren't necessarily good with tech.
I think AI slop code is going to be a much bigger problem than people anticipate. And the irony of it is that the solutions already exist... What doesn't exist is the mechanism to identify those solutions. But even the mindset needs to be corrected first.
So it's a waste of my time and tokens when the AI keeps making weak arguments which I proceed to destroy one by one.
Usually, what happens is it tries to debunk my statement but then I offer a rebuttal to all of its points, it then concedes to all of my points but each time it offers a new partial rebuttal... Which I proceed to crush... But it keeps coming up with increasingly irrelevant caveats. It goes on for quite some time and at the end I ask it to review our discussion and it admits that its original stance "was not as strong" as it initially claimed but it never fully concedes... Even though I literally destroyed all of its points and every partial rebuttal it tried to come up with... Meanwhile its rebuttals became increasingly nit-picky and distant from the original claims made...
It's like if I'm saying "the ship is sinking, look at all the water in the hull and look at all the water pouring in through that hole" and it's like "oh but this is a small hole and the pump can easily offset it" so then I say "What about this hole over here" "Oh but the pump can still offset the stream from both holes easily" and then I say "What about that third one? And fourth one? And fifth one?" And it's like "The pump can easily offset that" and I'm like "Common! The hull is full of water and the water level is rising, the hull is full of holes; it's not a stretch to suggest that the ship is sinking because of all the holes in it... I shouldn't have to point out the location of every single hole for my argument to start making sense!"
It's not proof by induction but it's probably as close as you can get to it for a topic which lies outside the realm of mathematics!
I don't even debug anymore on those projects. If Claude tries to add debugging logic in my code, I tell it not to and just provide additional information and it can usually find the solution faster that way.
This is when working on my own projects. When working on projects created by other people, it's a different story and I have to fight it constantly to stop it from implementing hacks and workarounds... It uses much more tokens to implement basic features. It's more work for both the AI agent and myself.
The project's existing code makes up most of the context so if the code is not great, you have to write long detailed prompts to set it on the right path. You have to make it clear that the existing code isn't good enough and your expectation is higher.
In this case, it usually gets better with more back-and-forth... At the beginning, it can't do anything because you keep pointing out a problem whenever it tries anything at all, but eventually, after a lot of criticism, it starts becoming more careful and adapting to your standards.
So yeah, even same person doing the prompting can lead to two very different experiences depending on who built the foundation.
So my conclusion is that the expertise comes from both the existing codebase and from the person doing the prompting... And TBH, I would say the codebase/foundation carries more weight than the person doing the prompting.
Pretty sure I could put an idiot on one of my codebases with Claude Code and they'd do a decent job.
Multi-pass rendering with multi-pass template placeholder substitution is very useful because you can have an entire complex component hierarchy declared in one place, very succinctly, in one file... Different components in the hierarchy can render data at any depth below them in the hierarchy (any level of nesting) without child components having to know about it. The child components just take the partially pre-filled template as their input and fill in remaining placeholders. The logic/markup is much more succinct and the components are much more flexible and composable. It promotes better separation of concerns.