555 karma · joined April 21, 2010
Building https://www.neovantik.lu
This is impressive as it is optimizing the effort on the low, but not too low hanging fruit.
LLMs compress the known shapes well and fit them to solve for a problem but they are still not good enough to build from scratch a large new lego piece which fits a full problem perfectly. And such a large lego piece might not be the most efficient solution either and might be difficult to prove so.
400 Bad Request is fine for a client developer who reads it once at design time and fixes the code forever. it's dead weight for an agent that must self-correct from the string alone. "Expected ISO-8601, got 03/04/2025" is now a functional part of the interface.
Not all agents are developers who can code their own interface.
In my experience, I find it to be exceptionally good at exploring and fixing the edge cases. Of course the output is not human maintainable for these fixes and needs to be heavily tests controlled, refactored or just accepted as being agent-maintained going forward.
Simon needs to resist the pelicans(and the django mindset) and Garry needs a new loop which can loop on itself without any human trigger so that the agents can "dream" better. Who knew that it was not just the models which could hallucinate.
It has 119 repositories.
Is this how AI slop looks like in code? Made for the agents, by the agents? Is this separation of concerns or context management with agents as a first class residents and humans merely acting as custodians?
Similarly lawyers/bankers were the ones who built in trust in capital, contracts, businesses and protection of investor rights. Delaware c corp is not an outcome of bad guys.
Finance, Law, VC guys were good too in the beginning but when the value/status change happens it attracts certain kind of guys who are average in talent but excel in demonstrating value and social management of the value/status.
Another change which has happened recently is that the economics of engagement farming have become common place wisdom as already proven effective for everything from selling books, personal brand, career skill/virtue signalling, staying relevant.
Due to this everyone is talking more without restraint and not keeping in their own lane of earned expertise.
Looking through wages and trying to find a ceiling(by time/effort) on the value creation by a human is one dimensional at best.
Claude is a better reader. I have to just tell it to read the docs/specs sometimes.
So either something is computing it or some exploration is happening at quantum level and we just see the final result.
When a new inference has to be done the query(q) is projected in the manifold space. This projection is dropped on the manifold and the gravity of the manifold gives an answer of q+1 length. Which(qw+i) is dropped qw+n times to output a final response of n length.
The gravity is created by repeated multiplication(of the weights/input) to find out how the projected embeddings should fall according to the manifold in the GPU.
Most of the conversational skill and perceived intelligence of these models in hidden in RL/system prompts.
Anthropic's target should be a codebase designed for agentic comprehension from the first commit. Here the codebase adapts to the agent. You can enforce conventions, structured metadata, semantic indexing, explicit dependency graphs. Whatever makes the agent's job trivial rather than heroic.
For e.g. ask any model "which class of problems and domains do you have a high error rate in?".
response: https://chatgpt.com/backend-api/estuary/content?id=file_0000...
result: FAIL
The ones who are frustrated are the ones who were interested in doing(whether good or bad) but are being told by everyone that it is not worth it do it anymore.