This sounds like arguing you can use these models to beat a game of whack-a-mole if you just know all the unknown unknowns and prompt it correctly about them.
This is an assertion that is impossible to prove or disprove.
I rarely have blocks of "flow time" to do focused work. With LLMs I can keep progressing in parallel and then when I get to the block of time where I can actually dive deep it's review and guidance again - focus on high impact stuff instead of the noise.
I don't think I'm any faster with this than my theoretical speed (LLMs spend a lot of time rebuilding context between steps, I have a feeling current level of agents is terrible at maintaining context for larger tasks, and also I'm guessing the model context length is white a lie - they might support working with 100k tokens but agents keep reloading stuff to context because old stuff is ignored).
In practice I can get more done because I can get into the flow and back onto the task a lot faster. Will see how this pans out long term, but in current role I don't think there are alternatives, my performance would be shit otherwise.
No, but they can take "notes" and can load those notes into context. That does work, but is of course not so easy as it is with humans.
It is all about cleaning up and maintaining a tidy context.
On the other hand Claudes tenacity, stamina and sustained speed is superhuman. The more capable models become the more valuable this is.
This is a joke right? There are complex systems that exist today that are built exclusively via AI. Is that not obvious?
The existence of such complex systems IS proof. I don't understand how people walk around claiming there's no proof? Really?
It is impossible to prove or disprove because if everything DOES NOT work fine you can always say that the prompts were bad, the agent was not configured correctly, the model was old, etc. And if it DOES work, then all of the previous was done correctly, but without any decent definition of what correct means.
If a program works, it means it's correct. If we know it's correct, it means we have a definition of what correct means otherwise how can we classify anything as "correct" or "incorrect". Then we can look at the prompts and see what was done in those prompts and those would be a "correct" way of prompting the LLM.
The fundamental fallacy you are exhibiting here is similar to saying that rolling a six sided die and getting a “6” means that you will always get a 6 any time you roll it. And that if you get a 6 and wanted a 6, you must have therefore rolled those dice “correctly” and had you not gotten a 6 that would have meant you rolled them “wrong.”
You know that is not true.
I don't know the exact internals of a car. But I can infer my car works by driving it.
>The fundamental fallacy you are exhibiting here is similar to saying that rolling a six sided die and getting a “6” means that you will always get a 6 any time you roll it. And that if you get a 6 and wanted a 6, you must have therefore rolled those dice “correctly” and had you not gotten a 6 that would have meant you rolled them “wrong.”
Bro we rolled that dice MULTIPLE times. It's not a one time thing. And the "rolling" of the die is done with a CHAIN of MULTIPLE qureries strung together. This is not one roll. It's multitudes of data points. Yes results can be inconsistent from a technical standpoint, but the general result converges on a singular trend.
We know that much is true: a statistic and that is at most all we can say about reality as we know it as science formalized can only give a statistic as an answer.
No, you can't infer that it "works." Only that it CAN work. The car may be poisoning you with carbon monoxide. Your rear brakes may have become disconnected (happened to me). The antilock braking system may have a faulty sensor that only fails at very low speed, leading to them engaging when making a normal stop, but also preventing the mechanic from seeing the problem, because he didn't listen to your bug report and instead tried to repro the effect with high speed panic stops (also happened to me).
If I use a product and have a good experience, I can conclude that SOMETHING must be going well, but not that EVERYTHING is going well.
This is reasoning about evidence 101.
This is called pedantitic reasoning. You look like a drowning person trying to stay afloat.
lets say i accept you and you alone have the deep majiks required to use this tool correctly, when major platform devs could not so far, what makes this tool useful? Billions of dollars and environment ruining levels of worth it?
I'd say the only real use for these tools to date has been mass surveillance, and sometimes semi useful boilerplate.
Then on Sunday I woke up and had claude bang out a series of half a dozen projects each using this GUI library. First, a script that simply offers to loop a video when the end is reached. Updated several of my old scripts that just print text without any graphical formatting. Then more adventurous, a playlist visualizer with support for drag to reorder. Another that gives a nice little control overlay for TTS reading normal media subtitles. Another that let's people select clips from whatever they're watching, reorder them and write out an edit decision list, maybe I'll turn this one into a complete NLE today when I get home from work.
Reading every line of code? Why? The shit works, if I notice a bug I go back to claude and demand a "thoughtful and well reasoned" fix, without even caring what the fix will be so long as it works.
The concepts and building blocks used for all of this is shit I've learned myself the hard way, but to do it all myself would take weeks and I would certainly take many shortcuts, like certainly skipping animations and only implementing the bare minimum. The reason I could make that stuff work fast is because I already broadly knew the problem space, I've probably read the mpv manpage a thousand times before, so when the agent says its going to bind to shift+wheel for horizonal scrolling, I can tell it no, mpv has WHEEL_LEFT and RIGHT, use those. I can tell it to pump its brakes and stop planning to load a PNG overlay, because mpv will only load raw pixel data that way. I can tell it that dragging UI elements without simultaneously dragging the whole window certainly must be possible, because the first party OSC supports it so it should go read that mess of code and figure it out, which it dutifully does. If you know the problem space, you can get a whole lot done very fast, in a way that demonstrably works. Does it have bugs? I'd eat a hat if it doesn't. They'll get fixed if/when I find them. I'm not worried about it. Reading every line of code is for people writing airliner autopilots, not cheeky little desktop programs.
> blows up in damn well near every professionals face
If you know how to use the tools, and know their limitations, you can generate vast quantities of useful code very quickly. If you can't manage this, its PEBKAC. You're saying you don't trust these to make more than minor changes, which might make sense if your code could kill people buy otherwise you're being overcautious or severely underestimating what these can do.
So, in professional environments, full auto is negligent to a point i hope it becomes a fireable offense. Like trusting lane-assist and adaptive-cruise in a car to handle full auto driving. It might even seem like it can, until the leading car disappears, until it hits a t intersection. You get me? Modern llms are lane assist and adaptive cruise, not full self driving. It frees up some of your headspace and attention but not all, in fact not even most of it.
It doesn't, that's ego-preserving cope. Saying that this stuff doesn't work for "damn well near every professional" because it doesn't work for you is like a thief saying "Everybody else steals, why are you picking on me"? It's not true, it's something you believe to protect your own self-image.
point me towards something complex which llms have contributed towards significantly without massive oversight where they didnt fuck things up. I'll eat my words happily, with just a single example.
I think it's fair to say that you can get a long way with Claude very quickly if you're an individual or part of a very small team working on a greenfield project. Certainly at project sizes up to around 100k lines of code, it's pretty great.
But I've been working startups off and on since 2024.
My last "big" job was with a company that had a codebase well into the millions of lines of code. And whilst I keep in contact with a bunch of the team there, and I know they do use Claude and other similar tools, I don't get the vibe it's having quite the same impact. And these are very talented engineers, so I don't think it's a skill either.
I think it's entirely possible that Claude is a great tool for bootstrapping and/or for solo devs or very small teams, but becomes considerably less effective when scaled across very large codebases, multiple teams, etc.
For me, on that last point, the jury is out. Hopefully the company I'm working with now grows to a point where that becomes a problem I need to worry about but, in the meantime, Claude is doing great for us.
> The vibes are not enough. Define what correct means. Then measure.
And how do you define correct feedback? If the output is correct?
1. Agent context with platform/system idiosyncrasies, how to access tools, this is actually kept pretty minimal - and a line directing it to the plan document.
2. A plan document on how to make changes to the repo and work that needs to be done. This is a living document pruned by the orchestrating agent. Included in this document is a directive written by you to use, update the document after ever run. Here also is a guide on benchmarking, regression, unit tests that need to be performed every time.
2a. When an agent has a code change it is then analyzed by a council of subagents, each focused on a different area, some examples, security, maintainability, system architect, business domain expert. I encourage these to be adversarial "red team". We sit in the core loop until the code changes pass through the council.
2b. Additional subagents to create documentation, build architecture diagrams etc.
2c. A suggested workflow is created on how to independently invoke testing, and subagent, etc.
LLMs are proving to be very much force multipliers of the kind of developer you already are, and of those who report a 10x increase in productivity they're probably all being genuine. Whether that 10x is of careful, thoughtful choices or reckless rough-shod slop though is really an artifact of the developers themselves. I've been saying from the beginning that your effectiveness with LLMs is roughly equivalent to your ability to get effective results out of a real team of human contractors.
It's really amazing, we've crossed a threshold, and I don't know what that means for our jobs.
> Another AI agent. This one is awesome, though, and very secure.
it isn't secure. It took me less than three minutes to find a vulnerability. Start engaging with your own code, it isn't as good as you think it is.
edit: i had kimi "red team" it out of curiosity, it found the main critical vulnerability i did and several others
Severity - Count - Categories
Critical - 2 - SQL Injection, Path Traversal
High - 4 - SSRF, Auth Bypass, Privilege Escalation, Secret Exposure
Medium - 3 - DoS, Information Disclosure, Injection
You need to sit down and really think about what people who do know what they're doing are saying. You're going to get yourself into deep trouble with this. I'm not a security specialist, i take a recreational interest in security, and llm's are by no means expert. A human with skill and intent would, i would gamble, be able fuck your shit up in a major way.
How do you know these are actual vulnerabilities? You just ran an LLM and it told you something and you came back to dunk on me, with zero context on the project.
Maybe you need to sit down and really think that you have no idea who you're talking to or what the project does. Next time you make a "omg this code is so shit" comment, include something more than "well my LLM says your LLM is bad" so we can have a discussion with facts rather than LLM-aided trashtalk.
EDIT: Out of curiosity, I've ran Kimi K2.5 on the codebase, and all the things it found are invalid, or explicit design decisions. So, next time you decide to tell someone their project "is slop" by running an LLM and relaying its verdict, consider a) the irony of what you're doing, and b) that the other person might know more than you about their own project that you spent "three minutes" running an LLM on.
You should educate yourself more before you go around slandering people.
This can be done via a signal/telegram message, via compromised plug-ins, via file uploads etc. At no stage does there appear to be, as far as I can tell, any attempt to mitigate or limit attack surface. That's just one of the problems. If that's by design, my bad, you do you king.
How is the attacker going to message the bot, genius? Do riddle me that.
You're hyper-focused on the front door, asking how an attacker would even message the bot, but you're ignoring the fact that modern attackers don't bother knocking. Read cloudflare's 2026 threat report, it's eye opening. Between automated session cloning and browser-based info-stealers (among many other modern headaches), the 'whitelisted user' is no longer a static, trusted entity. If a user on your list has their session token scraped via a malicious browser extension or a hijacked desktop app, the attacker effectively becomes that user. At that point, your bot doesn't see an intruder; it sees a 'trusted' account and hands them a loaded gun in the form of arbitrary SQL execution. Now, the problem is you aren't the only one with access to LLM's and obscurity never really was security, even less so now. An llm with credentials could easily probe its way through your bot's capabilities and connected data and exfiltrate everything.
So, the reason I'm calling into question your claims of having a 'safer' personal agent/bot/whatever is a matter of blast radius. A standard bot usually interacts with a restricted API or a set of hard-coded functions, so even if the account is compromised, the damage is capped. By giving an LLM the keys to the entire database, you've created a single point of failure that can result in total data exfiltration or a complete 'drop table' wipe, among any number of other nasty things. That's just _one_ issue in this project.
If you actually want this to live up to the 'safer than average' description, you have to move past the idea that a whitelist is a firewall. You need to distinguish between authentication and authorisation and implement defense-in-depth, starting with a database user that has zero permissions beyond simple 'Select' queries. You should be using a proxy that intercepts the LLM's generated SQL and kills any string containing 'Drop', 'Update', or 'Delete' before it ever touches your server, without some form of parsing/checking. Right now, you’ve built a powerful engine with no brakes, and telling people it’s safer just because it's on Signal is a dangerous misunderstanding of how modern exploits actually work.
Alternatively, fix how the project is described to be more accurate/honest than it is now.