Of course, there are other areas that can improve model output (Direction rather than open-ended assistance requests, using keywords + plugins that help, the "your output should include: " style prompting).
A few of us run almost the same exact setup at my shop (Base Claude Code w/ SuperPowers + a context repository) and the models are somewhat unhelpful to some, and give meaningful output to others. The only correlation I notice is that their prompts are no-good. Not from a meta "prompt" engineering standpoint, but from a general English 101 standpoint.
"dudde no i wanted the function to return 3 things. not like that. do it again"
VS something like
"Modify the "renderThreeVars()" function signature to accept another variable called "z" and add it to the return statement at line 64."
Obvious exaggeration, but you get the point.
I ask it all the time about whether X is feasible, how we can get started on Y, and to investigate issue Z.
It is working great for me in a >100k LOC project.
Perhaps this works less well with weaker models. I suspect the people who say Qwen 3.6 27B is working well, are using prompts like "modify the renderThreeVars() function in rendering.py".
The irony to me is that a lot of what I'd tell people about this is exactly what I would have told them about writing Stack Overflow questions.
For example, if I tell it that I want my app translated, it can plan for me what the recommended options are in my framework, what languages I should target for my app, and come up with a skill for a repeatable workflow.
I've noticed among my coworkers that we all have different amounts of trust we're willing to give the agent. That seems to manifest into some people only asking questions about the existing code, but never writing anything new with it. Others are willing to do limited targeted changes with the agent, but are unwilling to do things like let it make commits, or connect MCP servers, or really do anything that isnt fully understood by the human before setting the agent loose. Then I find myself, who has dove headlong into it all. I have skills that use MCP servers to check for pull requests, and give me summaries to give me more context for code reviews. I update Jira tickets in batches of 50+. I develop complete features exclusively through prompting the agent to do everything. I know I still have the responsibility to understand it at the same level as if I wrote it my self, and defend it and debug it.
I can easily see that superpowers would be wasted on most of my coworkers, simply because the benefits compound with the complexity of my ask. My coworkers aren't willing to hand off enough control to receive the benefits of superpowers.
Been working on some things that require more targeted, smaller scale changes and the base models do perfectly fine when given good instructions.
Superpowers and Compound Engineering are the two that I hear about around our team, both seem "fine" if you're into the fully agentic engineering "big change" future technology stuff.
From some very quick tests, I get the impression that having good grammar, punctuation etc. is really not important (although I try to do it anyway because I'm accustomed to trying to do it), but clarity and precision definitely are. And of course, if you have unknown unknowns, they do need to get figured out before progress is possible. But with the right mindset, that just means you need an extra turn or two, not that you're going to end up at a dead end. (... I guess that counts as a pun?)
The prompts were, predictably, really bad. Broken English, sentence fragments, vague requests, lack of context. Yet somehow, the users always got the answer they were looking for. It might have taken a few extra turns with questions from the model, but the end result was the same.
It's humbling, but a flowery, carefully crafted prompt is at best slightly more efficient than a "CAN A DOG BE EATIN SUN FLOWER SEED?" peasant prompt.
Unless the user wanted to know if a cat could eat sunflower seed or something.
ChatGPT kept the charade for all of one sentence. Then it dropped to talking about "your dog" the rest of the way. It even starts the final paragraph with "if, instead, you mean you (a human) ate them...", and finishes with the question "is this about an actual dog or yourself?"
Google did consistently refer to me as a dog, but its entire focus was on the steps "your human" should take, no advice for the dog itself.
In both cases, it looks like the first person roleplay was entirely inconsequential for the usefulness of the output. I think the current gen AIs have outgrown this trick and you can safely forget about it.
Why is would that be "bad prompt"? It is machine inpit, if machine can interpret fragment all the better.
They don't need to be told that you're asking about the safety of eating them, because they can infer that based on the fact that a very large percentage of any text linking dogs to sunflower seeds is obviously going to be about the safety of the dog eating them.
Even pre-LLM that would have been a perfectly sufficient google search for the same information.
Sure, but I think it's essentially just that people who are better at traditional non-agentic software engineering are better at agentic software engineering. The only exception would be individuals who avoid agentic coding due to skepticism, hostility, or lack of opportunity.
Are people really putting on their resumes that they are capable of reading and writing and appropriately defining and limiting context? That's all prompt engineering is - it's being able to communicate effectively and elucidate your objectives.
Congratulations to all you English majors out there, you're about to make $350K/year.
People put whatever buzzwords will get them through initial screening. My resume contains tons of banal shit like agile, automated testing, Linux, AI (since 2018), and design patterns.
Yes, it is. Source: had models inplement complex things from scratch and bullshit regardless of whether it was a one-line prompt or a detailed "SOTA witchcraft magic spells that are guaranteed to work"
It's like imagining you could "prompt" Richard Feynman to be smarter at Physics.
That is, for 99.9% of engineers, if you want the model to do a code review of your project, the best solution is to just ask Fable, "Hey Fable, do a code review of this project." Throwing in extra text like "think like a senior engineer", "ensure you focus on DRY principles, KISS, self documenting code, etc", doesn't make a difference.
These sorts of tricks used to work with dumber models, but now, like I said before, it's like thinking you can prompt Linus Torvalds into writing better C than he already can do.
> it's like thinking you can prompt Linus Torvalds into writing better C++ than he already can do.
Linus Torvalds, the inventor (and beloved dictator) of Linux, has always been quite harsh about C++ and why he rejects it for Linux kernel development. He’s not just been very vocal about it, but also brought up some arguments against the use of C++ that are worth reviewing in detail.
~ https://medium.com/@jankammerath/linus-torvalds-critique-of-...The models are beyond expert level in many areas at this point.
Do you really believe that adding extra junk to your prompt is going to make the model write code better than it does already?
Again, imagine going to Terrence Tao and "prompting" him to get better at Maths, do you think you can do it? What prompt would you give to him to make him produce better maths. Unless you're already a world-leading Mathematician I think you would find it hard.
It's not that the models aren't smart or whatever; we know they're extremely capable.
There's always going to be value in being able to clearly communicate what you want the model to do, especially when the context is lacking.
Extra prompt text isn’t to make the model smarter, it’s to pull attention towards what you want. “Do code review” is far different from “Make sure changes align with existing architecture” or “follow these enterprise standards XYZ”. The model just acts as an average of its training data, which may or may not align with your goals.