If you spend $8000 to generate an animated pelican riding a bike, then how much tracking does it really need?
Is the guy who spent $300,000 or so translating the FLT proof to Lean going to get a big Christmas bonus?
Rest assured, capitalist appears irrational in wasting money, but they certainly care more about profit.
I tried something similar and I remember it was still pretty dodgy in February.
If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review. If I had $100k to spend next month I could probably get through it, I'm running $2500+-api-equivalent a week at this point and I feel very token limited. Will be time for a 2nd or 3rd subscription soon for both labs I think.
Fable was a revolution, still learning how best to use it, 5.1 felt like a notable upgrade. At this point I launch a workflow with 10-20 minutes of interactive setup (and even that I feel might be too much), it runs for hours, and the PR is trivially mergeable (I still review every line, but 95% are just merge, maybe 4% are feedback needed, 1% are thrown away and regenerated, which implies I'm being insufficiently ambitious)
> If I had that many tokens/dollars I would be running canaries and adversarial verification in prod based on e.g. traffic replay, live fuzzing, all kinds of things to build confidence without direct human line-by-line review.
This part jumped out at me. There's something to watch out for here.
I recently had a funny experience. I delegated a major feature to an agent.
It turned out that it had implemented it precisely backwards, in a way which was pointless and which made things worse.
But it had written countless tests for the feature and all the tests were green.
I realised in that moment that even formal verification would not have helped, because it would simply have written a mathematical proof of the correctness of the incorrect feature...
The other thing I do, not as much as I should, but it's very powerful, is to generate spikes and deliberately throw them away to understand how to prompt better. Like I generated a swift version of the react native app I'm working on, and Alloy provers for the state transitions. None of it is production quality but getting great results that way is useful to scope future work.
What is this about? Could you give an example?
So when you're creating a (and hopefully b), you use one model with one context, and then you might use some other model to implement c. I kind of round robin the models and present them with seperate context - so for example I don't build a, b, and c together, I build a plus some b, then later go for a pass over b and c together, then maybe I used c to drive improvements in a, and I vary between openai and anthropic models as I do.
Heres an example set of minimal prompts:
a-focus:
implement feature <x> based on #ticket in github, be sure to reference the engineering standards documents and the swift and react native skills as needed
b-focus:
improve test coverage in the repo for <subsystem z> to ensure that <feature x> is covered completely, and fix any outstanding gaps or omissions in that feature as you go (standard references above)
c-focus (possibly in a separate repo):
You are creating a black box xcuitest to drive a physical phone for testing <feature x>, here is the user specification and known issues, create failing tests for each known issue and an overall robust suite to ensure any user facing or ui issues are caught
How are you running jobs unattended 24/7 without hitting your token limits?
Over 24h my token spend is <30$. Excluding tokens for review it's <10$. With the absurdly gigantic subscription subsidies and a reasonable workflow I suspect one could run parallel agents.
I'm not sure what the point would be though unless working on some kind of optimization problem -- it takes me days to review <24h of the agent's output. It's almost always near enough to correct to be shippable; though I do give it feedback and iterate until it's better than the code I would have written.
I'm at a point in my career where a small minority of my time is coding. The AIs can do in a day what would have taken me a week uninterrupted with acceptable (in some cases inferior prior to human feedback--but in some cases superior!) quality.
As I do not have 10 let alone 40 hours per week to devote to coding I think it increases the amount of high quality work product I can create with a given time investment. As I review it I merge small independent units and decompose the work.
All that is to say I don't really like it - but I suspect for most *well defined* coding tasks human produced code from highly experienced engineers will largely cease to exist in the next year -- getting cheap/relatively horrible models to produce good code is now straightforward.
OTOH I never use AI for any human facing communication outside of making my writing shorter. IMO AI slop "documents" are almost certainly a drag on organizational productivity.
I'm still wary of any unreviewed code - though my area of work is not tolerant of defects.
Agree on targets / verifiable indications of progress or success being a prerequisite for this being useful - although that covers quite a lot of SWE work.
For a task I left a local model running on overnight, only ~100k tokens were used because most of the time was just waiting on tests to finish, then waking up, tweaking a few settings and trying again.
So an overnight loop like yours wakes you only when it needs a decision, instead of you waking on a timer to check.
But in this case it would be triple LLM inception, one training another and a third one monitoring everything is done correctly :p