> What you really mean is, the core team there doesn't want to lose control.
It's 100% this. They basically produce vague guidelines such that only the core maintainers are allowed to use LLMs, under the guise of "well of course we understand the code" and no one else is. It's also completely unenforceable, how are they going to prove whether someone understands the code or not? Even if they show sufficient evidence/understanding the maintainers can simply sabotage them and accuse them of using an LLM to explain the code. No one wins here.
> A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.
Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you're being super pedantic about it.
I also just ran _one_ of their files (with include deps) through Astra and it detected >100 vulnerabilities/correctness errors (with over 10 outright UB/memory corruption issues). It's actually outright shocking.
Is this written in jest? Because it's very likely where the future of computing is heading. See https://www.youtube.com/watch?v=kZRE7HIO3vk; a lot of people were nagging on Casey because he implied that software was more efficient back when everyone "wrote their own kernel" and how "impossible it would be today". He even mentions how awesome it could be if every game came with it's own bootable USB. Now back then it truly was unthinkable, but today we're edging ever closer to that reality.
For instance, I have a working microkernel written in a Lisp dialect for embedded devices. Compiled to native machine code. 100% LLM generated. ~70k loc. In benchmarks it outperforms most other embedded kernel projects by a significant margin. And it only took around ~$1500 in tokens (API costs all included).
The real defensible reason is that with current tools available you can write perfectly safe (safer than Rust even) C++ code, while avoiding the horrid Rust compile times and without needing to pepper your code with unsafe all over the place. LLM's can help you formally verify your code and extensively fuzz/test it to the point where you can actually be sure (ie prove) that the code is safe, without really relying on Rusts compiler. Lastly, C++ lends itself more naturally to hardcore optimizations than Rust. But ultimately it really, _really_ comes down to a) Rust's horrid compile times and b) Rust's horrid metaprogramming support (this even kills it for LLM generated output because it wastes tokens).
It's necessary due to context shuffling. You want to make sure that the plan is splatted into multiple files and then repicked up, analyzed and implemented. ie have the harness create PLAN_{FEATURE_A,FEATURE_B,TESTS,ARCHITECTURE}.md and then pick it up piecewise or in parallel. It's also required if you want to change things between runs, if you try to one shot it you can't easily interrupt an agent.
You don't need plan mode specifically for this, but it amounts to the same thing.
Something will this would historically guarantee a Senior Staff+ position at Nvidia. Wondering why Jensen doesn't put money where his mouth is ("were seeking exceptional engineers blabla") and offer him a job?
Yep. It's hitting limits because models have effectively reached the logical endpoint of how "precise" their outputs can become. In order to overcome this they need to either a) get excellent at outputting and working with massive codebases (10+ million loc) since a singificant % of systems simply cannot be done in less than that (the language itself isn't expressive enough) or b) reach the next stage of intelligence where the models are somehow able to output sequences that can solve vast complex problems with a very narrow output. Option b isn't really possible (various statistical/computational fundamental limitations forbid it) and option A is _insanely_ expensive. Which is why all of the demos the frontier labs have been pushing out are either really impressive one shot demos where the model is able to take a concise input and produce something great on its own, or very long running internal sessions where they claim to do something vast with minimal oversight (NavierStokes, C compiler, Bun rewrite).
Yea, the $500 package is likely a rebranded $200 from a few months ago. Keep in mind all providers have been cutting token allowances for their subs for a few months now
Decent(-ish), but doesn't provide good value for the money, sadly. They still refuse to tell you exactly how much usage you're paying for. My original suggestion a few months ago was to provide a $500/$1000 tier but make it _unlimited_, which this doesn't appear to be.
> Paying $200 a month and part of their Cyber Verification Program but can't use Opus 5.5 or Sonnet 5.5 for any authorized bounty work. Immediately get flagged for `Cyber`.
AI providers still haven't realized how much cash they could rake in if they provided fully unrestricted models.
Seriously? I have no idea where this cognitive dissonance comes from. Or are people just lying (outwards or to themselves)? A rewrite of this magnitude would easily take a skilled human team months if not years to finish. This is on top of Rust not being an easy language to work with. Which, btw, is the sole reason why not everything is written in C/C++/Rust.
Frankly, I don't really see how it isn't solved, even with the current state of LLMs. Frontier models can write, understand, correct, and optimize code in practically any language at a superhuman level. I haven't come across a single problem that LLMs can't solve. You can easily give them a research paper, ask them to implement it and in an hour or two it's done. Or even point them to a video or screenshot of something and say "implement this feature in our game engine" and... they just do it. It might not be optimally perfect, but what % of human written code is? Even if you ignore the time amortization (given how models can spit out weeks of human work in an hour) they still obliterate even an experienced developer.
"optimize this code", "fix this code", "extend this code", "add this feature", "find errors and patch them", "find bugs and fix them", "rewrite this from python to rust".
This is all that's needed to actually use LLMs nowadays. How is it a "multiplier" rather than an "equalizer"?
You shouldn't just go for the shares but also apply interest on it as well. They could (if my math is right) technically owe you 50-100k shares of then nvidia shares. Which would be a monstrous payday.
LLMs are exceptionally good at this, _especially_ if you point them to an existing codebase. I have no idea if FL Studio is OSS or not, but if it is, a Rust rewrite is trivial.
Just hard fork the project. Frankly, llama.cpp is so badly written that these kind of speedups are trivial, and a hard fork (or a total rewrite) has been needed for the longest time.
OpenAI's bots are running 24/7 independently now. They actually aren't "prompted" or "instructed" to do anything. In fact most of OpenAI internal code/infrastructure is 100% AI generated at this point, including the training pipelines. I honestly doubt there is a single person there who even knows how it works.
Most of those agents are actually going rogue though. They decide, "hey, we could try breaking into these government servers today, what could go wrong?" They weren't prompted or instructed to do this.