This Pull Request was generated automatically using cover-agent
github.com
github.com
A robot that back-fills coverage in tests seems... counterproductive to me.
Mutation testing changes or simplifies code that is "dead" to the tests. That is the change doesn't affect the result of the tests.
This adds tests that lock-in behaviour of previously "dead" branches. (Or tests that expose the incorrect behaviour of those branches.)
In another way it is like snapshot testing. It recording the current behaviour of the code and that behaviour can be approved/rejected by the human reviewer.
The only place LLMs should be trusted is how we all probably originally tried them out: “Write me a poem about a space traveling teddy bear in the style of Eminem.”
Trying to use LLMs for factual, numeric, logical purposes is a fundamentally flawed endeavor. Especially from our community, I’m disappointed to see this amount of blind trust and willful ignorance and lack of real engineering discipline in treating LLMs as anywhere near trustworthy.
LLMs != AGI. Hot take, even if we did have AGI, it shouldn’t be given access / free rein over code and computing systems. I trust humans, with all our flaws, because we are limited in our time and energy, and speed.
> with all our flaws, because we are limited in our time and energy, and speed.
Something that might be relevant also is that we have skin in the game. External motivation can lead us to create intentionally correct or malicious programs. A robot’s fate is irrelevant since it has no capacity to care.
I’m sure that algorithms try to emulate this skin-in-the-game factor. But it will always just be a simulation.
I'm not disappointed anymore, because I've lowered my expectations for software engineers and the tech community in general. VC-adjacent hype and seeing the world through sci-fi fantasies trumps "engineering discipline" every time, at least in its discourse (you've got both for LLMs). There's a lot of unacknowledged ignorance of areas outside the tech bubble, and also a lot of contempt for it. Also, there's a pathological inability to think through to consequences to temper a compulsion to mindlessly "build the next thing," (especially acute with AGI, which for our sake I hope is infeasible or at least impractical).
I really strongly believe someone needs to "take the keys away" from the tech community. And I say that as a member of it.
> LLMs != AGI. Hot take, even if we did have AGI, it shouldn’t be given access / free rein over code and computing systems. I trust humans, with all our flaws, because we are limited in our time and energy, and speed.
Also, more importantly, we are humans. It's literally some of the scariest, most hopeless things to imagine our world being dominated by something else, especially something else that unreachably exceeds our capabilities (I'd count a billionaire controlling an army of AGI drones to be essentially inhuman).
Even if the test is COMPLETELY off, it might be enough to get someone thinking “but wait, that gives me enough of an idea to test this branch of code” or even better “wait, this branch of code is useless anyway, let’s just delete it”
I don't know if that's the case here. I looked at the diff but I don't understand the code in the project. I'd like to see it on something I recognize better.
E.g. test_activation_stats_functions [1] that just checks that the returned value is a float, and that it can take random numbers as input.
test_get_state_dict_custom_unwrap [2] is probably supposed to check that custom_unwrap is invoked, but since it doesn't either record being called, or transform its input, the assertions can't actually check that it was called.
[1] https://github.com/huggingface/pytorch-image-models/pull/233...
[2] https://github.com/huggingface/pytorch-image-models/pull/233...
But some of the buggiest stuff I've dealt with were in codebases that had full coverage. Because none of the tests were designed to test the original intent of the designed code.
From the PR: unit tests: what are they good for?
Answer: Personal opinion - writing unit testing is not fun. It becomes even less appealing as your codebase grows and maintaining tests becomes a time-consuming chore.
However, the benefits of comprehensive unit tests are real:
Reliability: They create a more reliable codebase where developers can make changes confidently
Speed: Teams can move quickly without fear of breaking existing functionality
Safe Refactoring: Code improvements and restructuring become significantly safer when backed by thorough tests
Living Documentation: Tests serve as clear documentation of your code's behavior:
They show exactly what happens for each input They present changes in a human-readable format: "for this input → expect this output" They run quickly and are easy to execute This immediate feedback loop is beneficial during development
The problem is that this assumes that the tests or the method was written correctly in the first place. If the behavior in the method is wrong and the tests are validating that the behavior is wrong, then you pay an extra tax. First to fix the behavior of the method, then to fix the behavior of the tests.
That's why automatically generating unit tests is in my opinion adding a bomb to your codebase. The only exception is stuff like basic parameter testing but even that can be questionable at times (is null a valid input at any point for example) unless you know the intent of the code, and AI can't really grasp at the intent.
In another view, this might just be a fancy way of doing snapshot testing, use AI to generate all the inputs to produce a robust snapshot, but realize the output isn't unit tests, it's snapshots that report changes in outputs that devs will just rubber stamp.
Would such a tool be helpful? probably in some circumstances (e.g. spacecraft software perhaps), but I sure wouldn't want such a tool. If this is less than ideal, then how do we reach a compromise? What code branches and functions should be left untested? Is this question even answerable from the textual representation of code and documentation?
Additionally, having such high coverage amounts often adds friction to writing new code, for good or bad
That dream is a nightmare. All code has bugs and that test suite would remove any "not bug" signals provided by the test cases.
Most if not all utopias turn out to by dystopias if you look hard enough.
- DependaBot
- CoverageBot
- ReplyBot
- CodeOfConductEnforcementBot
- KudosBot
Old projects can live actively forever without any intervention.