How does it feel to test a compiler?
medium.com
medium.com
E.g. if I want to test that `*` has higher precedence than `+`, I would write something like this:
assert_ast_equals(parse_expr("(a*b)+c"), parse_expr("a*b+c"))
assert_ast_equals(parse_expr("a+(b*c)"), parse_expr("a+b*c"))
You can rewrite the whole compiler if you want, but as long as you have some notion of a "parser", an "AST" and "two AST nodes being the same" this test will keep working.This is much more powerful than going into the parser internals and comparing get_operator_precedence('+') with get_operator_precedence('*') which is the default thing you would do if you're told to test every function after writing it.
John Regehr and his students have done impressively deep work in finding compiler bugs.
• https://blog.regehr.org/archives/category/compilers
• https://blog.regehr.org/ (some of his compiler-related posts aren't properly tagged)
• [PDF] https://users.cs.utah.edu/~regehr/papers/pldi11-preprint.pdf Finding and Understanding Bugs in C Compilers (edit I see someone else already posted this paper)
• A (very) little discussion: https://news.ycombinator.com/item?id=7728035
As examples: If you are creating a web API, only test against the exposed public routes of that API, instead of writing tests for internal helpers. If you are creating a GUI application, programmatically exercise the GUI instead of starting your tests halfway into the innards that respond to the GUI buttons.
Tests lock behavior into place, so you should only write tests for behavior that needs to be locked.
Unit testing enables hidden functionality to be tested. This can prevent future bugs when changes in the system suddenly uncloaks those functions, exposing them to higher level tests.
Unit testing is fine, if you do it for your public interface. If you are writing a math library then sure, unit test that `add(1, 2) == 3`. But if you just have an internal helper function for that, then think about if you really want to lock its existence and behavior into place, or if that would just hinder future architectural changes.
You can always test the exposed functionality that uses the helper and achieve full coverage of it that way. If you can't, then you have dead code.
Of course this is all a bit more nuanced. Past a certain size it might make sense to e.g. consider one modules interface to be public for the rest of your application and test it. But you can definitely overdo it and testing every single function you write (as I've seen people unironically suggest) is very likely detrimental.
A good language will provide clear boundaries such that it is obvious which is which.
Testing internals is almost always a code organization smell IMO.
"private" tests are more for like when you're having trouble figuring out an edge case failure and want to narrow it down to a specific helper function to aid debugging or if you need help coming up with the right design for an internal function. As before, we're talking exceptional circumstances. Rarely would you need such a thing. But if it helps, no need to fear it.
Either way, I'm not sure you would be looking for good coverage, only the bare necessities to reach the goal. Once settled, the tests are disposable. Organizing your project into internal libraries in case you encounter a debugging problem in need of assistance, for example, is extreme overkill.
I certainly wouldn't advocate for code reorganization for that case, but if there is a property that is important to maintain over time that isn't easily expressed by exercising the public API, it does suggest that reorganization is probably in order.
I agree it is not "tests" in the documentation sense, and I did mention that, but, regardless, you do seem to align on "throwaway". Now you have me curious, what kind of ephemeral "tests" were you imagining if not something similar to what I described in more detail later?
I agree with everything you’ve said up until this point. The add function should be locked down, whether or not it’s internal code. I’ve written some physics libraries for games I’ve coded in the past, and you can bet that I wanted to lock in the functionality of that complex physics code by working some examples by hand and then codifying them into unit tests!
I think you’re right at the end, everything is nuanced. Using the right tool for the right job is easier said then done :)
The other problem is that finding tests that cover everything from the public interfaces can be extremely difficult. Most test suites don't achieve full coverage even of theoretically reachable points in the code.
If they aren't then you are defending against situations that can never occur. In that case the check seems unnecessary, and I at least wouldn't make it a priority to cover it. Also, those checks are usually simple enough that one can reason about them easily, again making it less of a priority to have tests cover them.
> The other problem is that finding tests that cover everything from the public interfaces can be extremely difficult. Most test suites don't achieve full coverage even of theoretically reachable points in the code.
Yes, testing is hard. That is not a reason to write more detrimental tests though.
This insight of yours cannot help but save on costs. Be sure to suggest it to management.
Sounds like a good idea, actually. These places being the entire surface of your code that is exposed to the real world. So not unlike what I have been suggesting.
Do you have any particular example in mind that is not superfluous and really can not be exercised by the public API? I have a hard time thinking of any.
If you pin down too much in test cases, on the other hand, it will be hard to make changes to the compiler without breaking tons of tests.
I think, the best way to test a compiler is to have a large standard library or other body of code, including that compiler written in itself. Recompile the whole thing and then recompile it again with the compiled compiler and again. By the third iteration, you should have it a fixed point. That doesn't prove thing are correct, but it gives a lot of confidence, especially if the code base is large and uses a lot of the language. The second piece of confidence is that all that compiled code passes its tests.
Lots of places have a two tier system where the real developers write the code and those who don't make the cut test the code, with pay delta and an aspiration of being promoted out of testing.
Other places have a mandatory stint in testing for new developments as a way to get some headcount on the task.
Jetbrains don't do that. Or at least they didn't sometime before covid when I met a bunch of their test devs at a conference. The developers mostly doing testing were equals to those mostly doing product work. Possibly with a more extreme bias towards case analysis.
I don't think it's a coincidence that jetbrains treat their test team as peers to the others and that their software seems to mostly not fall over in the field.
(Spelling)
Everyone jumped on the testing bandwagon but writing code that is testable is a learned skill that nobody bothers to learn. Instead we end up with overly-mocked tests that—in practice—test that “the code is written like it currently is” rather than that it actually behaves correctly.
These kinds of tests actually provide negative value. Besides taking up time during every PR and build, they fire off on any attempt to clean up or refactor. Every time you edit the code you have to edit the tests which entirely defeats the point.
The solution ends up being, unsurprisingly, lessons learned from the functional world: don’t access or manipulate external state, operate only on direct inputs and only manipulate your outputs. As a result code ends up being dumb, short, and obvious. But nobody learns to write code this way, so two hundred line spaghetti functions are the norm.
I have seen functional and e2e tests prevent regressions far more times than I can count. The nice thing about those is that they force testing outcomes and don’t require “testable” code like unit tests do.
Don’t throw the baby out with the bathwater
At one job, our code had a lot of math, and the worst bugs were when our data pipelines ran without crashing but gave the wrong numbers, sometimes due to weird stuff like "a bug in our vendor's code caused them to send us numbers denominated in pounds instead of dollars". This is pretty hard to catch with unit tests, but we ended up applying a layer of statistical checks that ran every hour or so and raised an alert if something was anomalous, and those alerts probably saved us more money than all other tests combined.
Eventually I setup something simple to exercise most of these functions, as well as most other balancing code, plot their results and send the results to the balancing team a week or so. For example, using D2 terms, it would run 1M item drops at different magic find levels and plot the resulting item level distribution, graph the implemented level -> hp/mana/... curves and such.
That little thing caught so many implementation issues at first and we'd regularly have balancing poke us because someone "optimized" some code in there.
I’ve come to the conclusion that mocks at all are evil when done for anything except external services not under your control. And then, what you should be making is a trivial fake implementation of the other end, not just checking that specific methods were called.
Every testing environment I have seen since have looked like a joke but to be fair the testing budget of this place was bigger than what some companies spend on a whole product.
That is, you can have a solid testing team for pretty much anything. You have to empower them, though. Largely, a lot of what you get to empower them with is ability to put constraints onto the engineering team.
This could be crappy constraints that are far sweeping on the product. There is no need that it has to be that, though. The only constraint you have to have is stability of product interface. Which.... yeah, our industry doesn't do that so well. (Note, not stability as in "doesn't crash." Stability as in, same inputs that worked last year work today.)
And of course we still have manual exploratory QA, it’s difficult to replace that safety net.
Very little in this industry is done well. Tests are one of those things.
(1) They do not test at a level that is too low. If you test individual classes that do not have much logic in them, it probably is pointless. You are probably testing exciting facts like 'is the container type in my favorite language still capable of storing elements'. Mostly, it is best to test the interplay between a small number of classes on such a level that the attributes are such that the customer could recognize them as something that they value.
(2) Don't test at a level that is too high. Integration tests often take a lot of effort to set up, they are brittle, and they are slow. Have a small number of them but not too many. I.e., the concept of the testing pyramid.
In this thread one can read various assertions that sound like nonsense to me.
(1) Only write functional code: nice if you can get away with it but some code has the explicit purpose to manipulate things in the real worlds or is explicitly there to maintain some state. One might also watch https://www.youtube.com/watch?v=j71n33A0CkI&t=314s . The video is right. There actually is not much difference.
(2) Don't mock. Well, if your program is big enough to contain many classes and/or functions so that it would be too big according to criteria (1) and (2) above, it becomes impractical to test all of them together so you will need mocks. What can just be passed in as function arguments can be passed in as function arguments but what comes out as behavior may need to be mocked.
(3) Create test doubles instead of mocks. Use whatever is most convenient. Mocks record series of function calls. Test doubles maintain some state and one can see how that changes. Both can be good or bad depending on the situation. It seems pointless to have a preference separate from what you are trying to do
I've tamed a few untested code bases by writing huge integration test-ish unit tests, and comparing the output with pre-recorded answers. After that, you start adding real unit tests for fresh code, knowing the existing stuff is quite safe.
There's an assumption here that tests need to be run for each build. This rules out more laborious testing that could be exposing bugs not found by time-limited tests. For compiler testing, one can set up property-based tests that run for unlimited lengths of time. These can be set up to be always running in the background. Literally billions of tests can be constructed and run.
I did a presentation on that last year:
At one point, I ended up on a team that had virtually no QA and only did automated end-to-end tests on a complex product. After working through a lot of tech debt on the testing (took a few months), it was solid and caught many defects introduced by junior/new developers (easily missed in code review).
I now have the fortune to be in a different group with a similar mindset. I’m motivated to write tests for my own code just to prove to myself that it works properly in odd error cases too. My team prefers testing instead of “run fast and break things.” (And with multiple decades of experience, I humbly can state I cause significantly fewer bugs than in the past.) This also makes me write code in an easily testable manner. Trying to test already written code is usually a lengthy and painful exercise in frustration.
It’s also too easy to write error handlers that are NEVER actually exercised once, even manually — and surprise, they may not even work! People who never write tests do this quite often.
Perhaps people in these environments are made of different stuff but touching both sides feels almost essential to me.
Testing your own work also does nothing to protect against correctly implementing the wrong thing.
I’ve seen both and have actually seen much higher product quality when the QA team is smaller (or even almost non existent).
Turns out white-box testing usually fails to catch things only SMEs would know.
[1] During the 11 years I was there, there were only two deployments that were reversed, prior to new management. And the bugs were found during the immediate testing after deployment and we were able to reverse quickly.
[2] The Oligarchic Cell Phone Company.
And val team was creating fake product which was used to help to write tests without product (so tests were available faster, so deploy was faster), or just tests - all of this basing on spec
It worked well when it comes to bug/issues/inconsistencies catching
There are a couple effective techniques in the literature that might be useful here:
- Differential testing[1]: generate a bunch of random, correct, deterministic programs; run them under different compilers or under different compilation flags and check if the output of the program is identical
- Equivalence Modulo Inputs[2]: a class of techniques that can be used transform a program to a distinct program that is supposed to be equivalent to the original for a specific input. (shameless plug)
[1]: https://users.cs.utah.edu/~regehr/papers/pldi11-preprint.pdf
$ valgrind --leak-check=full --track-origins=yes ./lisp < test/test.lisp
$ cat test/expected.txt
What can I say, works for me.But as with any code, compilers have bugs too and sometimes they can be quite surprising.
That's not ambiguously standards-compliant, that is literally something that the standard has explicit mechanisms to indicate that's what the compiler is doing (look up FLT_EVAL_METHOD).
Some people would rather blame the compiler than their own code.
Others were shocked that a compiler could have a bug at all. Because most code is not corner cases, this isn’t a bad default assumption, really, but bugs do happen.
[1] https://www.coinfabrik.com/blog/why-the-fuzz-about-fuzzing-c...
Would that be the LLVM one wrapped in some rust API? From the last month or so accidentally rabbit holed by fuzz testing I am really liking the LLVM one but have a vague sense that I should probably try AFL++ as well.
The paper suggests it was a java implementation distinct from those two. The approach of turning the bytestring into a program is good, worth looking at the IR fuzzer in the LLVM suite - this looks like a fair reference for it https://arxiv.org/abs/2402.05256
It also disturbs me that the author mentioned the order at which sources are compiled to matter in the final result. It should never matter.
When we build software we should always make it in such way it’s trivial to write tests for it. If writing a test is easy, it indicates using the tool you wrote is also easy.
1. Functional programmers often write slow code. It turned out that my compiler was spending most of its time in my professor's code that while I'm sure was very mathematically pure, was a large consumer of immutable, short lifetime objects. Meaning under the hood mallocs. I should've valgrinded it but I'm certain it would've overflowed the counters (jokes)
2. If a comment spanned multiple lines in the resulting assembly, I could escape the comment and operate outside the bounds the professor setup, letting me use more assembly directives to solve the problems way easier. Ultimately we worked to fix that because usually it just means the student will try to compile part of a comment as assembly and that can be very confusing for the less assembler-error inclined. I used it for having a constant before a variable for type tagging. A 1 line solution. I believe the class's preferred way was putting the tag in a register and yadda yadda something that took a lot more finagling and effort. I did that maybe once before using my knowledge of the comment escape to do the arbitrary code injection.
[1] interpreted was written in the language we were interpreting, so as long as there were no typos or logic errors, the functionally was perfect vs running the code in the programming language. The compiler would return back a series of objects that wrapped assembly. For example, Add(R2, R2, R3). Usually pretty transparent. The framework we were given would then write out the .s file, I believe it would call some gcc or other thing, and we'd run the binary to make sure it worked.
“tailrec” is what you mention in function definitions to indicate that function is tail recursive, but it only actually worked when you called the same function itself in the return, and not any other tailrec function. The part of this which was idiocy was this would only manifest as a StackOverflowException at run time. (I found this as my language evaluation involved implementing a state machine idiomatically). If you are going to make tailrec only work as a while loop then have the compiler alert the programmer at build time, but this got all through their design and QA. Not exactly a well thought out process.
They have probably fixed this now, but instead I went off to golang, where the features are few but when they exist they are done properly.
The whole thing is attempting to hack together something that superficially competes with c# but just doesn’t have the substance.
I'm not aware of implementation details of either Clojure or Kotlin but if it is anything like F#, then I would really not hold tailcalls against either of them.
In order to effectively support FP languages, CIL specification defines tail. prefix for call, calli and callvirt opcodes, and CoreCLR fully supports it except scenarios outlined by the spec.
Does JVM bytecode specification have something like this?
Edit: Turns out F# compiler does additional heavy lifting besides emitting the prefix https://github.com/dotnet/fsharp/blob/main/docs/large-inputs...
the clojure approach (as i understand it; i've only written a few dozen lines of clojure) is to use an explicit `recur` operator, at the call site, for a tail-recursive call. implicitly that operator invokes the same function, not a different one, because that's the best you can do on the jvm. it's not a runtime error
i.e. someone familiar with tail recursion but not the specifics of kotlin would not reasonably expect any such restrictions. It is better to not have such concepts at all than half done in this way, as it destroys confidence in the rest of the language constructs.
You don't get to say "it supports tail recursion" and then only do so in the most narrow sense imaginable - i.e. the case that is trivial to turn into a while loop. If they can't make actual uses work on the JVM properly then don't claim to have the feature in the first place.
i agree that 'Kotlin supports a style of functional programming known as tail recursion.' is somewhat exaggerating kotlin's capabilities here