REPL Driven Design
blog.cleancoder.com
blog.cleancoder.com
It is my experience that good software comes from teams working together well, teams that are in some sort of harmony with each other AND the system they are building/using. They understand it and each other. Good software is not derived from a singular technique.
I do like test-first approaches when I can, I like SOLID. I probably agree with Bob on a lot, actually. I think they are good techniques to employ.
However, it has not been uncommon in my career to find people (and myself) getting dogmatic about these things and arguing over them. There are half baked attempts to change everything to an Entity. There are flaky test suites with a lot of mocks because someone is hot on TDD, but does not have the bandwidth to maintain the suite for the whole team.
This blogicle is another example. As pointed out by other posters, REPL driven development can fold VERY well into a test driven workflow. However, that's not enough, he's actively saying it's not as good and you need to do The One True technique.
Testing definitely has value, but when testing you always have to make the assumption that you can’t test everything, and then the trick is deciding what to test versus what not to. I think people forget this and focus on having every function tested, every interface having coverage, etc.
Not to say there is not still a place for manually written unit tests.
And I say all this as someone who has pursued this very idea of property testing side effects for years and in the end had to admit failure. It just is too much work get right (eg uuids being generated not unique, to name just one example), and your test runs become incredibly slow, which is a killer in many REPL-driven workflows.
Expressing your state in terms of pure functions of input events, as can be done with an FRP like Reflex, means you can test all your logic purely. You can be confident that you'll only request to launch the nukes when the nukes need to be launched.
Then you have to make yourself confident that, when the system asks the nukes to launch, it sends the right signal. But that's mockable by running the system attached to device that reads the signal and reports its correctness.
And making sure the nukes take off, aim correctly, and land on their target is not my department (says Werner von Braun).
The problem with that (tautological) statement is that too many people read it as: "you can't test everything, so don't bother testing anything".
I find it much more convincing if someone were to say, “test what’s easy” instead, that’s all.
1. stumble over differences from their normal approach
2. catch on and use it everywhere (Stockholm syndrome? Testing limits?)
3. finally find places it doesn't work and can apply judiciously
Given that in our profession we're constantly learning and experimenting, Uncle Bob is still doing that to this day.
That there are drawbacks to the REPL feedback loop is true. But they aren't the drawbacks Bob claims to have found here, because he ceased investigating at the point he felt his priors were confirmed.
A set of principles for program design, specifically OO (and some flavors of OO at that), but some of the ideas are general enough to apply to most languages and paradigms.
Testing is his singular solution for everything.
I think it was this video, but I can't watch it all the way through right now -- https://youtu.be/ecIWPzGEbFc
His lectures and books really are harmful. Beginners read them and get the impression that you can design good software by following a bunch of simplistic rules of thumb. In reality, you need to understand what trade-offs you're making when committing to certain design decisions and, even more importantly, you need to be able to discuss this with other people.
For example, it's not "small functions are good", it's "when you make a function smaller, here are the benefits and here are costs". Primitive, sloganistic thinking is the bane of software engineering profession right now.
Saying that his posts are harmful is a bit of a stretch.
So I agree with the parent ... his writings are in many cases harmful. They're a mix of good and bad, but often net harmful.
Look at how much growth happened in the last 15 years.
Why on earth would you think that?
Saying that his posts are harmful is a bit of a stretch.
Mentor your first few people who keep defending rookie mistakes with "But Uncle Bob said..." and you might change your mind about that.
Bob Martin says some sensible things, but he also says some things that are at best debatable and at worst outright contradicted by the evidence base, and he says all of them with great conviction and presumed authority so that his target audience can't tell the difference. That alone makes him dangerous as a source of advice.
I have been developing for 54 years. I have been through all manner of design philosophies, from Structured Programming, to Object Oriented Programming, to Agile, to TDD along with hack-it-till-it-works.
If you don't do the engineering, which he does not teach, you are going to struggle.
I believe this is a harmful post.
Another OO-evanglist who belongs to this camp is Ron Jeffries.
For a textbook example of how much of a snake-oil OO-evangelism is, checkout the sudoku solver fiasco involving Ron Jeffries and Peter Norvig (you'll find plenty of search results regarding this).
They are awful to work with and be around. All of those concept I mentioned have benefits in them, but not as a religion. Medium programmers lack nuance and deep experience and are dismissive of those that do, they see it as weakness. When their favorite technique doesn't take them to the holy Nirvana of software, they lash out at other, less zealous developers and blame their lack of discipline.
Completely shitty people - I just hope they grow out of it. But they are seductive in their intensity and (bad) managers love to hear how they have a simple solution to all their woes - rather than software development being inherently difficult (because you only have to do a specific hard things once, but that's another discussion!)
A (series of) REPL commands is just a test, where acceptance and rejection is manual: we look at the output, and keep changing the code until we get what we want.
At which point, we should be able to say to the REPL: yes, this result is correct, this is a test, and it's now called "Thing does what I want when I give it X and Y".
The runtime then saves this as a test, and will run it automatically when you ask it to. If there's a regression, it will boot back up into the interactive REPL so you can inspect it and fix it.
I've been writing a REPL. Turns out that's a lot of work, but I'm getting to the point where I can add this feature, which is the last one I really want to add before I go public with this thing.
I store the complete repl history, complete with results and session metadata, in a database.
Pruning a session's history down to only what's needed is a manual process, true, and laying the groundwork to provide that feature has been the main yak shave.
Before I get to my real comment, I guess a fair criticism of it would be “Well, where in the [development process] hierarchy do you actually want to live?”, which I guess is a longer discussion, but as a developer, you can’t become battle-hardened without going to battle. So we can nod to Carl Sagan[0] and say “I made this from scratch”, and know there’s value in that. At least occasionally running your REPLs by hand and taking the actual trip, instead of just wanting the destination. It may not be tenable as a full-time production development strategy, but that doesn’t mean it’s worthless.
Say what you will about TDD, but it does (often? sometimes? depends widely on how the tests are being written) produce useful, informative and functional artifacts. If you need to know how to create some entity or how to run some process, chances are there are tests that you can refer to that will show you how things are intended to work. This shouldn't take the place of comprehensive documentation, but it's a welcome counterpart.
With REPL driven development, there's nothing of the sort. If the developer has not provided documentation, you'll probably have to try to reverse engineer a given workflow based on the state of the system, which can be very problematic.
Now imagine that you can instantly load a REPL session with complete tests, already executed, hanging out in the session history.
Now you can change anything in the test and see what it does. That's the idea anyway.
Wouldn't there be production code that you can read? You're already working in the codebase after all.
I agree that a unit test might have a cleaner, distilled setup procedure though.
This is absolutely the result of poor practices and, documentation aside, if the developers had been using tests to drive out this behavior, I'd have had a sort of playbook that I could reference.
It's not uncommon for REPL users to write tests for their code. And nobody worth their salt would think that the REPL is a license not to document and operationalize their work properly! Your root cause seems to be some mix of developer laziness or inexperience, and poor business practice -- not really a tooling issue.
import IPython
IPython.embed()
I have used similar techniques to inject unit tests into the same pass as the integration tests. Basically at this point in the program, pause and run a bunch of unit tests with the internal state setup by the integration test. Requires less mocking, can differentiate between an integration and unit test failure.The (long) roadmap involves implementing the Jupyter protocol, adding a client library which can be injected into production code, and eventually using that common basis to extend it to the other languages supported by Jupyter, and any others which choose to do the groundwork.
Bringing in all of ZeroMQ is a lot of work I've been putting off, but I need it all over the place.
https://elixir-lang.org/getting-started/mix-otp/docs-tests-a...
I actually want this for both code and test development. For me the REPL is the first place that sees an implementation- multiple implementations, weighing various design factors- which then gets copy pasted into a source file and then REPL state reloaded (the Clojure "reloaded" workflow), and then the capability offered by the updated implementation exercised with test-like code in the REPL- which then as you describe should magically get into the test suite source for reuse.
Good luck. What language are you doing this work in?
When I work with Ruby, I often set things up with a script that contacts a method that I can call to trigger reloading of everything and a "replay all known good steps" while storing some state in globals. This is fairly easy in Ruby as you can redefine classes at will. It takes a bit more effort if you want to replicate the state more perfectly, e.g. by deleting methods that have disappeared etc., but that's rarely needed.
I then use "pry", and fire things up while I keep my code in an editor. I can test things in the pry repl, and when I find something that makes sense I add it as an additional method to my code. I usually cut and paste and keep the editor separate, but "pry" does allow triggering your preferred editor directly from the repl and also supports its own code reloading -- I prefer my own mechanism because it gives me more direct control, though.
Regularly I'll exit and run from a fresh state just to make sure I've not accidentally messed up anything with the reloads, but really that's mostly paranoia.
It's not a replacement for tests, but it tends to produce relatively easily testable code because it encourages code that allows you to quickly and painlessly execute steps manually, and so you quickly will e.g. encapsulate mocking steps and setup in suitable helpers etc. because it gets tedious to write them manually.
I like that a whole lot better than TDD, because TDD tests to drive (over)thinking things first vs. experimentation, I feel, because you invest too much in the test cases upfront.
The test cases are important, but it is very easy to end up with a design that is optimized for the tests rather than for your real world use cases. It doesn't have to be that way, and some people manage to do TDD without getting caught up in that, but for me at least, I've found it's much easier to be disciplined about making the code easy to execute in general when I work with it in a REPL on a regular basis, and build tests around that.
The big disadvantage, which is why I don't use it, is that you pre-commit to the "correct" function signature and response to some number of parameters. Optimized for the test, as you say.
When I'm working with a repl, it's easy to mutate and experiment with those things until I'm satisfied with them.
But then you still have to translate that into tests, or you get regressions. I'm lazy; I want to skip that part.
In short, these sorts of tests are super easy to create but become a massive friction at every level of the application since they depend on so much.
As an idea, maybe we could do... both? Explore at the REPL, and then commit to tests and clean up once you've solidified your approach (i.e. how most Clojurists likely operate).
(some-expression 1 1) ;; => expected result
(some-expression 2 1) ;; => expected result
Then I use a `let` to make changing the variable part easier: (let ((x 1))
(some-expression x 1)) ;; => expected result
(let ((x 2))
(some-expression x 1)) ;; => expected result
;; repeat
Then I turn it into a function: (defun my-fun (x)
(some-expression x 1))
And copy all my REPL experiments into the test suite: (test my-fun-tests
(is (= ?? (my-fun 1)))
(is (= ?? (my-fun 2)))
;; repeat
)
Since I actually tend to use org mode for writing lisp code, most of the first two sections are actually inside a block like: #+BEGIN_SRC: lisp
...
#+END_SRC
So I just have to edit this and turn it into the relevant test suite.If I don't have a REPL, I often write a lot of little throwaway tests to deal with intermediate bits of the code. They end up being redundant with the real tests. They're often really bad tests, in that they're tightly coupled to implementation details. I should delete them. And I often forget to. Which creates extra maintenance burden for myself and my teammates.
If I do have a REPL, I do all of that extra intermediate verification junk in the REPL. It's often more effective, it's usually less effort, and it always reduces the volume of unnecessary brittle tests in the codebase.
That's a simple lack of foresight on the programmer's side that is then blamed on the tools that were used.
Does anyone know of an open source codebase he's written or been a major contributor to? I know he writes a lot about clean code, but I've never seen production code from him.
Pretty sure he hasn’t written prod code in over a decade.
The general point that code should be clean is a good point to make. Writing a whole book about it seems a bit much.
Yup which is why we have millions of npm packages that do literally one thing.
What the npm ecosystem is doing is a completely different thing.
There’s no difference between module and function when you see that.
Until you really have to understand what is going on, like when you need to fix a bug or refactor. Then suddenly you to deal with the overhead of holding the branches of 20 levels of indirection in your working memory.
Also, this assumes the people you're working with are great at naming things. Which, from my experience, is almost no one.
Your job as a programmer is being able to walk this ladder of indirection up and down until you find what you seek and fix/add/remove whatever.
You'll always be dealing with way more than 20 levels of indirection when programming, even with a "hello world" program. If adding two more levels of indirection makes your mental model break down, you might consider upgrading your environment/workflow a bit to support this.
Regarding naming I agree with you. Good names are really, really, really hard to come up with. Entire books have been written on how to name things. Doesn't excuse us from trying though, the only way we can get better is by doing.
And it's not even entirely true. You can see that the top ten depended upon packages [1] in npm are the opposite of "literally do one thing". They are either full blown libraries like React or collections of related functions. I'm not claiming npm does not have a problem with too many one liner functions. But the most popular packages do avoid this practice, which is encouraging, and sets the bar for what you should do if you want to be at the top.
What other motivations could there be for being the author of many very small and popular npm packages when it comes to interviews and internet fame? There are probably a lot more reasonable answers than "influenced by Bob Martin".
https://www.npmjs.com/package/is-promise
7 million downloads a week is not "popular" to you?
npm is riddled with packages like these.
Just put my claim back in context, with an emphasis on the first sentence and you've got your answer.
While I do think that Bob goes overboard with his suggestions in Clean Code, I found it as a necessary counterbalance to encourage readable code in my team, where the methods tends to be gigantic blobs. I'm sure there are better books, but it was a good one.
For instance, one suggestion I did like from Clean Code is to separate exception handling from business logic. I don't always practice it, but it really helped in making my work-related code more readable.
A few years later a friend referred to this dismissively as “Programming by successive approximation.”
There is a lot of truth to this, but the distinction isn’t as sharp as it sounds: you still have to think ahead, and as you explore and expand you’ve typically always got something working, some test cases and invariants, and typically you have the benefit of working on live data.
Even though I do most of my work in compiled C++, my work model isn’t that different.
The usual workflow was something like
(comment
;; experiments go here
(defn add [a b]
(+ a b))
(= 2 (add 1 1))
(= 3 (add 1 2))
(= 99 (add 9 90))
)
Once you have what you want, you move the code out of the comment block. Whether you move it into a library or a test file is just dependent on the type of thing you're moving. Functions moved into library code. The samples that proved the happy paths (or unhappy paths) would go into test files using `clojure.test/is` (and typically a direct copy/paste).(I have no idea if an equivalent workflow exists in Elxir, but I've never seen it been done that way: just like people use ipython, they use iex in my experience.)
Maybe I'm just misunderstanding what people in Clojure are doing. Please feel to point me in the right direction.
Like, folks... REPLs are really nice when you want to figure out how a function or module or whatever works. Then yeah... you can take that knowledge you gained and apply it to your work.
Like.. how is this a deep profound revelation for folks?
a) Knowing both how to effectively write tests & how to effectively use a REPL.
and:
b) Knowing only how to bang around in a REPL enough to use it as a crutch to avoid writing any tests.
REPL is definitely not REPL.
[] I really don't think it is: reload(mod) comes with way more caveats than C-c C-k in CIDER.
You're making far too much of the distinction between different kinds of REPL. As Peter Norvig put it [1], as long as you have a "development process that allows for rapid development with continuous improvement", minor distinctions between languages [such as quibbles over how 'reload' works] really don't matter that much.
Squeeze all the power out of the tools that you use, and get things done -- ultimately that's all that matters. Religious wars like these were tiresome in the era of the Smug Lisp Weenie, and they haven't improved with age.
You ask how I could know that: I have used Python for a similar amount of time as you, and have been a Fellow of the PSF for about a decade. During that time I've attended dozens of PyCons across the world. My wife is the Executive Director of the PSF, so I've been the second person in and the penultimate one out for most of those PyCon USes, including the post-conference sprints. Previously, I was the Freenode group contact for and one of the most active contributors to #python. I've worked with a lot more Python programmers than most, and typically meeting them where _they are_: either pairing next to them or over IRC. People from all over the world and different experience levels and backgrounds. People who lived and breathed Emacs and are definitely aware of how SLIME/inferior processes work. I can't recall a single instance where someone said "eval this in a REPL" and the other party did that by calling reload() or exec from an editor. The vast majority of times someone mentioned reload(), it's been someone who built a long-lived process like an IRC bot, with hot code reloading via reload(), being confused that reload() "didn't work" because they kept seeing the old behavior. YMMV. This is all anecdata, sure, but I observe that IPython 5 casually broke inferior-shell compatibility and somehow that made it to release. It's possible that there are huge swathes of people who do write their own "trivial" reload()/exec integration, but that feels like the stronger claim to me than "there may be confusion here between what users of these languages typically mean when they say REPL".
I want to make sure I understand your argument correctly. Starting with the strongest form of the argument: are you saying you believe most Python/... programmers think "REPL" means "something involving reload" and not, say, ipython? I assume we can agree that seems silly, but perhaps not: and I think that's closest to directly refuting my claim. Unless I'm misreading your comment: I'm saying "different groups use this term to mean different things", and your comment is "come on they are practically the same" (ok, fine, we can disagree!) and "you can do reload()!" (sure, but I didn't say it's impossible, I said that's not what people generally mean). So, perhaps a weaker version of your argument: you're saying you think some nontrivial amount of Python users use a reload/exec-based editor integration? Enough that "REPL" is not suddenly a confusing term? So, something like what, 20%, 10% of users?
As for my argument, I think it was in two points. First was to challenge what evidence you had to support your claims. From your reply, I'm confident from that you're speaking from knowledge of many developers' experiences. Perhaps I am more of a special snowflake than I realized.
As for the second point, I'm basically restating Norvig's argument, as I think it's a good one. I'm not terribly concerned about distinctions between different REPL-ish tools. I do take your point that the distinctions can muddy conversation, e.g. when trying to relate experiences of (say) Jupyter users vs. `python-mode` users -- or worse, Lisp vs. Python.
But to paraphrase (co-opt?) Norvig, I think these are second-order differences. Important enough in a few contexts, but distracting in most. My thesis is simple: everyone should experience interactive programming, and should ideally incorporate it into their development process, because it makes development faster and more effective. Whether that means Jupyter, CIDER, SLIME, ghci, etc. is a local decision, based on where you're at and what your community is; I don't think "REPL" in particular is a special shibboleth of the Enlightened. The difference between the Java/C++/Rust set and the Python/Ruby/CL/Clojure set is far greater than any minor differences we might have within the interactive camp, and "internal" disagreements in front of the unindoctrinated tend to distract from the greater good, IMO.
Any way that we can help new developers to experience interactive programming is a positive step. Personally I think that's the first-order problem, and I think our collective, repeated failure to recognize that is a source of much pointless and religious argument. Perhaps that's why this conversation reminded me of those wasted days, and why I responded with more aggression than I should have.
I also happily stipulate that non-Lisps not working that way is largely praxis and not some fundamental language deficiency, and definitely strongly agree that interactive programming is a powerful tool we should get into more hands.
More interesting would be a correct solution, which was posted by Peter Norvig. The approach he uses is drastically different that anything TDD is likely to get for you.
TDD isn't going to find solutions to hard problems beyond bowling.
There is a stark contrast between agile development and actual software engineering. Agile works where the customer and the developer don't really know what is being developed.
And as much as he has lectured about software development, I don't get the feeling that he has actually developed a substantial quantity of serious software.
In one of his clean coder books, he argues that Java is not object oriented. Unclear if that is a widely-held opinion. If you look for serious high-quality software, they are not agile. Such as the software for the space shuttle, or qmail, or other formally-verified software.
I have a friend who worked with him and it was sort of Uncle Bob's way or the highway.
I think notebooks are useful (though only for exploratory analysis), but there has to be a better way.
Which gets back to the posted article. REPL-driven development permitted him to feel confident without doing his usual process, it didn't discourage decoupling, but it didn't encourage decoupling either. TDD, to be done well, almost requires decoupling your code to make it more testable.
The same may (I'm not sure I agree, but I've never used Jupyter notebooks) be true of using notebook-styled development approaches and modularity. If the workflow permits relatively easy (and mostly safe) development without encouraging modularity, then it's incumbent on the developer to bring that discipline into the workflow.
I used this methodology with a Slack bot, because doing integration testing with Slack is annoying when I'm not directly testing the Slack communication portion.
For example: https://github.com/cjsaylor/chessbot/blob/master/cmd/repl/ma...
Edit: It seems like there are a lot of comments surrounding him "abandoning" TDD. Not only did he not do that, he literally says in the conclusion of the article that he would not recommend doing it.
I think you need to have a language (or structure your particular program in a way) that is really made with live-coding in mind for it to actually give you any benefits, like Smalltalk or Clojure and similar. Otherwise it's just adding another step to reach the real program you're running.
Edit: using your example as an example here, hope you don't mind. A program made with REPL driven development would have way less code in the main function, in order for you to test the code outside with your REPL. The `StoreGame` function would accept a data-store + the argument you have now, so at development-time, you can pass it arbitrary stores. And similar changes.
You can still get proper editor support with Golang (even though it isn't designed for it). I've done so here with VSCode working with a "remote" delve session: https://github.com/cjsaylor/chessbot/blob/11e1059aa77fed84d2...
My point being, if you wanted to execute the following code: https://github.com/cjsaylor/chessbot/blob/11e1059aa77fed84d2...
You probably wouldn't be able to, without restructuring how you're handling your state that is being used there. A REPL driven made program usually is made that the runtime carries the state and exposes it, so you could just select that part and run it, and see the output of it.
I'm guessing delve is more like a traditional debugger that steps through each line, rather than an REPL, correct?
I may be misunderstanding the "true" REPL development here, but I think the points laid out by the article seems to match my experience to a degree.
It seems like a ripe target for an opensource library to accomplish this. I may explore this later on when I have more time to devote.
When everything's well again, I move the bits to a test case.
And you don't need to extract anything into functions, lisp-syntax kind of already gives you all the separation you need, so you can just execute code in place if you need it to.
And if you're doing both of those things above, then you almost certainly run some or all of your tests WHILE doing REPL-based development from your always-on REPL. In fact, that's one of the best ways since you don't suffer the JVM startup cost.
Why not compose with the repl, then write tests at the end?
The common fear hear is that “you be motivated to write the test, or the test will be less good”
For motivation I never found it to be a big issue, and for the quality of the test, writing it after often produces better results. You get more time to think about _what and how_ you want to test
In the long run it's not faster. However, the feature he was using REPL driven development for did ship faster. That's a valid tradeoff to make, as long as you're doing it consciously.
Isn't this just...using a computer?
I don't know; maybe I missed something.
For example
* Compile run debug. You write some code potentially with tests. Run it and either poke it manually or run tests to see what it did. If it didn't do what was expected you either edit it, write some more tests, add print statements, add break points or whatever and try again.
You are forced to create all state in one go and pay one complete edit build debug cycle for each thing you want to modify in isolation. This could be 30 seconds or hours.
* Minimal interactive development. You mostly write code as above but gain the ability to poke individual functions or pieces of functionality in order to aid your understanding when things don't work correctly. You write tests eventually or not.
* Repl driven development. You mostly write and interact with a live environment either by actually typing the code in a repl or writing it into a source code file and hitting a shortcut to evaluate as you make changes redefining it piece by piece in terms of the rest of the code and the state you have built up by prior evaluations. This persistent state and the ability to redefine parts of it are what separates it from writing some python and occasionally opening up ipython to poke the code.
Being able to poke and prod at a system is priceless for debugging and introspection.
But TDD is not design either, in my opinion.