Explorative Programming
blog.dziban.net
blog.dziban.net
I have a shell open almost all the time, and use it constantly for testing functions, writing queries etc. It drives me crazy watching other Devs change a query, wait for the server to reload and refresh the browser repeatedly, when editing in the REPL is so much faster.
Being able to `dir` and tab complete experimentally is so fast. (I also have this in the ipdb debugger... Also Django extensions has a debugger inside templates which is really nice)
Being able to `%edit 1-20` and have the repl session instantly in my editor to copy paste is fantastic too.
I also use a lot of TDD for tasks that fit that style, mostly where I'm writing business logic or data processing functions... but for exploring data / APIs /queries / data shaping, interactive just feels instantly productive. Once I've got the API or whatever it is to correctly produce the happy path, it's then easy to copy that into the codebase knowing what shape to expect, I can then turn that into unit tests and throw all the weird external data and edge cases at it with reasonable confidence about what should be happening underneath.
It also has a mode to print out all SQL statements it runs underneath your code - which is great for checking that you've not missed a join or something and got n+1 sneaking in
The REPL is so incredibly productive with immediate feedback that the long ORM expression actually runs the SQL query you were intending (or notice it does something you didn't intend).
Obviously I come from a lisp heavy culture, but I'm still stumped how many people stay in those high latency exploration mindset.
Software is made out of ideas, and code is just a symbolic representation of those ideas. Writing the code is the easy part; working out what ideas to represent is the hard part of programming. If your planning involves working out what ideas to represent, if you do that in greater and greater detail, eventually you stop planning the work and cross over into doing the work.
> File based organization, right from the start you need to start deciding what goes where. Of course I think most of us start with a single file, full of random functions and tests. But at some point this file becomes unbearable, impossible to find what you are looking for. At that point you have to make decisions and start splitting the file at logical points. Deal with what is private and what is not. It may not seem like much, but it's very annoying.
In a project I'm working on at home, I started off with an extremely naive implementation, the single file of functions described above, and when the messiness became unbearable, I started building object-oriented abstractions to make sense of it all. But the structure of those abstractions was discovered while making the mess - if I'd tried to build that at the start, I'd have built the wrong thing.
I'm very fortunate in that in both my day job and my recreational programming, I'm the only person working on that particular codebase at the time. (At work, I develop what are basically plugins for the framework-type-thing that's our main product.) Consequently, I have the freedom to make a mess and clean it up once I've muddled about in the solution space to discover what's going to work best, and so most of the solutions in the blog post are rather unnecessary for me. In the end, two months of work on my home project was thrown away and rewritten, which I could never have got away with at my job, but the project is all the better for it.
It’s almost like tending to something as it grows, discovering what the solution most wants to be with respect to the problem and the environment (e.g. programming language, server vs browser etc.)
I'm not disputing TDD as an effective strategy—I've seen it work wonders for some of my colleagues and I will still occasionally employ it myself when making additions/changes to well-established bits of code. But, for me personally, I feel I can stay in the flow and get all the same benefits by just doing step 1 ("Write a list of the test scenarios you want to cover"), implementing the code, then actually implementing those tests at the end.
You do you, though! As long as the code works and is easy to refactor, it’s all good.
[1] For more about writing sociable tests, see “Testing Without Mocks: A Pattern Language.” I’m on my phone but it’s easily googleable.
A better way to word it would be that I would have to rewrite both my code and the tests many times. When I'm in the exploratory phase, it's not just the implementation that's changing rapidly, it's the interface and structure of each piece too. I find it far more natural to discover as I go, even though that means I throw away a huge amount of the total code I write.
Consider the GP comment's example of having everything in one file to begin with, then splitting it out into separate pieces as becomes apparently appropriate. There's no unit test that would survive that process without equal levels of modification, even though the overall behaviour of all the pieces when combined never actually changes.
I suspect most developers have got a lot more of this stuff ironed out before they actually start writing code, but my brain doesn't work that way.
Thanks for the recommendation too, from the introduction that article looks really interesting, I'll be sure to check it out.
Also, if you haven't read it, check out Martin Fowler's Refactoring book. It shows how you can do things like splitting one file into multiple without rewriting.
What I do is test-drive the simplest implementation I can think of, which will often be one module or class, then factor out additional classes as I discover them. While I'm doing that, I'll leave my tests as is, and they continue to cover the code equally well. As I split things out, I'll typically discover more edge cases I want to support, and I'll test-drive those edge cases as well. I'll typically move any relevant tests at that time, possibly rewriting them to be more specific to the unit under test, but the important thing is that I don't have to, and I'll wait until I have the design figured out in my head before I do so.
I often feel like the code I write isn't something I came up with, just something I'm discovering as I go.
For me at least it's more like bang out the naive implementation mess then "discover" the proper structure of the abstractions in bed that night or in the shower the next morning.
It's a little bit of explanation of a project I'm doing as well as something I'd have liked to read years ago when I was sure I was just doing it wrong.
Would be cool if it targets WASM, this something I'm super interested right now, there are a lot of advantages on targeting WASM one of them is the interoperability with other components written in different languages, this will be when the component model[1] get's finished and implemented.
[0]http://unison-lang.org [1]https://component-model.bytecodealliance.org
It is actually one of the problems I have with common lisp, its impossible to do a semantic hash on arbitrary common lisp functions so I create more history that I need, since a reordering of independent let forms cannot be easily detected to preserve functionality
I've tried all of them, I think the closest thing I've seen to what you describe, which I also find very attractive, is the GT Smalltalk environment: https://gtoolkit.com/
Have you tried that? They call this idea "moldable development" as you can "mold" your environment to your needs.
Even though I loved it, I ended up not using it much, mostly because it's a bit too heavy to keep handy for exploration all the time when needed (it takes like 1GB of RAM even when idle!)... as I already can do most of that with emacs, which is much lighter, I just stick with it.
It's a simplified smalltalk (like pharo, forked from squeak), where the emphasis is on having a base system that you can reasonably fit in your head (think scheme vs common lisp).
Here you'll find much of what you described, including per-method revision history.
BTW, it comes with a rebuilt morphic gui with flawless scaling - unlike most smalltalk, this one actually looks crisp - even on large displays
I've always been searching for this mindset but I rarely read / hear about it.
Many points were the reasons why I am now using Julia to develop a large-scope full-stack application. Revise (package) allows to evolve the program interactively while preserving the system state. For instance, I can have an instance of a struct in the REPL; I change one of the modules, and the available behaviour of the instance can be immediately tried out without recreating the state. Or, when interactively developing a webserver, I find an error on one of the served pages. I changed the piece of code in the module, and it is immediately reflected when reloaded while keeping the existing state.
System introspection is easy as everything is available from the REPL. Default printing is available for all user-defined types, which can be specialised on a wish. Thus, printing an object or even getting to the precise point in the module call stack with an Infiltrator is always accessible.
One pain point regarding file-based organisation is that, in most languages, every file is within its namespace. In Julia, these things are orthogonal. You can have a single namespace organised over multiple files or multiple namespaces within a single file. This generally makes it easier to navigate a large namespace while not being bothered with making explicit imports, etc., when one is still in the exploratory phase.
I can’t really relate to tree-based exploration. Very few times do I find that the direction I have found is wrong, so I do a git reset to a previous commit. Also, there are a few sunk costs migrating from one approach to another when most of the functions are written to be pure.
Another essential piece for explorative development is types. Although some planning is required, it is generally worth it. In Julia, those are checked at runtime, which nevertheless significantly helps to locate issues quickly when doing test-driven development.
I think what the author describes is why Excel is so heavily used. It's often not easy to clearly define the specifications - but it's easy to play a bit with the data the user knows, and discover some flaws in the logic, corner cases etc. And also very easy to prototype processes. I almost miss it.
One of the ideas was to write code such that if you copy paste things into a different context, the syntax changes are minimal. It should be trivial to copy paste a block of code to turn it into a function. It should be trivial to turn a lambda into a full blown function. Type signatures should look the same whether they are function arguments or variable definitions etc.
Not advocating for Jai per se... But most languages are pretty bad at supporting this style, some are better.
The first and most important step is having a language without statements and only expressions.
But, as Jai needs a `return` in functions, it already failed the trivial test.
https://github.com/BSVino/JaiPrimer/blob/master/JaiPrimer.md
The only time I poke and play with the code is when I'm poking hardware to understand what it's doing, IOW when I'm playing with a relatively black box.
Writing new functionality is usually easy if it's self contained. If I know what inputs I will have and what outputs I want, then I generally have a pretty solid idea of how to derive one from the other before the first keystroke.
But if I find that the new functionality is turning into an internal API that will be consumed in multiple places for multiple use cases, things aren't so obvious anymore. You have a combinatorial explosion of design parameters: not just inputs, outputs, and algorithm, but also issues like ownership of data, ownership of the stack, which inputs are supplied at which point in the process, when is it legal to consume outputs and in what ways, push vs. pull interface, two functions for similar functionality vs. one with a flag. When the design parameter space gets this large, it's hard to reason about it ahead of time, and sometimes the most straightforward API doesn't reveal itself until you write code that uses it and realize that there's needless boilerplate or inconsistency that opens up room for logic errors. Then you go back and refactor to eliminate that. Casey Muratori calls this iterative process semantic compression [1]. I don't agree with all his software development takes, but I think he hits this one pretty squarely on the head. The key insight is that an experienced programmer can write good code on the first pass, but it usually takes a second pass to write great code.
The last time I did this, it was a completely novel framework for material simulations using boundary element method which happened to be my Ph.D. thesis. It did not completely turned out the way I wanted, but I was able to get the theoretical performance I targeted for (i.e. literally drowning the system in computation), nailed the data formats workflow we wanted to have, and it turned out to be pretty good. It can scale from laptops to high end servers, and on a typical laptop, it's 30x faster than a well tuned MATLAB implementation, while being at least as accurate. The formulae and algorithms we used for calculations were novel.
It was mostly designed on paper, with some whiteboard work, and implemented with zero memory leaks along the way, plus we hit the performance numbers we were hoping in second beta IIRC.
I'll shortly start on its second iteration, borrowing the high performance architecture, data types, but not the overall workflow. We'll try to make it even more user friendly.
So yeah, I do this during implementing APIs, libraries and novel programs.
However, your whiteboard and paper is my exploratory programming. I don't see it as it being that different.
I just visualize the problem on paper. You have this, you want this, and to go from A to B, we need to do this, etc.
Exploratory programming is nothing bad, or what I do is nothing good or exceptional. Humans' brains are wired differently, and as long as we arrive to the point we want, it's all good in my eyes.
There's no superiority in any of the methods. We're just cooking the same food differently. That's all.
It wasn't harsh, but it felt like being labeled as deluder, which I don't like a bit.
> (S)omeone hears of problem, gets coding immediately, and a very good result comes out.
No, that's not possible from my experience. You need that problem to ferment a little in your mind to understand the problem in its entirety, or more realistically to the point you know the domain. Experience makes the process faster, but not instant.
> We need a way of reasoning and investigating the problem...
Of course, this is why I exactly said that we need a way, and it can be any way, as long as it works.
Have a nice day :)
You are in a MUA composition windows and want to solve an equation? You have some bits that do so in your system, no need to switch "application", just feed the input to the relevant function and get locally the result. You just want to change something? Easy to do, it's just few bits of code.
This is the most effective and powerful use of a computer we have ever invented. It was dropped because no one who works in sw business want power users and easy to bend code, it's nature is collaborative, not commercial. It's nature allow people to understand and those who became "alphabets" do not like Alphabet's or Oracle's o Meta-information business and they can perfectly goes on without them.
How do you approach new problems?
While I haven't yet made this process very sophisticated - it's mostly string substitution, pretty prints and state machines - it has a natural way of tightening the screws on any specification, because each step presents challenges about which layer I want to use, and whether I want to break with locally idiomatic usage and take the approach of inlining something from above instead.
No, it works in a similar way for me. I can do abstractions to a certain degree in my head and plan ahead but also do need to work on a problem and trial things.
It's interesting that the article mentions Javascript as not having an immediate feedback cycle. Which is true enough for a lot of the modern Javascript stack. Although live reloading is also a thing depending on what aspect you are working on.
> Imagine a service that never stops, you diff the existing functions with the new functions, only updating what needs to be updated. If there is a change in the rate of errors then you automatically rollback the deployed functions.
I am pretty sure that this is already a concept in use by companies. Although the exact naming for it, I can't remember. It also highly depends on the sort of domain you are active in. The more critical your data is, the less likely it is that you want to risk using this type of development and testing model.
Not to mention that it seems lean in the logical fallacy of "only errors indicate faults" while the most devastating bugs often are those that don't cause clear errors. For example, when dealing with data your particular system might have no increased errors but due to a data change a system down the line might.
Another potential issue is that because you are developing this in an explorative manner, it also doesn't leave much room for a review process. Or any of the myriad of other processes and methods that normally should be in place to evaluate the thing under development.
In addition to the above, the way the article does describes the envisioned process leaves very little room for proper documentation. More specifically, it feels like a process where documentation will be more of an afterthought than it already is within many teams.
Off-topic for the subject, but using monospaced fonts for blog posts sucks for readability, imho.
The end remarks are more to let the imagination run a little bit free than anything particularly structured, the part about deployment doesnt really exist as it is in most companies, there is a concept of canary deployments, blue green, etc, but not on a running process ala Erlang or Elixir. Still, Im not sure if it is even a good idea since starting from a clean slate is usually good. I need to do more testing around this.
There is no interest in documentation while the process of discovering is going on, because the documentation would be stale as soon as you move on in the discovery process. Its not even a fault of the system, it just doesn't make sense. Once you find a good enough design, then you transition into settling it, documenting it, writing more thorough tests.
About the review process, this is mostly targeted at individuals and teams where there is a certain trust involved. But for knowledge spread that's why I also tackle doing this collaboratively, as I think it could even increase the effectiveness of this approach.
The design was pretty fast done with a template. I'll see if I have time today to test other fonts and see how it improves.
Thanks for taking the time to read
Personally I live digitally in Emacs, so I have some tooling for the present word, mostly hidden under Emacs or *nix system, but yes, it's still limited, less then Pharo since I can read my mails, org-attach my files and so on, but not integrate much more without all the modern integration issues of a an IT built to be isolated not integrated to sell programs before, services thereafter.
However I see a very slow and long trend: the more current IT evolve the more it tend back to the IT pioneer era, in the "modern past" all was widget-based GUIs, not even Microsoft recommend CLIs for certain works, most apps have WebUIs witch are limited DocUIs but with the same doc concept, NotebookUI are more and more widespread, here and there someone recommend LaTeX instead of WYSIWYG word processor, someone else recommend R instead of spreadsheets and so on. Perhaps in 50 years we will have Smalltalk, Lisp and so on again, at least at concept/paradigm level.
Mathematica also popularized the "notebook" style of dev which is also very nice for exploration.
If we stay on the Lisp side for a moment. Slighty different from Interlisp, then is Symbolics Genera. It's the OS and development environment of networked workstations for groups of Lisp programmers. The manual(s) has a chapter on the philosophy. I've put it online (the text is not by me, but from the manual): http://lispm.de/genera-concepts Then there is the idea to keep the developer "in the flow" while programming: http://lispm.de/symbolics-lisp-machine-ergonomics
Then there are the commercial environments, which have a similar development environment. Allegro CL and LispWorks. One marketing slogan there was: the cost of change to the software most only depends on the size of the change, not on the size of the software under development. Both have the advantage that they support application "delivery" on normal computers: Windows, Macs, UNIX, Linux, ...
Nowadays I think the combination of SBCL, SLIME + GNU Emacs (for the IDE), version controlled source code sites (Gitlab: https://common-lisp.net/project-intro ) , test suites, etc. makes for a very powerful development environment for groups. For example the SBCL implementation is a sizable and complex Lisp project itself, which is managed very well by the maintainers. There is a good build process and people have been porting SBCL, which is an optimizing native code compiler, to new platforms without much problems.
In the Lisp history there was always also the idea to apply exploratory development environments to other languages, improving on what various Lisp systems did. Example from the past, improving on Lucid CL: Energize, foundation for a C++ environment. https://dreamsongs.com/Files/Energize.pdf The development UI was Lucid Emacs, Objectstore was the code database, they had an incremental linker/compiler for C++.
This is what top-down programming is for (though, arguably, a lot easier for 'smaller' projects)
EDIT: oh sorry, didn't notice you asked specifically for JavaScript (not just Smalltalk in the browser).
- programming language not oriented around files
- PL compiler designed first and foremost for incremental complication / interpretation
- New source control that’s syntax-aware and watch not commit based
- New IDE that allows navigating the tree and space
And any of these are a reason for people to not use it:
- CommonLisp? Most people don’t like it. Could you adapt TS? Rust?
- If it’s not Git, can’t use GitHub
- Emacs? Could it be a VS Code plugin? A fork?
Where is a billionaire techie who could foot the bill to build this? It’s a mammoth effort.
The IDE was a hacked 50 line emacs major mode, it's really not that much work to get started. It has of course a lot of rough edges.
About the reasons not to use it, thats fine for me, for now this is for me and to serve as inspiration, I want to work with CL.
If anyone else wants to tacke it in another language, all the best for me :)
Git compatibility and output to files is something that the system should have, indeed