Mu: making programs easier to understand in the large
github.com
github.com
+ recording runs as tests
+ one library for all layers of abstraction
+ building in assembler to avoid the big runtime.
+ using literate programming techniques to manage coding in assembly.
+ scenario, "screen-should-contain" and "assume-keyboard". awesome.
+ spaces: nice primitive of the closure and similar concepts
+ attributes for meta programs. again, awesome.
+ labels and using [,] as labels. genius.
Suggestions:
* If there's a valid reason to call them recipes, go ahead. If they're actully just functions, please stick with what everyone knows and gets. Same thing with "reagents","reply", "ingredients" and "products". Update: I did read your blog post about this. still think the regular names are better in the long run.
* I dont know where you're going to with multiple types of the "number:list" kind, but if so, the map must look the same, not lispy.
In all, I really like where you're going with this. Kudos! I find a lot of resonance of ideas I've had for a long time now.
The terminology is definitely a work in progress. It's mostly for my attempts to teach programming using Mu. I noticed that mathematical words intimidated some students. I tried to write up my rationale here: http://akkartik.name/post/mu. But I've certainly started mixing up the terms with my student, so it might not last.
I can see how the use of the term "arguments" could be confusing ("parameters" or "inputs" would make more sense), and "threads" is a rather tenuous metaphor for how scheduling works within a kernel, but all of the others mirror their real-world meaning pretty well. I'd be more worried about students needing to unlearn "containers", "ingredients", and "reagents" when they start reading material from outside of Mu, talking to other developers, or learning calculus (which uses functions, sets, and arrays).
Functions are what something is for, as opposed to form. It isn't natural to think about their inputs and outputs.
It's quite possible my solutions are too blunt and problematic, but these seem like real problems.
Interesting, as I tried a quick brainstorm, I thought of a group of kids sitting together playing with the same toys. Actions that only one could do on a toy at the same time. Trying to convey the problems or exclusivity. Mentally came to problem of sharing. Then that people often borrowed toys temporarily then the owner checked up on them.
(lightbulb) Rust has a borrow-checker. And ownership. Now I wonder where they came up with that haha.
In Ada, "tasks" makes sense, but only because it is an abstraction that makes use of threads while managing all memory access and separation from the rest of the process.
How you like THAT! Maybe need a similar analogy closer to whats going on but the spirit of it seems accurate. Maybe factory workers on assembly lines or offics workers at desks.
Procedure however might be a good enough name, less connoted than Recipe.
Quoting the Elements of Programming by Alexander Stepanov:
"A Procedure is a sequence of instructions that modifies the state of some objects; it may also construct or destroy objects."
With the definition of Object as:
"An Object is a representation of a concrete entity as a value in memory"
As for functions, I think the term makes sense, even with the non-mathematical definition, and especially in the context of OOP:
> an activity or purpose natural to or intended for a person or thing. synonyms: purpose, task, use, role
The "purpose" part doesn't make much sense with the way that we use the word, but it's certainly an activity. For example, a function of a stove is boiling water, its input would be liquid water, and its output would be steam & hot water. In OOP, "stove" would be the name of the class, and "boil_water" would be the function.
From that observation, I think one should conclude that using analogies from cooking, as you do with recipe in the kitchen sense is not the best idea.
https://gist.github.com/uucidl/df81b6ab0fe65713f5ba
My notes:
- creating an array w/ create-array is possible, however I could not find how to take its address and pass it around
- one cannot get the address of an invalid position of the array, so one cannot construct bounded ranges (STL style) using a begin/end pair of addresses
- the interpreter strongly suggests to use refcounted pointers (address:shared) instead of addresses even when one does not borrow the memory
- could not figure out how to write a test scenario that uses checks named memory addresses rather than ordinal memory addresses (memory-should-contain instruction)
memory-should-contain doesn't currently support named locations, sorry. Part of the problem is that with spaces a name can have many different addresses in different functions. So my tests write stuff I want to check in raw numbered locations, and the first 1000 addresses are reserved for tests so that names can never be clobbered.
I looked at Rust's borrow checker for a bit but wasn't smart enough to understand how it works or transplant it easily. I also noticed that it wasn't smart enough to deal with things like doubly-linked lists without reaching for ref-counting. So I figured I'd keep things simple and just use ref-counting for everything. That way I punt on all the complicated static checks in favor of a simple runtime one.
The rule is: new returns shared:address, and get-address and index-address (and maybe-convert) return address. Use shared:address to pass things around between functions, and reserve non-shared addresses only for short-term operations, usually mutations. Since non-shared addresses are not dynamically allocated, there's no possibility of use-after-free so they don't need to be refcounted, and you can copy them around as much as you like.
Use-after-free and related memory corruption is really the only thing I'm concerned about protecting my users from. Memory leaks I plan to have tools for, so that programmers can identify and break them down when memory becomes a concern, but not worry about until then. I call this "zero-developer-cost abstractions" :) It feels less restrictive and more dynamic than Rust.
Edit: I just took a look at your code, and it feels perfectly idiomatic. Nice job. Only issue I found was that you forgot to specify the outputs of find in its header. So its calls end up doing some runtime type-checking. I should probably raise a warning in this situation. Thanks again.
I like that you have users and thus will be able to test things out and get some genuine feedback.
This might however introduce a bias guiding your language and standard library in a certain direction (I think it's inevitable and what's happening in all languages anyway) .. Therefore I would add some of your peers into the mix.
Otherwise I really like the directions you are exploring, as I've come to very similar conclusions at this point in my career: - being able to modify a system through safe additions - not focusing too much on local details of a system
I'm not sure what I'm looking at?
It's a new programming language.
edit: The source seems to be structured intentionally in a recommended reading order. See the comment at the top of
https://github.com/akkartik/mu/blob/master/000organization.c...
Layers of code are filled in as you read down the list of files. It's an interesting concept, similar to literate programming. (And FWIW, these source files are C++ code for the compiler, not an example of Mu code.)
The author also has a blog post up here: http://akkartik.name/post/mu
More details on my flavor of literate programming: http://akkartik.name/post/wart-layers
a) It makes the build scripts more complicated, which means they'll be more likely to break on some poor noob, and that when they break it'll be less likely the noob will be able to tell how to fix it.
b) Invariably the codebase accumulates dependencies between directories that are uneconomic to reorganize. At least a flat directory has a shot at under-promising and over-delivering.
c) Maybe people do it to make the place look neat, they way I used to 'clean my room' by stuffing all my dirty clothes into drawers. But codebases that are messes at a deep level are less likely to be cleaned up if they look clean at a superficial level. And all codebases eventually turn into messes, the way we've done things so far.
But they can also be used well, and the promise of a meaningful directory structure is really appealing.
Here's a question, given all of what I understand about your layered way of structuring code: how easy is it to spin off logical portions of a project? If my project accretes its own HTTP server and I want to turn that into a stand alone library, could I? Would it be possible to "rewrite history" so that the HTTP-related changes never appeared in the layers to start with?
Thanks!
$ ./mu factorial.mu
To run a more complex app, you just give a directory name rather than a filename: $ ./mu edit
Files inside the directory are loaded in numeric order just like at the top level. You can also run just a subset of the layers for the editor: $ ./mu test edit/001*
$ ./mu test edit/00[12]*
$ ./mu test edit/00[1-3]*
(Notice how the number of tests/dots grows at each layer.)Since the layers are just regular files there's nothing stopping you from rewriting history as much as you want. That ability was precisely what I built them for.
I've only recently started using directories, so I'm sure there's stuff here I haven't considered. Feedback most welcome.
So, no, there's a difference. I'll also note that Daniel Bernstein of qmail & NaCl often used a filesystem how some use databases under the philosophy of "Why create new, complex functionality when you have well-tested code that gets the job done?"
You're right that there are points beyond which a single file becomes too unwieldy, but with a good editor that limit is quite high for me. Maybe 10k LoC. Directories can grow even larger before they start having problems.
Designing a programming language is kind of like a 'rite of passage.' I think most of us have at least thought about how to make languages better.
Thanks for your kind words! I'm gratified that somebody understood what I'm trying to do in spite of my crappy writing skills.
I like your idea, and I think there's room for improved introspection in debugging.
At the same time, I think that if a program can't be understood without being executed, then it's already lost (in terms of readability).
Certainly, for example, if you have a threaded program, you need to be able to visually inspect and verify that there will be no deadlocks or race conditions when the program is run, because you won't be able to test every possible condition in a debugger.
(Mu scenarios can't insert context switches yet, but it's planned.)
Perhaps it would help to think of the dichotomy as between the rules and the state space of inputs they handle, rather than between reading and running. Seemingly simple code can often hide surprising subtleties. Why is this line written just like so and not thus? How does everything turn out just right in this one situation? Tests help to record the right questions for the reader to ask.
> I can't imagine visually inspecting and verifying anything if I didn't write it to begin with.
It's a skill, you need to develop it. You develop it by doing it :) > I wouldn't know where to begin thinking of possible ways to break it.
For deadlocks, look at every single lock used. Show that one of the Coffman conditions doesn't apply, and you've proven that there can never be a deadlock.For race conditions, look at every piece of shared memory (or other shared resource). That is where race conditions always occur.
For understanding a program, the key is to understand the structure. That is why people put things in subdirectories, to help communicate the structure of the program, and show which things are closely related.
> Do you really look at every single piece of shared memory?
Yes > When you inherit a large codebase from someone?
If the codebase has lots of threading errors, then yes, it's the only way. If it's a large program, it can take months to go through and check every piece of shared memory (and removing threads along the way, removing shared stuff); but the alternative could take years.In one case, I moved every single lock/unlock to the top of a file so I could quickly see all the locks and unlocks, and that each lock had a matching unlock, even in error conditions.
I think I'm not concerned so much about inheriting a codebase with lots of threading errors. It's more about a codebase that's almost perfectly right, but where I'm scared to change anything because I don't have a big-picture model of the concurrency in my head..
Do you have any code samples I can try to learn from? For example, I'm not sure how you would move lock/unlock pairs to the top of a file from different functions/scopes. Unless you were doing some sort of literate programming as well?
> How did you get into such issues? Were there classes where you got started, or was it all on the job?
I took the usual college classes dealing with concurrency (in my school it was in the OS class), but it took several years in the industry to really feel confident with threads (it took less time to feel confident with networking). I wrote down my knowledge (for what it's worth) in this book: http://www.amazon.com/dp/0996193308 > I'm not sure how you would move lock/unlock pairs to the top of a file from different functions/scopes.
That solution might not work in every case. It worked in the particular case I was referring to. > It's more about a codebase that's almost perfectly right, but where I'm scared to change anything because I don't have a big-picture model of the concurrency in my head..
Hmmmmm, that's an interesting question. Usually when the codebase is written, the person who wrote it had an idea in his head that, "this is how things will be locked to avoid problems." I try to figure out what that idea was.Sometimes there is like a critical 'zone,' where a thread acquires the lock when it enters, and releases when it leaves. For example, it could acquire the lock when it enters a class method, and releases it when it leaves the class. Then the class becomes the critical zone.
Maybe learning to think of 'critical zones' is the most important skill to understanding the big picture?
I'm glad I kept the conversation going so I could find out about your book. Purchased.
> n * n-1 # does what you expect
Very good idea!
[1] Colorized version of https://github.com/akkartik/mu/blob/379244666c/010vm.cc
Could you elaborate on the target audience for the doc you linked to? I am unable to make out whether that doc is intended for serious programmers in other languages coming across Mu for the first time or whether its intended for new comers to programming.
I can already feel my carpal tunnel acting up!
I actually don't consider Mu to be a language. (I'm the author.) It's a low-level starting point to explore ways to make the standard OS primitives more testable. I don't need it to look nice. Higher level languages can come later, once the foundations are in place.
I actually am growing increasingly suspicious of all our navel-gazing about syntax. We spend all our time in places like HN thinking about how to make small screenfuls of code look nice, but all that effort hasn't really translated into helping me understand the large-scale structure of random open-source codebases more easily.
Perhaps one way to think about it is that syntax helps insiders keep a codebase in their head. I'm more concerned with ways to help more outsiders import a codebase in the first place. Because people leave and move on, and it's very hard for software projects today to improve once their original authors move on. More on this: http://akkartik.name/post/readable-bad
(I'm susceptible to carpal tunnel, but it hasn't become a bigger issue. I take wrist breaks.)
[1] I can even imagine doing the translation in machine code, so that Mu then becomes self-hosting without needing any of the current stack. I might go down that road..
This. The pain point that mu addresses is huge. I haven't really looked at the project in much depth yet, or played around with it, but I love that you are tackling this issue. I have spent way too much time wandering around a big code base trying to build mental models of what's going on, and thinking to myself the entire time "there's got to be a better way."
You've effectively created a sort of portable assembly language, but it's more verbose than usual assembly language ("sub eax, 3").
but all that effort hasn't really translated into helping me understand the large-scale structure of random open-source codebases more easily.
I think the root of the problem is that software is being written to be more complex than it could/should be, so the solution is to encourage reducing complexity.
Exactly. The verbosity is mostly because of my teaching project. I think teaching assembly can be just as ergonomic as teaching a high-level language. Indeed, most of us from a generation ago learned programming using assembly. It just needs to get a little more ergonomic.
the solution is to encourage reducing complexity.
Yup. The problem is that human beings have a _terrible_ track record at managing complexity. It's not just every software project ever; think about the creep of bureaucracy in ancient China, or the creep of legislation in ancient Rome. Everytime we've created a repository of rules it's gotten complex (and then gamed by smart operators). I think the problem is that it's very hard to justify removing a rule, so such repositories grow monotonically. The only way to periodically prune unnecessary rules is to first track why you created them in the first place (https://en.wikipedia.org/wiki/Wikipedia:Chesterton's_fence). Hence: tests! I think they are the great advance bequeathed to the human race by software. And they're far more broadly applicable outside software -- though I have no idea how to do this applying. I wrote up some speculative ideas about it at http://www.ribbonfarm.com/2014/04/09/the-legibility-tradeoff a couple of years ago, but that piece isn't very clear.
substract a 3
or even: (- a 3)Programming languages (at least the not-so-academic kind) need to mesh with the problem-spaces being solved, and there are definitely some contexts where most of the work is encoding and applying arithmetic rules.
Anything more than that is for wimps. Brain is the only language I ever program in.
> [...] think of a tiny improvement to a program you use, clone its sources, orient yourself on its organization and make your tiny improvement, all in a single afternoon.
If we alter that to "three programs you use [...] all in a single afternoon", we can plonk that into my resume.