The Hardest Program I've Ever Written (2015)
journal.stuffwithstuff.com
journal.stuffwithstuff.com
This reminds me of the Tao of Programming 3.3, at http://www.mit.edu/~xela/tao.html , the relevant part of which I will copy here:
There was once a programmer who was attached to the court of the warlord of Wu. The warlord asked the programmer: "Which is easier to design: an accounting package or an operating system?"
"An operating system," replied the programmer.
The warlord uttered an exclamation of disbelief. "Surely an accounting package is trivial next to the complexity of an operating system," he said.
"Not so," said the programmer, "When designing an accounting package, the programmer operates as a mediator between people having different ideas: how it must operate, how its reports must appear, and how it must conform to the tax laws. By contrast, an operating system is not limited by outside appearances. When designing an operating system, the programmer seeks the simplest harmony between machine and ideas. This is why an operating system is easier to design."
I always disagree with the people that dismiss "CRUD-apps" or languages/tools made for "simply (ja!) crud code!".
I do CRUD for life. Is NOT simply. Ever.
Sometimes, for relax, I toy with language design. Maybe I must do some assembler or APL. Certainly easier that try to implement the crazy scheme of "discounts" of my customers!
Occasionally the owner brings up that maybe we should just write our own POS, to which I very quickly back-pedal and let him know that's not a good idea. As much as the actual inventory tracking isn't that hard compared to what we've done, the accounting portion is not something I want to touch with a ten-foot pole. Let someone else make sure the accounting reports we do our taxes from don't have odd bugs, and spread the gains of bug-testing and fixing across thousands of customers. I've already got plenty on my plate before trying to take on a massive project with huge downsides like that.
> Certainly easier that try to implement the crazy scheme of "discounts" of my customers!
Any billing system that can correctly account for all the various ways businesses and sales departments want to display, link, sell and promote services over time quickly approaches a complexity which is hard to to explain, and without conscious effort also quickly becomes the domain of a few select people that have bothered to actually dive into the mess, leading to silo'd expertise.
Additionally, if you try to make it easy for non-programmers to use, you quickly end up with a DSL which while it may get the job done admirably, is necessarily powerful and expressive enough that you end up having pseudo programmers to write and read them anyway.
Yes! Now I truly converting that "toy with language design" to actually "maybe use it for my apps!".
P.D: And yes, I know is good to use any lang already made, like Lua or TCL or similar. But CRUD-apps are not very well served by most languages today. If wanna know what kind of language is almost ideal check it FoxPro (or the dbase family)
No lie, I've gotten more than one good book recommendation (including "Tao of Programming") from '/usr/games/fortune'. Also, if you find a book fascinating or highly informative, skim the bibliography for more like it.
http://www.catb.org/jargon/html/koans.html
I especially like the power-cycle one, due to spending plenty of time in a lab debugging circuits.
I've been using an implementation of it [2] with a lot of success for pretty printing Rust code [3].
[0]: https://homepages.inf.ed.ac.uk/wadler/papers/prettier/prettier.pdf
[1]: http://belle.sourceforge.net/doc/hughes95design.pdf
[2]: https://hackage.haskell.org/package/prettyprinter
[3]: https://github.com/harpocrates/language-rustAutomatic code formatters don't need to be perfect or complete; they simply must format good code well. If bad code (as many of the difficult examples were) formats poorly, that's just another reason for people to write better code.
But one of the biggest use cases of a code formater is for understanding bad code. If it can't format bad code, that use case is gone.
More recently I've been trying to learn Chinese, and one of the features of Pingtype is to put spaces between words.
To my surprise, this article linked to a Wikipedia page about line wrapping, which says that line wrapping in CJK is unsolved.
https://en.wikipedia.org/wiki/Line_wrap_and_word_wrap#Word_w...
"Most existing word processors and typesetting software cannot handle either [personal names or compound words]."
My method works, but I don't know who to give it to. This is Hacker News and there's people from all different backgrounds here, so I'll just throw it out there - if anyone is interested, please contact me.
This is not how things are done. Just put your code and documentation on the public Web, then submit the link to Hackernews or other relevant forums.
If you want to be successful in communicating your improved method for word segmentation to the world at large, extract the code, publish it, and document the relevant details what makes it better and how it compares against existing solutions.
I am a hobby computational linguist. "Contact me" is an outmoded concept, most of us are doing our collaboration in the open.
I would simply indent all chained function calls and be done with it. Eg.
return foo(param)
.then(bar)
.catch(err => {
logger.error(err);
return -1;
});On the other hand, Go's style guide doesn't have a line limit, so its formatter doesn't have to solve tricky problems like this. Programmers can add line breaks by hand if they get too long. (A nice side effect is that auto-formatting your code doesn't change line numbers.)
- I think most of them doesn't recognize multi line string literals, which is difficult if you consider the case that you can have "" in comments, comment symbol in "" string literals and line breaks. The only way to deal with it is to scan linearly with context.
- It's tricky to wrap a long line: + some points in a long line are more suitable for breaking points in logic level + but sometimes you want less lines and not to break too often, even it's more clear in logic. The lines could be just some parameter list that will be both well represented in one column or multi columns. + with nested code the natural indent position could be at the far right, which make each line very short if you stick to 80 columns rule.
After quite some efforts my code can deal with all the comments, multi line strings, all the operators I known (I need to separate unary and binary operators to determine whether to insert space), but the script take several seconds to run, and I haven't start to deal with indent. I probably can save some time if I do more optimization, but I don't have time to finish it now.
This python formatter talked about its algorithm, worth a read.
And if it's not fast enough, you can use one of a number of relaxations to a standard LP and call it a day.
This is why the language should be designed with a formatter in mind from the beginning, as Go was designed. Just enough mustaches to make formatters accurate and fast. How many possibilities should there be? Exactly one.
Sounds like this might be a good candidate for some AI methods which are not intimidated by such large search spaces.
"If the line splitter tries, like, 5,000 solutions and still hasn't found a winner yet, it just picks the best it found so far and bails."
This sounds like it would give no more guarantees of a deterministic result than AI methods would. They also try their best, only search a subset of the search space, and bail when user-defined criteria such as time, iteration, or space limits are reached.
Another approach could be to use both the author's hand-crafted algorithm and an AI method. There are at least two ways to go about that. The first would be to run the hand-crafted algorithm and pass its evaluation of the best solutions to seed an AI method (for example, as the seed populations in an evolutionary algorithm). The second would be to do the opposite, and run the AI method first, and seed the hand-crafted algorithm with the best solutions the AI method discovered.
Given that, I don't see any reason why the algorithm should be non-deterministic just because it gives up after a set number of iterations. As long every time it sees a given string it tries the possibilities in the same order, scores them the same way, and stops at the same place then it should be deterministic.
You're right that if the algorithm described in the article will always get the same result given the same input it will be deterministic in a way that the kinds of AI methods I'm thinking of (the stochastic kind) will not. That sort of non-determinism will be a problem if you expect your code to be formatted the same way every time it goes through the formatter.
For example, one could use a subset of the Linux kernel source code as training, and another (unseen) subset as testing.
Highly rated projects on github might work too.
Even if training the splitting policy from scratch is too hard, learning would be a good way to establish cost weights on different kinds of rule violations.
ninja edit: I mostly jump remembered the picture of Robert with his/a dog and the text, “Hi! I'm Bob Nystrom, the one on the left.”
Naive question here: What is so hard? It can be solved with dynamic programming, right? Doesn't he even link to solutions of the problem?