Build Your Own Text Editor
viewsourcecode.org
viewsourcecode.org
Having something that outlines the key features and components and which ignores the important but complicated edge cases assists in keeping the attention focused.
Now if there are annotation within the source code, that would be truely incredible.
It is Windows specific but very informative.
Corrode is absolutely incredible. This file is literate Haskell, which means there's more documentation than code (I guess), and it transforms C into Rust.
That's the intent, but to the compiler, the meaning of literate Haskell is that comments are the default, and only lines starting with > contain code
By the way, is there are similar resource for building your own relational SQL-based database?
http://aosabook.org/en/index.html
(The books are CC and free to read on-line :))
I'm just through the few first chapters of the first book, and I must say it's absolutely amazing. Each chapter gives some understanding of the thought process people designing (and iterating on) a known open-source project had.
Yes. This is exactly the kind of thing I was once looking for, only I didn't know how best to phrase at the time. Think I've since found some of the details I needed then, but will look into these nonetheless and see what I missed, if anything.
Thanks!
As long as the first commit isn't something like 'import to git' or 'add the code' (which seem to be tragically frequent) I find VCS a huge help here. The problem is that large OSS applications are (tautalogically) large. VCS allow cutting it back, and showing the evolution.
https://github.com/DigitalMars/me
and translated to D:
The first thing I do with a new system is port ME to it. I'll use vi to do what's necessary to get there (usually get ssh working so I can edit the files with ME on another machine).
I've worked out how to do syntax highlighting on it, but never got around to doing it.
Back in the 80s, I handled configuration by having ME directly patch the ME executable. (This was a trick I learned from the old ADVENT Fortran game.) It was marvelously simple and bulletproof to do that.
Unfortunately, programs that patched their own executables became huge no-nos as malware took off, and I had to abandon that.
Mine was a matter of grouping the global variables for configuration together. Take the address of it, compute the offset of that address in the .exe file, and write.
How far did you end up getting with the C editor? :)
Another useful resource I've relied upon in the past, dates back to the 1990s: Freyja, which is Craig Finseth's emacs-like editor written in C.
Here is a list of features:
* deletions are automatically saved into a "kill buffer"
* ability to edit up to 11 files at once
* ability to view two independent windows at once
* integrated help facility
* integrated menu facility, with help on all commands
* can record and play back keyboard macros
* supports file completion and limited directory operations
* includes a fully-integrated RPN type calculator
It was designed for MS-DOS with the Cygwin terminal library.I found the architecture to be very clean, and it is well explained in Finseth's classic book ("The Craft of Text Editing"). The book is worth reading even if you never touch the code: http://www.finseth.com/craft/
It uses a multi-buffer architecture roughly similar to Walter Bright's text editor (see sibling posting). (I knew about Finseth's editor years ago, but was not aware of Bright's work until now, thanks Walter!)
The Freyja source can be downloaded from: http://www.finseth.com/parts/freyja.php
Look for the link to "freyja30.exe" which turns out to be freyja30.zip (not an executable).
In hindsight, I would have a bit more caution programming text editors. I started tweaking and modifying Scite years ago, it was very interesting but it was no small undertaking and I came to understand why Neil advised in the support forum, to customise it using the inbuilt Lua scripting. Im still using this 6 year old customised version of Scite that I never managed to sync with the latest version, and it has 10 thousand lines of custom Lua facilities like file encryption, navigation panels, multi-edit mode etc.. which I wrote and stabilised a few years ago. I rarely venture to alter it now that I am at last comfortable with it, but its going to need serious attention sooner or later...
... it sounds like you kind of wanted emacs. One of the most impressive things I find about emacs (especially since semi-proper packages became a thing) is just how easy it is to get stuff that is 5-10-15 years old working on it. No word of a lie, it's amazing how they've managed to break so little over the years.
I really must get round to sharing the source but it needs a bit of preparation.
I forked antizez's editor to add Lua support, and posted it here in the past:
That works well for me, even with the minor omissions, and is well-paired with my console-based mail-client - again scripted by lua:
The syntax highlighting functionality alone is a combination of something as complex as what a compiler/parser does, and you have to do it almost in real time.
So if you want to color every different part like operators, symbols, braces, with a different color, and you try to do this on C++, it is not going to be a small task...
Also something you learn about text editors is the rope data structure. https://en.wikipedia.org/wiki/Rope_(data_structure)
For a lot of modes, Emacs does this using regex.
Regex'ing languages like HTML or C++ can be quite mind breaking... I think.
Some of the standard features from codemax have to be rebuilt, but I'm still very impressed. Oh and before somebody plugs their favorite editor and why I should be using that instead, if it doesn't know DataFlex it won't help ;)
This looks like a great tutorial to work through with it :)
One other thing is that the {.compile.} pragma in term.nim is not needed anymore, as Nim's stdlib has added a lot of those in[0], but it does show how easy it is to bridge between the two languages (and I'm not much of a C developer!)
Under 50 lines for a text editor written in K by the language's author. Way beyond my present understanding, but the promise of very small, powerful code is incredibly attractive.
But what does it actually mean in terms of programming? How maintainable is "small code"? How readable is it? (Well, you answered that question already.) How hackable is it?
It's a curiosity, and a really fun one! But it's not practical.
But super-compact code could be very practical/readable/hackable, as long as there's the prerequisite knowledge. Fewer symbols means less complexity to parse. This works, provided that those symbols map to powerful operators that can be really understood and effectively combined to produce the desired outcome.
And there's no need to scroll: perceive everything in one glance!
Without trying to be too trite, what does this question mean in terms of English?
> How maintainable is "small code"?
Once you get the hang of it it's as maintainable as any other code base
> How readable is it? (Well, you answered that question already.)
There is a learning curve, but the code is actually reasonably readable. It takes time to get used to many operations occurring in one line but there are benefits (eg you can see everything that the CTRL-Z function for undo does at a glance).
> How hackable is it?
Again, I'm not sure what you're asking here.
The editor is very bare bones so isn't a great example of production code. In a real system generally people are a bit more verbose.
It's just that a large percentage of languages are quite similar in how they're structured. Since Pascal and at least until Java/C#, most mainstream languages ended up roughly doing "one thing" per line. Quite often one function call or simple mathematical operation. Then each function or block of a larger one does one larger thing. And so forth.
Code and/or languages that break that paradigm are often confusing to "switchers" and thus often abandoned or maligned way too early.
But most often, they're just scanned differently.
Assembly would be the opposite end. It basically takes "do one thing per line" to the extreme. But quite often, experienced asm programmers scan the program by blocks, as some patterns are quite common or you find some constant/string to attach your focus to, then continue from there into the details.
Forth programmers obviously read slightly differently, due to the high level of decomposition and the stack based nature (fewer parameters). Lisp looks a bit weirder at the first glance, but I wouldn't even say that it's read all that differently from the Algol family, if written imperatively enough.
Functional code, in almost any language, often has to be read differently, as a lot of things can happen in one line.
APL and some DSLs (regular expressions for example) are the opposite end of the spectrum. But line length would be the wrong axis to judge things, the amount of operations isn't necessarily that much smaller.
Data exchange (arrays vs. stacks vs. function parameters) is the bigger change, as is symbolic density (APL symbols vs J shorthand vs. function names vs. HiHowYouDoingIAmAnAbstractJavaBeansFactoryConsumerImplementationNiceToMeetYou)
1) A standard library / set of operators that are a very good fit for the typical domains it works on.
2) Minimizing symbol length.
E.g. I translated one example I saw into Ruby, and ended up with something of similar length once I 1) implemented equivalent methods, 2) dispensed with all idioms for how to write Ruby and went for single character variable names and method names etc.
In other words, there's nothing particularly "magic" there.
K code gets to where it is largely by because its author and users are willing to violate every convention from other languages in terms of how to write and structure code in pursuit of a philosophy that is fundamentally different in terms of e.g. focusing more on code size.
That could be a good thing, but I'm not convinced that the extremely sparse code is worth the (to me at least) extreme lack in readability.
Code golf can be fun in any language, but it seems few of us have gone back to writing other code and decided it's worth aiming for code that small. In fact, I've more than once rewritten code to be longer because it made it easier to read.
That said, there are good parts in terms of language constructs etc. that'd be worth learning from. If only it wasn't so incredibly annoying to decipher the code (yes, I'm sure it gets faster when you get used to it).
understatement of the year.
I've been messing around with a little text editor component myself in golang:
https://github.com/briansteffens/tui/blob/master/editbox.go
For a SQL editor project:
It is surprisingly tricky to get right. There used to be a POSIX function in the c standard library - "getpass" to handle this but the implementation was not thread safe and thus it was depreciated from posix spec - the only portable way I know to do this is to "roll your own code" using termios - not ideal.
see: http://man7.org/linux/man-pages/man3/getpass.3.html note the comment: "This function is obsolete. Do not use it" - there is a heap of potentially vulnerable code out in the wild that uses getpass.
I've often thought of coding one for fun, with no intention to share it, just for the purpose of having a long-term project that evolves along with my skills. I've never made time for it, but I still consider it once in a while.
I actually need to get off my duff and code what I really want, which is a version of what I remember the best parts of the Norton Editor implemented as an emacs mode. I have great memories of writing pascal code in ne.com, so it is a big nostalgia item for me.
Interesting that you mention this. I loved the Norton Editor. I have wanted to for years add a more Norton feel to the text editor that I write.
I wonder if I can still get a copy to play around with and run in a virtual machine. Any idea?
Googling "norton editor manual pdf" lead me to a few old copies of the manual, so that is what I have used as a guide in my work, but most of it is just driven by how I remember it.
For coding though? I usually use emacs. It's just so damn customizable..
I have been careful to always carry a recent copy of my .emacs.d around with me on a thumb drive.
Example: I have an iPad version where I can use an Apple Pencil and handwrite my code on the screen. To me it is very useful when my wife is driving me someplace and I still want to work. Recline in the passenger seat and code.
It is working pretty well for me. I used OpenAI and basically started asking my friends to hand write samples for the NN to learn from. They write them various levels of neatness and various slants, font size, etc. I store the handwriting strokes as a series of points and convert to digital text when they are ready. I save the handwriting along side the digital version. It is possible to hand write, then come back days later and your work is still there. Maybe I could make a video to show if it is useful.
Committing to git currently works, push works, but sometimes says it fails. Still looking into this.
Color syntax highlighting works with your handwriting, but in the "worksheet" (an interactive shell) it does not yet.
It took me forever to unlearn the Wordstar key bindings.
I'll have to look it up.
Unlearn Wordstar key bindings? Isn't that sacrilege? :-)
When my computer illiterate aunt decided to write a book back in the early 90s, I set her up with a minimalistic QEdit, a few bat files to perform backups, versioning, etc automatically and a floppy disk for each day of the week plus a daily backup one. Simple instructions and process, simple editing setup, and two years later the 670-page book was finished and published, and she was ready to actually learn how to use a computer.
So a big thank you for creating such an excellent tool!
I've used that editor daily to write source code, mainly in Clipper, and I think some of my coworkers used that too.
But then I moved on and knew vi, and now I can't take my hands from the center row :)
ProgEdit: https://sourceforge.net/projects/cpstools/
SynEdit, the base of the editor: https://github.com/SynEdit/SynEdit (confession: I was only a user of SynEdit, not a developer).
Don't do it.
bool buffer_prepend0(Buffer buf, const char data) { return buffer_prepend(buf, data, strlen(data) + (buf->len == 0)); }
Two function calls plus arithmetics plus step-return.
Depending on the debugger interface(eclipse with gdb, vs, etc), one will have to press some form of Step 1 to 4 times to advance the line. Depending on the complier that was supposed to provide correct debug info (GCC vs msvcc vs xlc) and bugs, these steps may or may not work.
It is short, but when you do multi platform code, a real pain to deal with.
Consider aix and Linux multiplatform code for one of my past projects. My choice would have been: 1. Step and maybe hit a bug, msg the complier team, proceed writing register values on paper. 2. Break up the line, recompile. Few minutes on Linux, over half an hour on aix.
This is a pretty common way to write C, though, it's not something specific to this particular codebase. You just had a non-standard use case where you were constantly running into low-level bugs. (In an embedded platform?) If you aren't in that domain anymore, it's worth revisiting the trade-off. In most domains, gdb's continue command is super useful.
I will in few hours, thanks.
Great work.
I suppose that would be more along the lines of a hex editor, which is probably interesting to implement too.
Great site, great presentation, thanks for digging into it with so much detail.
But, if you've done any Java or Objective C, you can probably infer the gist of it without serious C study.
Hanson's C Interfaces and Implementations is another great book that would better prepare you for this.
What kind of data structure are you using for edits?
The guy who made this (Jeremy Ruten, I think) should be commended for taking something complicated and boiling it down into accessible terms.
I hope we see more documents like this on HN in the future.
https://github.com/antirez/kilo/blob/master/kilo.c
I wonder how mouse clicking and scrolling in editor works? Would love to make html/js/css for terminal node module. I think that would be fun.
But at least that is more of a LEGO style of putting different bricks (plugins) together.
$ echo 'cat > "$@"' >editor
$ chmod +x editor
$ GIT_EDITOR=./editor git commit -a
stuff
[master b2d3915] stuff
1 file changed, 1 insertion(+)
You'll need to hit enter before ^D.You can add simple Emacs-like keybindings by changing that into
$ echo 'rlwrap cat > "$@"' >editor
Or just learn ed: $ GIT_EDITOR=ed git commit --amend -a
256
1c
stuff and more stuff
.
w
271
[master c5092a6] stuff and more stuff
Date: Thu Apr 6 12:58:18 2017 +0200
1 file changed, 1 insertion(+)
(the numbers are ed telling me how much was read and written; `1c` means change the first line; `.` means I'm done inserting (go back to command mode) and `w` means write/save; exit with ^D) $ printf '#!/usr/bin/env sh\ncat > "$@"\n' > editor
But the second point is more valid. Ed is the standard editor. (global-set-key (kbd "C-d") 'save-buffers-kill-terminal)You may have a broken version of the font locally. If you download the offline version of the tutorial, it'll come with all the fonts it uses and hopefully will work.