Asmrepl: REPL for x86 Assembly Language
github.com
github.com
Here's an anecdote I heard once about Minsky. He was showing a student how to use ITS to write a program. ITS was an unusual operating system in that the 'shell' was the DDT debugger. You ran programs by loading them into memory and jumping to the entry point. But you can also just start writing assembly code directly into memory from the DDT prompt. Minsky started with the null program. Obviously, it needs an entry point, so he defined a label for that. He then told the debugger to jump to that label. This immediately raised an error of there being no code at the jump target. So he wrote a few lines of code and restarted the jump instruction. This time it succeeded and the first few instructions were executed. When the debugger again halted, he looked at the register contents and wrote a few more lines. Again proceeding from where he left off he watched the program run the few more instructions. He developed the entire program by 'debugging' the null program.
Everything needs a catchy name, so I call it Debugger Driven Development.
I write the first few lines of a function, however much I feel sure about, then add a dummy statement at the end and set a breakpoint there. When the code stops at the breakpoint, I can see exactly what my code did and what data I now have on hand.
I use that new knowledge to write the next few lines, again until I get to something I'm unsure of or where I'd just like to get a better view of the data. Set a breakpoint there and view that new data.
Repeat as needed until the function is done.
At my work we have many internal APIs that are "documented" but the documentation is fairly lacking. With the debugger, I can see not only what the API claims to do, but what it really does with my actual input.
I am bummed that so many developers today eschew debuggers. I even read an article recently along the lines of "These famous programmers don't use debuggers, and you shouldn't either". Why would anyone want to talk people out of using such a useful tool? It makes no sense to me.
https://dev.to/kmruiz/working-with-your-running-program-an-i...
That would be because of what I call "long vs. short-term". If you're only thinking of the next few lines and doing that a lot, you will have effectively trained yourself out of looking at the bigger picture. As someone who taught programmers, I've seen what "debugger driven development" code looks like (because that's how some of them will try to start writing code.) It's not pretty. There's a reason a lot of highly productive (but not necessarily famous) programmers consider debuggers as a last-resort tool: writing code that needs debugging should be a rare occurrence.
From 30-year-old memory.
Yes, we actually had to do stuff like this. When I started my first job in the late 1980s, it was normal for hard disks to have a little label with a list of the (known) bad blocks on the drive. (Hand-written by someone at the factory on the first (circa 15-20MB) disks I used; later, on things like big ESDI disks, dot-matrix printed.) In some formatting tools, you had to manually enter them; formats took a long time, so eliminating retries on bad blocks could save half an hour.
Novell Netware came with its own low-level formatter called `COMPSURF`: COMPrehensive SURFace analysis. Dozens of people would be sharing a server's hard disk, so data losses would be extra-bad -- and might well bring down the server, losing everyone's work.
Note: the assumption in the early days of Novell was that workstations didn't have hard disks of their own and booted off the server too, making a LAN tens of thousands of £/$ cheaper than giving everyone their own HDD.
Running COMPSURF before you installed took hours. Server HDDs were big -- hundreds of megabytes! Scanning all that took ages.
https://devblogs.microsoft.com/cppblog/c-edit-and-continue-i...
https://docs.microsoft.com/en-us/visualstudio/debugger/edit-...
https://docs.microsoft.com/en-us/visualstudio/debugger/hot-r...
Here is a demo of its latest state from the VS 2022 launch event, https://youtu.be/8SP1w7i8r-Y?list=PLReL099Y5nRc9f9Jpo1R7FsdH...
If you are on Windows and need something in a console, a nice colorful asm repl is available WinRepl [1] which is similar to " yrp604/rappel (Linux) and Tyilo/asm_repl".
mov ax, 13h
int 10hI am yet to find a modern graphics programming environment that is so comfortable and easy to use as this.
I was taught as a kid to program simple graphical demos using peek and poke in basic. Then in assembler. In either case, stupid me got colored pixels on the screen after a few minutes of work. Kids these days, how do they start? Please, don't tell me "matplotlib" or I will cry myself to sleep.
https://about-prog.com/en/articles/python/hello_world_sdl2
Its basically, grab a window, show it, grab a render/draw buffer update it and make it visible.
I don't really find the base C version much more complex, although it does have a bit of boilerplate around init/window creation/grab surface/display surface/etc. I'm not sure I would consider that particularly complex. Sure SDL can get complex when you start trying to use GL/etc but if all you want is a buffer to write bytes that become pixels its pretty straightforward IMHO.
I would say its roughly the same level of complexity (if not less) than HTML canvas+JS.
I wrote a DOS TSR program (remember those?) which would pop up a window when you pressed a key sequence and present you with an ASM86 REPL.
You could selectively 'save' pieces of code, and then when you exited the window, it would paste the saved code as inline assembly code (a hex byte array surrounded by some Turbo Pascal syntax) into your keyboard buffer - the assumption being that you are running the Turbo Pascal IDE, of course.
The TSR itself was written in x86 assembly, which added a level of complexity. I would have given and arm and a leg to be able to do it in a high-level language like Ruby.
In the later days of DOS, programs grew so big that they wanted all of that 640 kB to themselves. Optional-extra type TSRs went out of fashion and DOS (first DR DOS 5, then MS played copy-cat with MS-DOS 5) gained built-in memory managers to load necessary TSRs (e.g. CD, mouse and keyboard drivers, disk cache, etc.) into Upper Memory Blocks.
UMBs were a 386 thing: you used a 386 memory manager to map any unused bits of the I/O space in the PC's memory map (i.e. from 641 kB up to 1 MB) as RAM. Anything that wasn't being used for ROM or memory-mapped I/O, you could put RAM there and then load TSRs into these little chunks of RAM -- 1 or 2 dozen kB of RAM each.
https://en.wikipedia.org/wiki/Upper_memory_area
Yes, we were that desperate for base memory. It didn't matter if you had 2 or 4 or 16MB of RAM, DOS could only run programs in the 1st 1 MB of it, and only freely use the first ⅓ of that first meg. All the rest could only be used for data, disk caches, and other non-executable stuff.
A side-effect of having a 386 memory manager, for real DOS power users, was that fancy 3rd party ones like Quarterdeck QEMM could also offer multitasking. Quarterdeck sold a tool called DESQview that let you run multiple DOS programs side-by-side and switch between them -- radical stuff in the 1980s.
But once you had that, you didn't need TSRs any more.
It is a step up from having front-panel switches (https://en.wikipedia.org/wiki/Front_panel)
On early mini- and microcomputers, those sometimes had to be used to enter the boot loader (https://en.wikipedia.org/wiki/Booting#Minicomputers, https://en.wikipedia.org/wiki/Booting#Booting_the_first_micr...). That was a step down from mainframes, which could automatically read in a program to run at boot.
It wouldn’t surprise me much if there were people alive today who still have some muscle memory to rapidly enter such a boot sequence for an Altair.
People forget that in 1984 information wasn't a click away.
The problem with owning a Hong Kong-made 286 clone in 1984, and using pirated software, is that it was extremely hard to learn things. I was limited by the books at my local "Waldenbooks" computer section, which was about 20 books. Computer shopper and Byte magazine were kinda helpful, but I learned very, very slowly. It wasn't until I entered college that I started learning rapidly, but the focus wasn't on PCs (it was still MTS mainframes). It took until my first job writing 16-bit drivers that I finally started learning the nuts and bolts of MSDOS.
http://web.archive.org/web/20051211022146/http://www.btinter...
This reminds me of a fun project I once did, writing an x86 assembler in Lotus 123, using lookup tables. On the odd occasion when it worked, it was immensely fulfilling.
IRuby is the Jupyter kernel for Rubylang:
iruby/display: https://github.com/SciRuby/iruby/blob/master/lib/iruby/displ...
iruby/formatter: https://github.com/SciRuby/iruby/blob/master/lib/iruby/forma...
More links to how Jupyter kernels and implicit display() and DAP: Debug Adapter Protocol work: "Evcxr: A Rust REPL and Jupyter Kernel" https://news.ycombinator.com/item?id=25923123
"ENH: Mixed Python/C debugging (GDB,)" https://github.com/jupyterlab/debugger/issues/284
... "Ask HN: How did you learn x86-64 assembly?" https://news.ycombinator.com/item?id=23930335
It really tough to write x86 assembly.
You could probably visualize the operand stack and opcode sequence, but it wouldn't be quite as "flashy" as x86's state transitions look when visualized here.
But you probably can run it on an M1 anyways, since Apple's Rosetta will do the dynamic binary translation for you under the hood. YMMV.
The app is entirely written in Ruby. So, it might run on Apple M1, but only if you're running an x86 Ruby interpreter through Rosetta.
In general, the way you handle translation of machine code tends to resolve around compiling small dynamic traces (basically, the code from the current instruction pointer to the next branch instruction), with a lot of optimizations on top of that to make very common code patterns much faster than having to jump back to your translation engine every couple of instructions. The interactive generation this article implies is most likely to be effected with use of the x86 trap flag (which causes a trap interrupt after every single instruction is executed), which is infrequent enough that it's likely to be fully interpreted instead of using any sort of dynamic trace caching. In the case of x86 being generated by a JIT of some sort, well, you're already looking at code only when it's being jumped to, so whether the code comes from the program, some dynamic library being loaded later, or being generated on the fly doesn't affect its execution.
This basic sort of support is needed for any application that targeting x86 that uses any form of dynamic code generation, which is probably a whole lot more than most people think (even some forms of dynamic linking utilize small amounts of generated code, due to being more efficient than calling a method though a pointer to a pointer to the method).
i'd venture a guess that the rosetta jit stuff probably does some kind of prelinking.
kinda makes me wish i had an m1 mac to play with...
No, pages and the executable bit are something that the processor knows about.
Rosetta implements x86 execution bit semantics.
It does this by invalidating translated pages when the system call to set the execution bit is set.
Which bit do you not understand?
How do you think for example the JVM works today on Rosetta?
So your model of how Rosetta works is off - the translation would need to support remapping the original code page read-only regardless of whether the x86 code did so, and letting a subsequent write invalidate the JIT cache of that page, instead of relying solely on the emulated process to implement W^X.
Most JITs do execution an icache flush, and Rosetta does catch it to invalidate their code.
For example https://github.com/openjdk/jdk/blob/master/src/hotspot/cpu/x...
Otherwise, how do you think it works?
It does if you wrote instructions from one address and execute them from another, which is why they use a flush.
> Rosetta emulates this correctly
Maybe you know more than I do, it my understanding is it does not emulate it correctly if you do not flush or change permissions.
How do you think it detects a change to executable memory without a permissions change or a flush?
char *buffer = mmap(NULL, 0x1000, PROT_READ | PROT_WRITE | PROT_EXEC, MAP_ANON | MAP_PRIVATE, -1, 0);
*buffer = 0xc3;
((void (*)())buffer)();
*buffer = 0xc3;
((void (*)())buffer)();
The region is RWX, and code is put into it and then executed without a cache flush. This requires careful setup by the runtime, and here's how Rosetta does it, line by line:1. buffer is created and marked as RW-, since the next thing you do with a RWX buffer is obviously going to be to write code into it.
2. buffer is written to directly, without any traps.
3. The indirect function call is compiled to go through an indirect branch trampoline. It notices that this is a call into a RWX region and creates a native JIT entry for it. buffer is marked as R-X (although it is not actually executed from, the JIT entry is.)
4. The write to buffer traps because the memory is read-only. The Rosetta exception server catches this and maps the memory back RW- and allows the write through.
5. Repeat of step 3. (Amusingly, a fresh JIT entry is allocated even though the code is the same…)
As you can see, this allows for pretty acceptable performance for most JITs that are effectively W^X even if they don't signal their intent specifically to the processor/kernel. The first write to the RWX region "signals" (heh) an intent to do further writes to it, then the indirect branch instrumentation lets the runtime know when it's time to do a translation.
Think about code that is modified without jumping into it, such as stubs that are modified or certain kinds of yield points.
One way how this could be implemented was the way mentioned above: By making sure all x86-executable pages are marked r/o (in the real page tables, not from "the x86 API"). Whenever any code writes into it, the resulting page fault can flush out the existing translation and transparently return back to the x86 program, which can proceed to write into the region without taking a write fault (the kernel will actually mark them as writable in the page tables now).
When the x86 program then jumps into the modified code, no translation exists anymore, and the resulting page fault from trying to execute can trigger the translation of the newly modified pages. The (real, not-pretend) writable bit is removed from the x86 code pages again.
To the x86 code, the pages still look like they are writable, but in the actual page tables they are not. So the x86 code does not (need to) change the permission of the pages.
I don't know if that's exactly how it is implemented, but it is a way.
How you are disagreeing with me, then? The actual page table entries that the ARM CPU looks at will never mark a page containing x86 code as executable. x86 execution bit semantics are implemented, but on a different layer. From the ARM CPU's POV, the x86 code is always just data.
The implementation of AMD64 is in software. It knows about page executable bits. The 'x86' code knows about them.
Again, how do you think things like V8 and the JVM work on Rosetta otherwise?
Where did I claim anything else? The thing I claimed the x86 code does not know about is the pages that contain the translated ARM code, which are distinct from the pages that contain the x86 code. The former pages are marked executable in the actual page tables, the latter pages have a software executable bit in the kernel, but are not marked as such in the actual page tables.
> Again, how do you think things like V8 and the JVM work on Rosetta otherwise?
Did I write something confusing that gave the wrong impression? My last answer says: "x86 execution bit semantics are implemented, but on a different layer".
maybe arm pages with an arm wrapper that calls the jit for big literals filled with x86 code are, or arm pages loaded with stubs that jump into the jit to compile x86 code sitting in data pages are... but if the arm processor cannot execute x86 pages directly, then it wouldn't make a lot of sense for them to be marked executable, would it?
> how would it know to translate the instructions that are being generated on the fly interactively?
Just answered your own question.