Introduction to ARM Assembly Basics
azeria-labs.com
azeria-labs.com
If so a link would be fantastic!
https://static.docs.arm.com/ddi0403/e/DDI0403E_B_armv7m_arm....
http://infocenter.arm.com/help/topic/com.arm.doc.dui0553a/BA...
There is also the old cheat sheet:
http://users.ece.utexas.edu/~valvano/Volume1/QuickReferenceC...
I think you've inadvertently just helped make my case for me.
I don't have a link to share unfortunately, the only way i could get access was by going to ST directly.
The chapter after that even includes instruction encoding.
Maybe you were looking for something else?
For documentation of the devices in a particular SoC you'll need the reference manual from the SoC manufacturer -- they of course vary in how easy it is to find those docs.
And you can output assembly from C programs
gcc -S <source file>
You can disassemble a binary to see how it actually looks. The resulting binary is much larger than your assembly code. objdump -d <binary file>
Many disassemblers will show you friendlier output than objdump. I use ht editor (packaged in Debian based distros as ht), an open source clone of Hiew. In ht, press F6 -> select image, and you will have an easy to follow disassembled version of a binary that you can edit, if you happen to know opcodes.What do they do? I haven't seen anything like this before.
I see some explanation in https://azeria-labs.com/arm-conditional-execution-and-branch... but the reason why branching must work this way isn't explained.
edit: I see "Branch with Exchange" switches the processor from ARM to Thumb in http://www.embedded.com/electronics-blogs/beginner-s-corner/.... But I don't see why switching processor modes must happen on branch instructions.
There are branching instructions that can't change instruction sets for two reasons: 1. A direct branch can have more range if you can assume it's going to a 4-byte aligned address. 2. For compatibility with code written for ARM7DI and other really old pre-thumb processors. Since the original direct branch instructions ignored the bottom two bits of the target address, new direct branch instructions were added rather than change the behavior of the existing ones.
This is distinct from x86 32 vs 64 bit and ARM 32 vs 64 bit: in both those cases there's really a different processor mode with extra registers and so forth, and switchover is correspondingly more involved.
I started getting fascinated about computer architecture a while back, but then I saw how dead embedded programming was in my area.
If you or anyone else is intrigued, we are hiring! https://nvidia.wd5.myworkdayjobs.com/en-US/NVIDIAExternalCar...
I do have to drop little snippets of asm in to code for bare metal stuff, but generally looking at the output of the compiler to see if it and I are on the same page is my main use.
However, as we may well be about to switch our IoT device to ARM from AVR, this might be a useful primer...
It's something id like to try (however in c++) but I'm not sure how to do it in a smart/not-too-ugly way
For saturating arithmetic and other stuff we used compiler intrinsics, which freed us from handling register allocation, stack management etc by hand. On that processor there weren't special instructions for saturating arithmetic but a flag was used instead, the compiler also kept track of that one too.
We did read the assembly result and tweaked C code until assembly looked like what was expected, though.
Talk to old video game veterans [waves]. We wrote tons of assembly because there wasn't much choice. But these were largely 8-bit processors, the compilers weren't any good and the code space was constrained. But I was talking with a guy recently who said he'd written hundreds of thousands of lines of 68K assembly, and I have no idea why you would do that because 68K C compilers, Pascal compilers, anything compilers were pretty good even back in the benighted 80s. Well, better than assembly.
Of course, once you flip into C you're still not in an environment where you have much of a runtime (I kept having to explain to a contractor why he couldn't do heap operations in an early boot phase, much less expect the results to be addressable later).
Even in a very code-space sensitive project, I started off with a tiny bit of assembly, then made everything more or less functional in C, then went back and hand-coded routines as we needed to get bytes back: http://www.dadhacker.com/blog/?p=1911
My level of fascination with an architecture can be dramatically by the quality of the available tooling. If there's nothing then that's fine, green fields are great fun. But if the tooling sucks (TI and your DSP software, I'm looking at you) or is horridly expensive, then I'm usually going to look for excuses to use something else.
I used to work next to someone programming Tilera manycore systems in assembler, because that's a sufficiently weird architecture that you need to do that in order to see any benefit. This is probably why manycore has never really taken off.
Easy to miss, particularly if you are scanning through.
I enjoyed the read!