Bit banging a 3.5" floppy drive
floppy.cafe
floppy.cafe
Consider what would happen if a data sector contained 12 * 0x00, 3 * 0xA1, 0xFE, ... then this implementation could mis-sync. On real hardware you'd have to be even more unlucky to have the CRC match -- which also isn't checked in this implementation. Using this as-is on real floppies could result in reading corrupt data. It would be possible to construct a floppy that would read correctly on real hardware and sometimes mis-read data with the current SW implementation.
The problem is that the 0xA1 bytes on the disk are special. The 0xA1 bytes are MFM encoded with a missing clock pulse -- making them "0xA1 syncs" that don't match an 0xA1 data byte.
This is what I think dragontamer is alluding to in another thread -- you cannot properly decode the header unless you also recognize the missing clock pulse. So it is important to do the clock recovery in order to notice this.
The special encoding is also present for 5.25 and early hard drives using MFM encodings.
While I no longer have an MFM controller, I believe I still have an RLL controller in my pile o' stuff.
This explains the 2M utility that allowed storing about 1.8mb on a floppy disk. It was fun playing with it.
This bit is particularly relevant: https://www.pagetable.com/docs/visualize_1541/sector.png See the solid bit at the end of the sector, just before the next header? You could squeeze a few more bytes in there, but if the drive motor is just slightly too fast, it'll overwrite the next sector. That's why there's a gap, tolerance for timing variation.
Most floppy drive technologies wrote blindly, guessing where they were on the disk based on timing estimates since the controller last saw a sector header. This is also why disks needed to be "formatted". Not just in the sense of writing the file system data structures, but writing out all the sector headers. This had to be done all at once with the same drive, due to those small timing variations.
https://porterolsen.wordpress.com/2016/06/15/accessing-mac-f... https://c65gs.blogspot.com/2023/10/reading-amiga-disks-in-me....
Microsoft itself was shipping software on ordinary PC floppies formatted for 1.68MB https://en.wikipedia.org/wiki/Distribution_Media_Format
Denise is the video chip. Paula handles interrupts, audio, the floppy data signal (control logic signals are in CIA) as well as the uart.
Management just didn't let their engineers/architects get things done.
Specifically, AmigaOS's floppies write one sector after another, with a single gap in the entire track.
Whereas the IBM PC format has a gap in each sector. This is because RAM was more expensive back then, thus holding an entire track in memory would have an associated cost.
A major factor - I'd argue even more significant than cost of memory - was that the IBM PC used an off-the-shelf floppy disk controller chip (NEC uPDC765A), which was hardwired to support the industry-standard IBM floppy track format (which had evolved from IBM's 3740 mainframe data entry system, introduced in 1973), and didn't support the Amiga's custom track format. Whereas, the Amiga could do this because it didn't actually have a floppy disk controller-the functionality of controlling floppy disks was in part contained in their custom ASICs, and in part implemented in software by the CPU. Unlike the Amiga, the original IBM PC eschewed proprietary ASICs in favour of off-the-shelf chips, in order to minimise time-to-market.
Ironically, it is possible to read arbitrary formats with the debug read track operation. Yet, it needs to find one triple sync word somewhere in the track; a controller limitation.
Unfortunately, the Amiga standard track format didn't account for that, and uses the same sync word (the 4489 one) but double.
It could have been designed to use a different sync word, and include a triple 4489 at the track start, but they didn't think about it at the time.
Some tricks bit-banging the controller allow for writing arbitrary tracks.
It is also possible to read arbitrary tracks, if there's two floppy drives and a standard ibm pc formatted disk is present in the other one, by switching the drive after the controller has started reading.
The track format was designed that way in the early 1970s. RAM was likely one reason for it, but another was that the 3740 used sector gaps as record boundaries; it was commonly configured so each 128 byte disk sector held a single 80-column punch card worth of data. 3740 format floppies support deleting sectors (by using a different sync word in the sector header) so you could delete database records. PC floppy controllers supported deleted sectors too, even though almost no software used them (some copy protection schemes did, but duplicating deleted sectors isn’t hard once you know they exist.) Part of the motivation for 128 byte sectors was likely the fact that it was the smallest power of 2 that could fit a whole punch card.
Also, a whole 3740 track was 3338 bytes (26 sectors of 128 bytes), which was a lot of RAM in 1973; in 1981, a whole PC floppy track was 4096 bytes. 4KB was a lot less expensive in 1981 than in 1973, so by then it would have been less of a motivating factor than when the track format was initially defined.
Then there were things like Spiradisc (https://en.wikipedia.org/wiki/Spiradisc), which created incompatibilities by design.
I guess the alignment on that one particular drive must have been ridiculously far off baseline.
Just the absolute worst drives ever released to the public, but it was that or tapes and while most will agree that the 1541 was a slow FDD, waiting 30+ mins to load a tape (don't forget to flip it) was much worse.
Then copy protection companies decided that the very best way to CP software was to create unreadable/unwritable sectors and let the drive slam the head against the arrester trying to read it, because they didn't have to replace the drive. That was a you problem.
For real! Do they mean 2MB or 2Mb ?
The motors in the disk drives could be controlled directly, and you could pack the tracks tighter by stepping the motor just a little bit less than you were supposed to. And in theory if you did it right, other disk drives could read it.
'In theory' is carrying a lot there. I tended to find 1.5something to 1.6something worked and anything higher rarely ever did.
"There are 80 tracks on your average 3.5" floppy drive. You can select a given track by pulsing the STEP pin and combining it with the direction select pin."
The first few generations of consumer hard drives adopted some of these same techniques.
It was extra sectors not tracks, although is looks like some people added another track or three at the edge of the disk.
Writing the 1 depends on the precision of the magnets and I suspect they just weren't able to utilize the media better, with the tech of the time.
Now we use lasers. ;)
https://m.youtube.com/watch?v=qSehRwClXNk
Tldr, a bit more than 1.7MB.
Analysing the circuits, I saw the controller chip was wired such that use of the DMA chip was software configurable, so I thought beaut, I'll write and test the first iteration without DMA, then add and test that after.
Couldn't get it working. Scratched my head for a while until, while discussing it with one of the hardware engineers, was told that while the controller chip had been wired software configurable, the floppy drive itself was hard-wired to use DMA. If only I had the spec for that I would have figured it out!
So added usage of the DMA chip… success!
I assume even if you were using something that looked like a standard floppy drive, it was using a higher level interface, or was including something like the controller ISA board in the "interface".
It turned out that "vendors not providing datasheet" has forever been the problem for hardware developers.
And on that page the "make sure the compiler didn't inject 10,000 lines of boundary checks" bit told me everything I needed to know about what language the project was written in :lol: - here's the link to the driver: https://github.com/SharpCoder/floppy-driver-rs
(Side note: I'm glad to see the Teensy continuing to get love; I adopted it back when it was at v1 and v2 as it was just such a complete no-brainer of a better choice than the Arduino stack everyone was using back then. I think now there's even an Arduino-on-Teensy software stack, but I've moved to just using STM32 directly even for just fun home hacks and have greatly enjoyed coding for that target in rust.)
I don't think it'd be suitable for this task, though. The Pi may be faster, but accessing GPIOs from userspace is slow.
Note that you don't have to run Linux, you can just stick to the Arduino eco-system, which is far more limited than what a full Linux environment would offer you, the Pi Pico is cheap for what it does and gives you an option with lots of memory in the footprint of the smaller Arduino's.
The write gate basically turns on the electromagnet in the head, which will do exactly what you'd expect that to. Early floppy drives' documentation actually came with schematics which show this more clearly.
Early hard drives based on ST-506 also have a very similar interface.
That said, Teensy 4.0 is 600 MHz ARM cpu, so there are 1000 cycles even between the shortest transitions.. some overhead is fine, the project is not exactly cpu-starved.
I also wonder if author has considered using a peripherals for precise signal capture? Something like timer in capture mode feeding into DMA buffer would allow hardware signal capture with very high precision and without any dependencies on exact instructions emitted.
In my work projects, I can use the dangerous code like this - because we have compilers and libraries frozen, unit and integration tests, a complex testing process. We can do all the right efforts to ensure the things work, even if solution is intrinsically unreliable.
In my personal projects I write some stuff and start using it, the testing is minimal, and toolchain versions is "whatever platformio decided to pull up today". I'd hate for my project to break just because I rebuilt it to add the new feature and meanwhile my compiler got upgraded. So I'd definitely abuse SPI port or something to get things reliable.
can you explain this a bit more? When you say "transition" are you talking about an individual transistor moving from on to off or vice versa?
As described in the page, there are multiple signals changing when operating floppy ("track 0", "write gate", "data", etc..). Of them the fastest one is "data", so that's what I am going to focus on.
The 2nd page of writeup [0] says:
A short transition (S) will nominally have 2us between bits, and represents 0b10
A medium transition (M) will nominally have 3us between bits, and represents 0b100
A long transition (L) will nominally have 4us between bits, and represents 0b1000
So we are looking at 3, 4 or 5 microseconds between bits. To get this in CPU cycles, you multiply this by clock frequency - google can help you with units, searching for "2 microseconds * 600 megahertz" [1] shows the answer, 1200, right away. I've rounded this down to 1000, as there are two transitions per pulse and it is all very approximate anyway.And then you use your embedded knowledge to assign meaning to the number: the CPU is ARM, so 1 instruction/cycle is a good approximation (it could be more due to dual-issue or less due to jumps). So you have like 1000 instructions. Each function call in language like C or C++ might be a 5-20 instructions overhead, and you probably want to read that pin at least 10 times to detect both transitions. The tightest loop is also going to be a dozen instructions or less (read gpio, mask, compare, maybe jump out, increase, compare timeout, loop)
So.. you can do it in C/C++ easily if your main loop involves no function calls (and you have no interrupts). If you use functions to read, your timing is going to be tight and those functions are better be super-optimized, you will be asking your compiler for a lot. Higher level languages like lua/micropython are out of the question (at least for that loop). And as I learned from reading this, rust is also out of the question, although I wonder if there are some unsafe primitives which do not do any checking.
(and yes, there are transistors changing in the background all the time throughout the process, but I really don't care much about them, they are on too low of the abstraction level)
[0] https://floppy.cafe/mfm.html
[1] https://www.google.com/search?q=2+microseconds+*+600+megaher...
Nit: That's definitely the wrong approach though IMO.
So you want to accomplish two things:
1. Clock recovery -- Figuring out the timing of a signal
2. Decoding -- Figuring out what that signal means
These are two separate steps and should be done separately, be it in code or hardware. Though advanced protocols combine both into a single step, the older protocols (UART / Floppy / etc. etc.) had these two concepts separated into two different steps.
You won't have 2us between bits: but instead 2.01us or 1.99us between bits, etc. etc. Clock-recovery mechanisms means that even in the face of worst-case timing differences, your code remains resilient.
Decoding is the step you've done here, but it should be done after clock-recovery.
-------------
Traditional clock recovery methods are phase-locked-loops (in hardware), or various XOR-loops (in software) to try and figure out the timing of the clock from the 0-1 and 1-0 transitions.
----
The traditional UART (ex: 9600 baud or 115200 baud) is ~16-ticks per bit. (IE: a 9600 baud UART needs to look at the signal 153600 times per second. A 115200 baud UART needs to look at the signal 1843200 times per second). The 16-times per bit helps you "center your aim" for the transition. You then typically aim at the center-3 timeslots (ex: count number 7, 8, and 9) for when to send and/or read the signal.
--------
That being said, your analysis for "how many instructions you have per timeslice to read the data" is correct. I just feel like adding that the clock-recovery portion needs to be definitely addressed.
After all that is what the FM in MFM implies. There is an obvious parallel with the simplest approach to demodulating FSK (or for that matter DTMF) in digital domain, which works by counting/timing zero transitions of the signal.
The UART receivers are similar in that there is no clock recovery, with the assumption that the clock is stable enough that any kind of frequency error or drift will be insignificant for the relatively short (usually 10bit) frame. The oversampling is there to align the sample point with middle of the symbol and the majority voting from multiple samples serves to average out effects of spiky noise that may be superimposed on the signal.
You need to discover the edges of the clock, and make sure you read _AWAY_ from those edges. The bits are not well defined on the clock edges. Even with a 100% accurate clock, if you're reading on the edges you'll be very unreliable.
UARTs aim to read on the "center" of bits. (If there are 16x reads per bit, then the "center" is on reads 7, 8, and 9). You'll want to stay away from reading on timeslot#1 or timeslot#16.
Obviously there is a bit of history and the whole system was originally implemented electro-mechanically, which is the reason for things like two or more stopbits (it creates time for the mechanism to settle to the reset state) or even the concept of NUL character.
Well yeah.
It's still clock recovery though and the algorithm to find that edge is the same as clock recovery algorithms in general (lastVal XOR thisVal) to find that edge. All decisions by the UART receiver are based on what happens on that edge.
Lets take a proper encoding scheme, like 8b/10b encoding. The difference is that while UART has 1x opportunity per frame, 8b/10b has multiple opportunities per frame (worst-case 111110 or 000001 as longest string of 0s or 1s) to recover the clock.
Yeah yeah, DC Balance and other such niceties. But from the decoding perspective / hardware+software that decodes the data perspective, 8b/10b is just finding the edges and trying to read from the middle again. Just faster, tighter-tolerances and other benefits compared to the braindead easy/simple UART (but with less... good... clock recovery built in).
I don't see how so basic a call would create conditional branches that would have to be predicted. Calling the function is an unconditional branch-and-link, after which it should just be doing a load and returning. (It's LLVM with the equivalent of -O2, it's not going to be doing anything weird.) Unless the return address isn't cached in these processors?
Chuck Peddle of 6502 fame did a PC in the eighties called Victor 9000 https://en.wikipedia.org/wiki/Sirius_Systems_Technology#Vict... where he implemented special software driven (6502 controller just like in Commodore drives) ZCLV/GCR floppy drives. 1.2 MB using ordinary 5.25 DSDD disks (500/360 KB capacity in PC). 100% capacity gain from switching to 80 tracks, another 50% on top of that due to ZCLV/GCR. Standard PC 1440KB 3.5 inch HD floppy using same strategy would start at 2160KB.
There is this great "Discovering Hard Disk Physical Geometry through Microbenchmarking" blog post by Henry Wong visualizing whats going on on the platters https://blog.stuffedcow.net/2019/09/hard-disk-geometry-micro...
Whats crazy to me is learning early HDDs didnt use ZBR either, manufacturers were leaving 50% capacity gain on the floor. 1989 Seagate ST157a is comparable to HD floppy, denser tracks at 650 vs 135 tpi, but almost two times lower Bit density :o 10000 vs 17434 bpi. ST157a could lose one platter, or gain half capacity 40->60MB by going with a little bit more complicated ZBR controller (switching bitrate per track).
Just a few bucks and sports a 12 Mbit/s USB interface (and wifi for pico-w).
Most of the effort was in entering the text of the labels manually or restarting and switching between capture formats when a Mac or IBM disk cropped up. FluxEngine uses the expected format to decide whether the track most recently read contains errors; and if so, re-read the track 3-5 times, quintupling overall read-times.
All that's needed now is a floppy handling robot, a macro for recording the physical disk and label descriptions, and maybe a container format for all the capture products; thumbnails, cover art, metadata, flux images; that emulators and disk utilities can read, and that the community can re-distribute easily.
Still need a lot of work on my pile of old Amiga disks.
No wonder floppies were sooo daaarn sloooooow.
Relative to the datasette (C64) I was used to.
- https://decromancer.ca/greaseweazle/
- https://www.cbmstuff.com/index.php?route=product/product&pro...
I think it'd be interesting to connect a high sensitivity / resolution sampling probe directly to the analog output of the drive heads. You could do software-defined signal processing to potentially recover damaged data. These USB-based tools are getting the signal after being amplified in the analog domain and processed by the drive's electronics.
That wire though..
It seems they could have gotten twice the speed by having two read and two write pins, one one additional pin.
But that aside, like dusted I too wonder just _how_ hard it would have been for a company like, say Apple, to demand the extra circuit. Might not have been worth it for just 2X speedup.
Now I want to build a 10X floppy RAID ...
Ask for it to be integrated and now you go from being able to buy drive from almost anyone to only one vendor.
It is odd it never improved and stayed as simple as it was until the end of floppies. The power of history I guess.
After that it’s all backwards compatibility. No one wants a new interface.
I remember the first time I saw a floppy pin out, I was dumbfounded. I expected it to be far more complex like IDE. I didn’t realize the term “floppy controller” would be so literal.
Of course in retrospect it makes sense. IDE = Integrated Drive Electronics. I assume the hard disks that came before worked more like the floppy and their controllers did more work too.
I guess no one cared enough to enhance the poor floppy drive until we got to things like LS-120 that I assume flat out couldn’t use the old interface and never took off anyway.
You nailed it. The first hard disk interfaces for IBM PCs were basically identical to the floppy disk controller just faster/bigger.
https://en.wikipedia.org/wiki/ST506/ST412
The ST506 interface between the controller and drive was derived from the Shugart Associates SA1000 interface,[5] which was in turn based upon the floppy disk drive interface,[6] thereby making disk controller design relatively easy
But as they started wanting to make disk reads faster and faster, the drive part had to get smarter and smarter
The ST412 disk drive, among other improvements, added buffered seek capability to the interface. In this mode, the controller can send STEP pulses to the drive as fast as it can receive them, without having to wait for the mechanism to settle. An onboard microprocessor in the drive then moves the mechanism to the desired track as fast as possible
At that point the drive is starting to getting smart enough that you might as well integrate the whole controller on there.
The original versions of IDE/ATA were basically just an extension of ISA, and the first "IDE cards" were just a bridge to the ISA slot they went in.
You also had "hard cards" where the hard disk itself was integrated onto the ISA card... https://en.wikipedia.org/wiki/Hardcard