Fixing the Ethernet Board from a Vintage Xerox Alto
righto.com
righto.com
Back in the day (mid 1980s), when I was going through an electronic engineering apprenticeship with a flight simulator company and many boards were end-to-end TTL, it was generally understood that power-hungry 74S series TTL were used for a reason (generally speed) and should not be substituted lest 'funny things' happen.
Mind you, there was one time when debugging a glitch led to one chip being replaced with the same type from a different family - something like a 74LS replacing a 74F - and the propagation delay difference was enough to fix an edge-case timing issue. Due to project constraints, the root cause wasn't tracked down and so the board bill of materials made it clear that the 'odd one out' was correct and should not be changed.
PS: Ken - I think I still have a quantity of 74S TTL 'pulls' from the day, so if all else fails sourcing a part, look me up!
[1] Read as 'I know just enough to be dangerous, but not enough to understand why what I'm doing is dangerous.'
I once traced hundreds of thousands of dollars worth of bunk boards back to some RAM chips substituted on the basis of being a "drop in replacement" by the procurement department. The timing diagrams revealed the replacements possessed different timing characteristics. Enough so that the chips would occasionally leave the bus floating when the memory was read from the host side.
Is it even possible to do this kind of debugging on modern hardware? Or is this a lost art?
EE's still have to diagnose hardware problems. Keysight will be happy to sell you diagnostic equipment for investigating problems with your 100Gbps ethernet devices. Your bank probably won't care to fund the necessary loans, though.
[1] https://www.youtube.com/playlist?list=PLkVbIsAWN2lsHdY7ldAAg...
It's so neat to watch a pro in a completely different domain. I know some of the basic concepts, but it's mesmerizing to watch him reason through voltage levels and pinpoint a failed resistor. That's magical to me.
Bad idea to eat cereal while watching his videos... That was hilarious.
I love that he leaves in all the little mistakes. Most videos are polished and presented. Tutorials are nice, but you get so much more raw data from watching this. Plus it feels like having a conversation.
Looks like Ken has some videos like this too: https://www.youtube.com/watch?v=adEr2aRwHnI
In fact, in the video I mentioned, someone who was watching his livestream pointed out that he installed a polarized capacitor backwards...
"That's not unprofessional... This is unprofessional... "
That made my day! Thanks for pointing this out.Oh well.
To some extent this is easier today (but not might be so in few years) because you can buy realtime DSO that is fast enough to capture almost any computer interface used today with usable resolution (albeit you can probably find some successful small-ish hardware startup that would cost you less than such scope)
The art in this kind of debugging is IMHO in using woefully inadequate test equipment to come up with meaningful idea of what might be broken. And this is still done today, with TV repair shops fixing OLED TVs with nothing but DVM and RLC bridge, me looking at 400Mbps+ serial LVDS on customer site with 50MHz portable DSO (they didn't have anything else) and such.
Most storage tube oscilloscopes supported single sweep capture in bi-stable mode. That is, you could clear the screen, reset the timebase, and capture the trace of a single trigger event. In bi-stable mode, the image should last hours. Some models (like the Tek 7313) even allowed the top and bottom of the screen to be set to storage or live independently.
On top of that, most scope camera shutters could be set up to automatically open during the forward seep (called "single sweep" on Tek cameras). That allows you to set up the scope, and walk away while you wait for the rare event to occur. Tektronix even sold delay lines (big coils of hardline coax) with trigger pick-offs for the express purpose of displaying before-the-trigger waveforms.
No question that DSOs make this all easier, but saying it was impossible is a bit of a stretch.
I recently debugged packet issues after 100gbps port speed change to 25gbps, to 10 gbps issues. Breakdown the issue to sub-system components - check phy, mac loopback, check link status on both side of the QSFP connections before and after speed changes, check counters (A LOT of them), enable debug packet to send to CPU with special command. Check all the VLAN, port settings, configuration commands over and over again. Enable debugging on kernel driver to track down every bytes/bits of every packet. Use tcpudmp on linux socket layer when one gets to that point.
Instead of oscilloscope, today's SOC does have a lot of counters inside that help one identifies issues.
For complex issue, one does get tremendous high when the issue was ID and resolved.
I find getting a little high really helps with long debugging sessions.
Also, I have on occasion observed the magic of an extender card causing a flaky card to function properly.
If you get stumped, you might want to depopulate some non-essential cards to open up a space to get your hands and probes on the card under investigation, while it is directly plugged into the slot. Excuse me, I mean get your hands and probes in there while the system is powered off.
(answer: "already be rich")