HNHacker News
TopNewBestAskShowJobs

azonenberg

244 karma · joined March 27, 2014

submissionscomments
azonenberg··on Memory Mapping an FPGA from an STM32
I have a H735 on a retired board slated for decap so we'll find out once I open it up.

Do you know if it's fabbed in house, TSMC, or Samsung? I've seen ST silicon from all 3 foundries but the only thing I've seen stated publicly is 40nm. When I get it opened up it should be easy to tell, TSMC and Samsung processes have distinctive features on them that I recognize by sight.

azonenberg··on Memory Mapping an FPGA from an STM32
I have encountered issues with QSPI (mostly caused by the annoying prefetch queue) which is why I am switching to the FMC for FPGA interfacing (i.e. not using OCTOSPI). That was the whole point of this experiment, validating FMC as a replacement for my legacy OCTOSPI based MCU-APB bridge. I have a previous board using QSPI reliably in indirect mode (i.e. not memory mapped) but found it was full of pain when memory mapped specifically in writes. So that firmware memory maps it for reads but switches to indirect mode for writes. And has cache disabled.

So far I have it working quite reliably (my test firmware does a loopback test with 100K reads/writes of a 32-bit register at the start that I had written with intent of using it for link training of the PLLs to optimize read/write capture timing but never ended up using as such) and my iperf test can push tens of thousands of packets per second without issue.

azonenberg··on Memory Mapping an FPGA from an STM32
Multech (multech-pcb.com) is my preferred manufacturer these days for high end stuff. I've done six layer HDI any-layer via stackups, ten layers with filled via-in-pad, RO4350B, TU872SLK, flex, 75 micron trace/space, etc. And that's nowhere near the limit of their capabilities, I just haven't needed higher end yet.

I have some 25/100G stuff in the pipe for probably some time next year that I plan to make with them too.

Their website undersells, I get the impression most of the actual sales contacts are word of mouth. I talk to my sales rep by skype mostly (the alternatives are expensive international phone calls or wechat).

The really cool thing is that you get a 10+ page QA report with every order including measured copper/dielectric/soldermask thicknesses, hole sizes, ionic contamination measurements, and a ton of other metrics. And they send the TDR strips and polished cross section with every order as their way of saying "look, we actually did the QA, double check our measurements if you don't trust us". (I actually have repeated some of the measurements to spot-check and got results within a few percent of their QA department, no surprises there).

And they don't make silent gerber changes or anything. They do a full CAM review and send you working gerbers and a list of suggested DFM tweaks for you to sign off before beginning manufacture. If something doesn't look right you have a chance to say "wait there's a problem".

For example, one time they wanted to make a really large width adjustment for impedance on some RF traces that I had carefully modeled in an EM solver. But they didn't make a bad board without telling me, they flagged it on the CAM review and we went back and forth before realizing the mistake was on their end (they had calculated impedance assuming solder mask over the traces, while they were actually exposed copper). They re-ran the numbers which then closely matched my simulations, I signed off on the modified design, and the board was manufactured without issue.

azonenberg··on Memory Mapping an FPGA from an STM32
S1000-2 is quite cheap and lossy (Df 0.016), slightly better than Isola 370HR (0.021) but nowhere near the stuff I usually use. At my usual Chinese board house it's one of the lowest cost substrates available for prototypes since it's always in stock and there's no need to special order.

For higher end digital work I typically reach for Taiwan Union TU872SLK (Df 0.009) which also has a better range of prepregs and glass styles available to help minimize fiber weave effect. Still quite a bit lossier than e.g. RO4350B but far less expensive and if you have decent equalizers on your SERDES the difference is typically not significant unless you're making some kind of humongous backplane. I get wide open eyes with just a tiny bit of post-cursor emphasis on the TX FFE at 10.3125 Gbps on TU872SLK for my typical shortish high speed tracks (FPGA to SFP+ cage).

azonenberg··on Memory Mapping an FPGA from an STM32
H735 is one of the single core SKUs. Just a 550 MHz M7.

Would not surprise me if the M4 was there and fused off (i.e. same die as multicore H7 offerings), but it's not active.

azonenberg··on Memory Mapping an FPGA from an STM32
The projects in question include things like a 48 port gigabit Ethernet switch with packet datapath in the FPGA, and dual 10/25G SFP28 uplinks. You're not doing that on a MCU. Also higher end oscilloscope work (e.g. 10 Gsps 12-bit JESD204B)

But a STM32 is more than sufficient for the management interface on both.

azonenberg··on Memory Mapping an FPGA from an STM32
The intent is for the high performance datapath to live entirely in FPGA (and the project you're probably thinking of is switching, not routing).

The MCU is for control plane only. Several hundred Mbps between the control and data plane is more than enough for a SSH management CLI and poking registers on the FPGA to move a port to a different VLAN in response to a CLI command or add an ACL rule or something.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
I had some ideas for quotas but never got to the point of fully building that part of the system.
azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
The CPU used in the thesis is a barrel processor. There is a forcible context switch every clock cycle, equal priority round robin for now although more sophisticated prioritization schemes could be used.

The advantage of a barrel scheduler is that you never can have more than one pair of instructions in the pipeline from the same thread at a given time, so data hazards are impossible and all of the checking/forwarding logic can be entirely absent from the CPU.

You lose single-thread performance with this vs more conventional hyperthreading, but it's much simpler to implement.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
I was targeting industrial control systems, medical implants, and other critical applications. So a lot of the constraints of desktop/mobile/server computing just don't apply.

The assumption was that you'd end up with something close to a microkernel architecture, but without any all-powerful software up top. A purely peer to peer architecture running on bare metal.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
This architecture also borrowed a lot from the exokernel philosophy, which would actually have provided significant speedups by cutting unnecessary bloat and abstraction.

My conjecture was that this would make up for most of the overhead, but I didn't have time during the thesis to optimize it sufficiently to do a fair comparison with existing architectures.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
This is a SoC, addresses of on chip devices are fixed when the silicon is made.

The source address is added to a packet by the on-chip router as it's received. You can send anything you want into a port, but you can't make it seem like it came from a different port without modifying the router.

This model allows you to integrate IP cores from semi-untrusted third parties, or run untrusted code on a CPU, without allowing impersonation.

Let me put it another way: I ran a SAT solver on the compiled gate-level netlist of the router and proved that it can't ever emit a packet with a source address other than the hard-wired address of the source port. Any packet coming from a peripheral into port X will have port X's address on it when it leaves the router, end of story.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
My guiding design principle was to take a trendline from monolithic kernels to microkernels, then keep on going until you had no kernel left.
azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
My long term plan was actually to do a full formal verification of the CPU against the ISA spec and prove no state leakage between thread contexts, but I didn't have time to do that before I graduated

I deliberately went with a very simple CPU (2-way in order, barrel scheduler with no pipeline forwarding, no speculation or branch prediction) to minimize opportunities for things to go wrong and keep the design simple enough that full end to end formal in the future would be tractable. Spectre/Meltdown are a perfect example of attack classes that are entirely eliminated by such a simple design.

I was targeting safety critical systems where you're willing to give up some CPU performance for extreme levels of assurance that the system won't fail on you.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
Keep the hardware design as simple as possible, formally verify everything critical, then run the rest as userspace software and hope you got all the critical bugs.

The point wasn't to move everything into silicon, it was to move just enough into silicon that you no longer needed any code in ring-0.

As an example, the memory controller's access control list and allocator was a FIFO of free pages and an array storing an owner for each page. Super simple, very few gates, hard to get wrong, and easy to verify.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
That was the point of the research: throw away how everything has been done and explore what a clean slate redesign would look like if we had the benefit of 30-40 years of hindsight.

The easiest way to implement transparent paging would be to have a paging block that exposed the memory manager API (allocate, free, read, write) and would proxy requests to either RAM or disk as appropriate. But since I was targeting embedded devices, it was assumed that you would explicitly allocate buffers in on-die RAM, off-die RAM, or flash as required. The architecture was very much NUMA in nature.

The prototype made for my thesis only allowed the "terminate" request to come from the process requesting its own termination. It would be entirely plausible to add access controls that allowed a designated supervisor process to terminate it as well.

Remote memory can be inspected, but only with explicit consent. I had a full gdbserver that would access memory by proxying through the L1 cache of the thread context being debugged. (Which, of course, had the ability to say no if it didn't want to be debugged).

The goal was to have a system that had very small number of trusted hardware blocks which did only one thing, then build an OS on top of it microkernel style - except that all you have is unprivileged servers, there's no longer any software on top.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
Which is why the interconnect fabric adds the source address based on who made the request.

Even hardware IP blocks don't have the ability to send packets from arbitrary addresses. You can send to anywhere you want, but you can't lie about who made the request. Which allows access control to be enforced at the receiving end.

That "something you are" as well as "something you know" factor prevents any capability from being spoofed, since the identity of the device initiating the request is part of the capability.

azonenberg··on Antikernel: A Decentralized Secure Operating System Architecture (2016) [pdf]
NoC addresses are implicitly part of the handle. So for example, a pointer isn't just physical address 0x41414141. The full capability is "physical address 0x41414141 being accessed by hardware thread 3 of CPU 0x8000" because when you dereference that pointer, it creates a DMA packet from the CPU's L1 cache with source address 0x8003 addressed to the NoC address of the RAM controller.

If malicious code running in thread context 4 tries to reach the same physical address, the packet will come from 0x8004. Since the RAM controller knows that the page starting at 0x41414000 is owned by thread 0x8003, the request is denied.

The source address is added by the CPU silicon and isn't under software control, so it can't be overwritten or spoofed. If the attacker were somehow able to change the "current thread ID" registers, all register file and cache accesses would be directed to that of the new thread, and all memory of their evil self would be gone. Thus, all they did was trigger a context switch.

azonenberg··on Hardware Reverse Engineering Course
We tried to record a few lectures but had technical difficulties with audio quality, they were totally garbage so they didn't get posted.

To my knowledge this is the only course of its kind that has ever been taught, I had to write all of the slides from scratch since I couldn't find any lecture notes to base mine on.

Sadly I've graduated and moved on to other things, but am still actively working in the field (see my recent conference talk https://recon.cx/2015/slides/recon2015-18-andrew-zonenberg-F...)

azonenberg··on Hardware Reverse Engineering Course
You got U4, J3, P2, U2 right.

U5 was a ROT13(SCTN) but your answer is close enough to get credit.

U13 was a ROT13(SGQV HFO pbagebyyre), you can tell since it connects to J2. I don't recall what U1 is but you can ask marshallh (retroactive.be) since he's the guy who designed the board.

Source: I'm the instructor and wrote the quiz

azonenberg··on Why Apple's iPhone encryption won't stop NSA
> The article that the author cites the possibility of the "Secure Enclave code being able to read the UID key"; as comex mentioned yesterday [1], this isn't true.

We don't know that, we just know Apple says it's the case and nobody's broken it yet. Without a full reverse engineering of the Secure Enclave firmware, plus the IC, there's no way to know if there's a hidden backdoor, bug, or debug mode allowing the data to be read.

azonenberg··on Why Apple's iPhone encryption won't stop NSA
Exactly. I wrote it to set the record straight: Apple's crypto is intended to guard against a limited class of attacks, and "the KGB is after me" is not one of them.
azonenberg··on Why Apple's iPhone encryption won't stop NSA
If the key is kept across a battery replacement or repair procedure, then it's going to be hard-wired/fused into the chip. SRAM needs to be powered constantly to retain data.

Credit cards/smartcards include self-destructs that will erase the nonvolatile memory (flash) in certain cases if power is applied while a tamper signal is asserted. They cannot erase data while in the "off" state. One of the problems with fuse-based memory is that it's easier to dump off the silicon than, say, Flash.

Although I haven't decapped an A7 yet (as soon as I get my hands on one, rest assured I will) adding flash to an IC fab process is very expensive and adds somewhere around a dozen new masks, so OTP fuse memory (which doesn't need any new masks) is typically used instead of flash for on-die ID codes etc.

azonenberg··on Why Apple's iPhone encryption won't stop NSA
Disk encryption will only stop a targeted attack (someone physically gets their hands on your phone) anyway. All this does is raise the cost by a bit.
azonenberg··on Why Apple's iPhone encryption won't stop NSA
You're assuming I'm not affiliated with siliconpr0n.

I'm actually one of the main contributors to the site and took a lot of the photos on it, just not that particular one. John (my friend who actually admins the server) is fully aware of the situation and just raised the resource limits to counter the DoS. If either of us uses an image somewhere that we expect to stay online for a while, we make a point of leaving it in place when reorganizing directory structures etc.

azonenberg··on Depackaging the Nintendo 3DS CPU
See my lecture notes: http://security.cs.rpi.edu/courses/hwre-spring2014/Lecture9_...
azonenberg··on Depackaging the Nintendo 3DS CPU
> It's doubtful you'd have a functional chip

I have a couple of fully functional examples on my desk that say otherwise. If you actually want to keep the device usable to the point that you can still solder it to a board, then you normally preserve most of the package, bond wires, and leadframe which requires more care during decap. For this particular specimen we didn't bother because we just wanted the ROM.

Here's an example of a fully functional decapped device soldered back to a board: http://i.imgur.com/UebB3FO.jpg

My lecture notes at http://security.cs.rpi.edu/courses/hwre-spring2014/Lecture3_... go into more detail on various methods, chemical and otherwise, for decapping with and without preserving the leadframe.

← PreviousPage 2 of 2