Tools for Learning LLVM TableGen
blog.llvm.org
blog.llvm.org
Emphasis on factors out as in even if the result is eye wateringly complex, it's much better that that complexity is in another declarative file. This factoring also allows reuse by other tools like debuggers and linkers.
Emphasis on a lot of detail because 500,000 lines of instruction, register, scheduling, ... information is encoding a lot of detail. Processed, --print-records generates 87M bytes of text just for the X86 backend.
Take the simple BPF backend. It has 5 TableGen files:
BPF.td
BPFInstrFormats.td
BPFInstrInfo.td
BPFRegisterInfo.td
BPFCallingConv.td
TableGen streams these files and processes the unloveable TableGen language over them producing a collection of records. The backend's CMakeFile.txt tells TableGen to generate .inc include files. Each of these -gen flags invokes a TableGen backend to generate a specific include. tablegen(LLVM BPFGenAsmMatcher.inc -gen-asm-matcher)
tablegen(LLVM BPFGenAsmWriter.inc -gen-asm-writer)
tablegen(LLVM BPFGenCallingConv.inc -gen-callingconv)
tablegen(LLVM BPFGenDAGISel.inc -gen-dag-isel)
tablegen(LLVM BPFGenDisassemblerTables.inc -gen-disassembler)
tablegen(LLVM BPFGenInstrInfo.inc -gen-instr-info)
tablegen(LLVM BPFGenMCCodeEmitter.inc -gen-emitter)
tablegen(LLVM BPFGenRegisterInfo.inc -gen-register-info)
tablegen(LLVM BPFGenSubtargetInfo.inc -gen-subtarget)
So here's the deal. Yeah, TableGen is ungainly but then the problem that it solves is truly massive. If you were to redo it now with what we know now, the architecture would be very similar and the language would be only slightly better. TableGen solves a really hard problem, maybe not perfectly but well.Some parts are, of course, a matter of taste (compiling in backend code rather than having them be separate programs consuming an easy-to-parse canonical form; admittedly that’s in line with the LLVM ethos). But also, honestly, the language—late binding and all—seems a bit like a more ad hoc version of CUE. If I were to look for a tool for a TableGen-shaped problem, that’d probably be the direction I would be thinking in.
The docs have been improving, https://llvm.org/docs/TableGen/
I also like these resources:
Lessons in TableGen FOSDEM 2019; Nicolai Hähnle https://fosdem.org/2019/schedule/event/llvm_tablegen/
- What has TableGen ever done for us?, http://nhaehnle.blogspot.com/2018/02/tablegen-1-what-has-tab... - Functional Programming, http://nhaehnle.blogspot.com/2018/02/tablegen-2-functional-p... - Bits, http://nhaehnle.blogspot.com/2018/02/tablegen-3-bits.html - Resolving variables, http://nhaehnle.blogspot.com/2018/03/tablegen-4-resolving-va... - DAGs, http://nhaehnle.blogspot.com/2018/03/tablegen-5-dags.html
If you want to generate code, define a better language and write a good compiler.
TableGen tries to be similar to C++, but is different enough that it is annoying.
And god help you if something goes wrong, error messages are straight up misleading.
It seemed to me that TableGen was trying to be a pure functional language more like Haskell than C++.
The core dialects are upstreamed so that people don't reinvent them, but all tools to utilize any of it aren't. Some of it isn't even open source. For example, nothing else in MLIR uses PDL right now.
Now, when you're writing an LLVM backend, these tables are immediately consumed by another tool that tries to write a very complex state machine based on these tables, and the formats of the tables that said tool expects is essentially undocumented and confusing, and the error messages tend to range from confusing to, well, straight up misleading. But technically that's not the fault of TableGen, that's the fault of that TableGen-based tool.
My point was about the 'how'. See the following example code from the manual and tell me it doesn't resemble C++ if you squint.
class PersonName<string name> { assert ... string Name = name; }
class Person<string name, int age> : PersonName<name> { assert ... int Age = age; }
def Rec20 : Person<"Donald Knuth", 60> { ... }
It has been ages since I delved into GCC internals, think GCC vs egcs days, but if I am not mistaken it was either DSL based on C macros, or some variation of GIMPLE.
Any feedback is welcome !
It specificies how the whole set of internal data structures for instructions actually map into real CPU instructions, and also guides the register allocation algorithms.