Pyramid3D Real-time Graphics Processor (1997) [pdf]
vgamuseum.info
vgamuseum.info
They were simple enough that the system integrators or board partners would actually be the ones writing the drivers for them as the company was just selling them the chips with the datasheets and manuals with instructions on how to interface with them via PCI and how to program them, that's it.
NVidia were famous for being the first to in-house the driver development themselves instead of their board partners which gave them the edge on driver quality and performance.
Shortly after, Nvidia started producing the whole board and those value-added pure grahiccs companies like Herculues went out of business.
I'm not aware of any GPU that explicitly supported such primitives, but you actually get these flat trapezoid for free with certain designs of a triangle rasteriser. A triangle is rasterised with 3 edge equations, and it's useful to explicitly define a start and finish y position. A flat trapezoid is just two edge equations with the same start and finish y positions (and I assume their line primitive is the same thing, a but with a single edge equation)
While I'm not aware of any GPU with explicit support for these flat trapezoids, the Nintendo 64's RDP takes edge equations directly (usually generated by the RSP microcode), and if you feed it malformed edge equations, it can render certain non-triangle polygons, though most of them aren't useful.
But one of the malformed polygons are the same flat trapezoid primitive. And while playing around and attempting to design a better RSP microcode, I discovered that these flat trapezoids appear to be useful for a fast polygon clipping algorithm. My goal was to replace the N64's slow depth buffering with portal-based clipping on the RSP, but there is no way to check if it's faster without an actual RSP implementation.
For your N64 experiment, you say there's no way to check if it's faster without an actual RSP implementation - was there something stopping you from doing so? I've seen other people write custom RSP code in present day.
* HP later patented this idea (expired now).
While exposing these lower level internal primitives typically did not make sense for general purpose graphics accelerators, some in-house embedded implementations did actually go further and only supported flat trapezoids, relying on the CPU to preprocess more complex shapes. For instance, the "SOLO" ASIC used in second-generation WebTV boxes took this approach [1] (among other interesting cost-cutting measures such as operating natively in YUV color space).
[1]: http://wiki.webtv.zone/misc/SOLO1/SOLO1_ASIC_Spec.pdf (warning: 200 MB scan)
"The list price for the (Dynamic Pictures) Oxygen 102 was $1495 in 1996, later reduced to $399."
[1] https://www.vgamuseum.info/index.php/cpu/item/1067-dynamic-p...
The tech was real, they just had bad luck with manufacturing partners. Eventually they pivoted to 2D graphics accelerators and that became the successful business. ATI later sold it to Qualcomm, I think, where it became part of their mobile graphics stack.
This article is a bit disjointed in places but helps fill in my understanding: https://hardforum.com/threads/bitboys.1973024/