HNHacker News
TopNewBestAskShowJobs

lkcl

74 karma · joined November 29, 2008

submissionscomments
lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
yes. it is however a test ASIC. therefore it has no on-board boot ROM, and has to have programs uploaded to it over JTAG.
lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
we use nmigen (python-based OO HDL) which through yosys generates verilog as an automatic step.

180nm is still by far and above the world's most heavily-used geometry, because the price-performance (bang per buck, however you want to put it) is so extremely high.

an 8in wafer is USD 600 and that's extremely low. any power MOSFET, power transistor, diode or other high current semiconductor you absolutely don't want small "things" (detailed tiny tracks) you want MASSIVE ones.

why on earth would you waste money on tiny features, it's like using the latest 0.15mm 3D printing nozzles to 3D print a massive 300x300x300 mm cube that's going to be used for nothing more than a foot-stool. you want a 1.2mm nozzle for that!

then any processor below 300 mhz, you can get away with 180nm. need only an 8 mhz 8-bit or 4-bit washing machine or microwave processor, or something to go in a cheap digital watch? 180nm is your best bet: you'll get tens of thousands of < 1 mm^2 ASICs on a single wafer which means you're well below $0.05 per individual die.

a 28nm 8in wafer would be about... 10x that cost, you'd end up with exactly the same transistor (or 8 mhz 8-bit processor), why would you pay more money for what you don't need?

btw the real reason why there's a chip shortage: the Automotive industry, who are cheap bar-stewards, wanted even lower than $600 per 8in wafer so they went with 360nm and cruder geometry. that's equipment that's even older than the 1990s, like 40+ years in some cases.

so then the stupidity hit, and they stopped ordering. then 18 months later they phone up these old Foundries and say, "ok, we're ready to start ordering again". and the Foundries say, "oh, we switched off the equipment, and it cooled down and got damaged (just like that massive Electric plant in S. Australia that was de-commissioned, the concrete cracked when they switched it off, and it's completely unsafe to start up again). you were our only customer for the past 30 years, so we scrapped it all. you'll have to now compete with the consumer-grade smaller geometry Fabs like everyone else".

which is something that none of the Automotive companies have told their Governments, because then they can't go crying "boo hoo hoo, we can't make chips any more at the price that we demand, waaa, waaaa, i wannnt myyy monneeeeey"

and now of course they can't use the old masks, because those were designed for 360nm and cruder geometries, they have to redesign the entire ASIC for 180nm and that's why you can't now get onto 180nm and other MPW Programmes because the frickin Automotive Industry has jammed them all to hell.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
thx phndrenad2. funny i just searched "chisel gpu" and found two: https://github.com/jbush001/ChiselGPU https://github.com/Chlorophytus/broccoli

half-float we'd like to do by using a dynamic SIMD-aware 64-bit ALU that has auto-partitioning. we do however already have an actual FP16 implementation https://git.libre-soc.org/?p=ieee754fpu.git;a=tree;f=src/iee...

or more to the point, one that is compile-time configureable with one parameter (bit-width), so the same HDL does FP16, FP32 and FP64. i'd like to make that dynmaically-SIMD-configureable but it'll take some base work in nmigen to do without massive code-explosions.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
you can get a pretty good idea right now, the simulator is functional and the unit tests include explanations in english:

https://git.libre-soc.org/?p=openpower-isa.git;a=tree;f=src/...

i'm currently in the middle of a rabbit-hole exploration of being able to do in-place RADIX-2 FFT, DCT and DFT butterflys, the target is a general purpose function to cover each of those, in around 25 Vector instructions.

not 2,000 optimised loop-unrolled instructions specifically crafted for RADIX-8, another for RADIX-16, another for RADIX-32 ..... RADIX-4096 (as is the case in ffmpeg): 25 instructions FOR ANY 2^N FFT.

btw if you're interested in "real-world" SVP64 Vector Assembler we have the beginnings of an ffmpeg MP3 CODEC inner loop:

https://git.libre-soc.org/?p=openpower-isa.git;a=blob;f=medi...

that's under 100 instructions, more than 4x less assembler for the same job in PPC64. and 6.5 times less assembler than ffmpeg's optimised x86 apply_window_float.S

you will no doubt be aware of the huge power savings that brings due to reduced L1 cache usage.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
yes, so the "normal" way that GPUs work is: the architecture and the ISA are so staggeringly optimised they're completely incompatible and incapable of running standard (general-purpose) workloads. no MMU, vast wide SIMD engines, massive numbers of parallel memory interfaces that run really slowly but can handle (when added up) vast bandwidth far in excess of "normal" processor memory, and so on.

on top of that, because it's an entirely separate processor, to get it to do anything you actually have to have a Remote Procedure Call system, operating over Shared Memory!

oink.

so the process for running a GPU shader binary is as follows:

step 1: fire up a compiler (in userspace) step 2: compiler takes the shader IR and turns it into GPU assembler step 3: the userspace program (game, blender, whatever) triggers the linux kernel (or windows kernel) to upload that GPU binary to the GPU step 4: the kernel copies that GPU binary over Shared Memory Bus (usually PCIe) step 5: now we unwind back to userspace (with a context-switch) and want to actually run something (OpenGL call) step 6: the OpenGL call (or Vulkan) gets some function call parameters and some data step 7: the userspace library (MESA) "packs" (marshalls) those function call parameters into serialised data step 8: the userspace library triggers the linux (windows) kernel to "upload" the serialised function call parameters - again over Shared Memory Bus step 9: the kernel waits for that to happen step 10: the userspace proceeds (after a context-switch) and waits for notification that the function call has completed...

... i'm not going to bother filling in the rest of the details, you get the general idea that this is completely insane and goes a long way towards explaining why GPU Cards are so expensive and why it takes YEARS to reverse-engineer GPU drivers.

in the Libre-SOC architecture - which is termed a "Hybrid" one, the following happens:

step 1: the compiler is fired up (in userspace, just like above) step 2: compiler takes the shader IR and turns it into *NATIVE* (Power ISA with Cray-style Vectors and some custom opcodes) assembler step 3: userspace program JIT EXECUTES THAT BINARY NATIVELY RIGHT THERE RIGHT THEN

done.

did you see any kernel context-switches in that simple 3-step process? that's because there aren't any needed.

now, the thing is - answering your question a bit more - that "just having vector capabilities" is nowhere near enough. the lesson has been learned from Nyuzi, Larrabee, and others: if you simply create a high-performance general-purpoes Vector ISA, you have successfully created something that absolutely sucks at GPU workloads: about TWENTY FIVE PERCENT (one quarter) of the capability of a modern GPU for the same power consumption.

therefore, you need to add SIN, COS, ATAN2, LOG2, and other opcodes, but you need to add them with "reduced accuracy" (like, only 12 bit or so) because that's all that's needed for 3D.

you need to add Texture caches, and Texture interpolation opcodes (takes 4 pixels @ 00 01 10 11 square coordinates, plus two FP XY numbers between 0.0 and 1.0, and interpolates the pixels in 2D).

you need to add YUV2RGB and other pixel-format-conversion opcodes that are in the Vulkan Specification...

and many more.

but, we first had to actually, like, y'know, have a core that can actually execute instructions at all? :) and that's what this first Test ASIC is: a first step.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
they both have the same connections on the outside (they both have the same "netlist") and you can use the exact same SPICE model (a transistor-level simulation) but usually they're entirely empty inside.

so the VLSI tool can still Place-and-Route them, you can still creaate GDS-II Files, but if you send them to the Foundry, the Foundry will look at you like you have two heads or something and won't talk to you again.

that said: some Foundries have their own Symbolic ("ghost") Cell Libraries, which they send you. you run the VLSI tools with those, then when they get the GDS-II files they SUBSTITUTE the REAL cells for the ghost Cells... and then put that into the Fab.

they do this because they're so paranoid they don't even want you to know what's inside their "Symbolic" (ghost) Cells.

Foundry Symbolic Cells are invariably available only under NDA.

sigh.

which begs the question, how the hell is any information is going to leak out from a completely empty Cell, and unfortunately the answer is: quite a lot. number of layers, what the "stack" is of those layers, distance between tracks, width of tracks, and so on, and the PDK also has to include via sizes and so on anyway.

this starts to give you some idea of the levels of insanity we had to workaround, to meet our Audit and Transparency objectives.

bottom line is until we can bust through these final layers of NDAs, customers who really want to verify the complete GDS-II Files are also going to have to sign a Foundry NDA.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
we'll be going as far as is practical and pragmatic with the actual hardware, and still actually meet user-expectations. firmware, bootloader, OS, drivers, BIOS: definitely.
lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
> Also, as your webpage states, you signed TSMC's NDA:

FALSE. again. i do not work for LIP6. i do not work for Chips4Makers. i am an independent *LIBRE* Developer. i have NEVVERRRR signed a Foundry NDA and, having a background involving security analysis and Reverse-Engineering, it would be suicidally and monumentally stupid and counter-productive for me, personally, to do so.

please try to not conflate matters (twice in succession) that you haven't checked or read properly. the best thing to do is to ask questions, such as:

"You're a Libre Project. that has significant implications that everything is entirely Libre. I notice however that you say that someone signed a Foundry NDA? what impact did this have for you? did it stop you from releasing any source code as per obligations of LIBRE Licenses?"

and then i can answer positively and in a friendly way rather than having to publicly waste both my time and that of readers in first unpicking the mistakes, embarrassing you in the process (which risks a public confrontation that annoys everybody even more), and it all goes to hell pretty quickly after that.

answering the question above that you didn't ask: as you know there are about five layers of NDAs in the Silicon Industry.

we've managed to bust through three of those, and so have managed - as a LIBRE Team - to fulfil our obligations both to our funding body, NLnet, under their Privacy and Enhanced Trust Programme, and to Libre/Open Hardware developers by releasing all HDL under LGPLv3 Licenses

     https://git.libre-soc.org
and using Libre-Licensed VLSI tools

and using Libre-Licensed Cell Libraries

now, the TEAM THAT DEVELOPED the VLSI tool - signed a TSMC NDA.

      NOBODY ON THE LIBRE-SOC TEAM SIGNED THAT NDA.
also, Chips4Makers - the developers of FlexLib - signed a TSMC NDA

      CHIPS4MAKERS != Libre-SOC
we are three separate and INDEPENDENT teams, working together, to tackle an insane situation, at different levels. i'll say it again:

      LIBRE-SOC HAS NOT SIGNED AAAANNNYYYYY FOUNDRY NDAs.
are we clear about that, now?

there happens also to be another team, Libre-Silicon, also funded by NLnet, who are developing an actual Libre VLSI process and actually developing a mini home-grown Fab.

then there is another NLnet-sponsored project, working with the Libre Silicon team, to develop another Libre-Licensed Standard Cell Library, that is targetted at Libre-Silicon's PDK (when it's available)

  https://nlnet.nl/project/LibreSiliconStandardCellLibrary/
however neither of these are ready, so we went with the pragmatic route, after exhausting all other options: the parallel track.
lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
you'll be fascinated to know that we picked a python-based (Object-Orientated) HDL - nmigen - for exactly this reason.

we've developed a dynamically SIMD-partitionable-maskable set of "base primitives" for example, so you set a "mask" and it automatically subdivides the 64-bit adder into two halves. but we didn't leave it there, we did shift, multiply, less-than, greater-than - everything.

https://git.libre-soc.org/?p=ieee754fpu.git;a=blob;f=src/iee... https://git.libre-soc.org/?p=ieee754fpu.git;a=blob;f=src/iee...

can you imagine doing that in VHDL or Verilog? tens of engineers needed, or some sort of macro-auto-generated code (treating VHDL / Verilog as a machine-code compiler target).

the reason for doing this - planning it well in advance - is because we're doing Cray-style Vectors (Draft SVP64) with polymorphic element-width over-rides. yes, really. the "base" operation is 64-bit, but you can over-ride the source and destination operation width.

the reason why we're using our own Cell Library is actually down to transparency. we want customers to be able to compile the GDS-II files themselves, fully automated, no involvement from us, no manual intervention.

ironically, as an aside: Staf's Cells are 30% smaller (by area) than the Foundry equivalents.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
yehyeh, or you bought a winmodem over 15 years ago, someone told you "hey you have to upgrade to windows 10", it downloads over your 56k Dialup winmodem, reboots... and... no drivers. yes this really happens: ThinkPenguin stock TTYACM USB modems and their biggest customers are Rural people in the USA who are too far out to get broadband!

LIP6 does actually have a fully NDA-free silicon-proven Cell Library, called nsxlib, it's been used in 360nm and 180nm, the 180nm was done by a Japanese University. i think i may have mentioned this already, it's a small town with a 2(?) micron foundry, they make it available to people anywhere in the world entirely for free, it's for training the employees of the town, because it's so old and basic it's hard to mess it up. so they want people to submit designs that the trainees can learn how to fab, before they move on to the more expensive equipment.

but, really, use Chips4Makers, he has 360nm available, EUR 1750 for 20 MPW chips in QFP, i believe.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
ultimately what we'd like to see is entirely NDA-free PDKs even for 12nm and below, and you can run the VLSI tools and generate the EXACT GDS-II yourself, then yes, de-cap the processor and do a digital comparison.

before you even get to that stage, you run the Formal Correctness Proofs and unit tests on the HDL, so that YOU have confidence that the HDL which you're about to generate the GDS-II files from is actually correct and does the damn job.

example of a Formal Correctness Proof for the fixed arithmetic Power ISA pipeline:

https://git.libre-soc.org/?p=soc.git;a=blob;f=src/soc/fu/alu...

runs with symbiyosys, so you end up running SAT Solvers like yices2 and z3.

basically we absolutely do not want to be the people you come to and say, "can we trust your ASIC?" and like Intel they lie to you and say "of course!", we want to say, "don't bloody well ask us, go run the damn tools yourself! oh, btw, if you want help with that we charge USD 5k per hour"

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
thanks :)
lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
no, it's pretty basic, and implicit: it's the (newly-created) "Scalar Fixed-Point Compliancy Subset) - i added a bit to the wikipedia page last month about them https://en.wikipedia.org/wiki/Power_ISA#Compliancy

it's 64-bit, LE/BE, and it's implementing a "Finite State Machine" (similar technique to picorv32, if you know that design). this because we wanted to keep it REALLY basic, and also very clear as a Reference Design, none of the "optimised pipelined decoders and issuers" that you normally find, which make it really, really difficult to see what the hell is going on.

bear in mind this includes SVP64: https://git.libre-soc.org/?p=soc.git;a=blob;f=src/soc/simple...

if you go back several revisions, the non-Vectorised version is like... 400 lines?

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
we used an entirely Libre-licensed VLSI "compiler", which takes HDL as input and spits out fully-completed GDS-II Files.

the problem with this particular irate individual is that he's assumed that because TSMC's DRC rules are only accessible under NDA that automatically absof*** everything was also "fake open source".

idiot.

sigh.

clearly didn't read the article.

whilst both Staf Verhaegen and LIP6.fr signed the TSMC Foundry NDA, we in the Libre-SOC team did not. we therefore worked entirely in the Libre world, honoured our committment to full transparency, whilst Staf and Jean-Paul and the rest of the team from LIP6 worked extremely hard "in parallel".

the ASIC can therefore be compiled with three different Cell Libraries:

* LIP6.fr's 180nm "nsxlib" - this is a silicon-proven 180nm Cell Library * Staf's FreePDK45 "symbolic" cell library using FlexLib (as the name says, it uses the Academic FreePDK45 DRC) * the NDA'd TSMC 180nm "real" variant of Staf's FlexLib

i was therefore able to "prepare" work for Jean-Paul, via the parallel track, commit it to the PUBLIC REPOSITORY (the one that's open, that our resident idiot didn't bother to check existed or even ask where it is), which saved Jean-Paul time whilst he focussed on fixing issues in coriolis2.

it was a LOT of work.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
HDL source code: https://git.libre-soc.org/?p=soc.git;a=summary

Coriolis2 source code: http://coriolis.lip6.fr/

Chips4Makers FlexLib Cell Library based on FreePDK45: https://gitlab.com/Chips4Makers/c4m-pdk-freepdk45/-/releases

Automated Layout scripts for generation of GDS-II Files: https://git.libre-soc.org/?p=soclayout.git;a=summary

please do try to get your facts right and not mislead people by making false claims, eh?

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
> The POWER/PowerPC ISA is still widely used in safety-critical avionics

and in the Mars Rover, which is a radiation-hardened 133mhz 32-bit Power ISA system.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
yes, Chips4Makers http://chips4makers.io will help anyone who wants to do a 360nm ASIC, the costs are ridiculously cheap. like... EUR 1750 for 20 MPW samples, something mad, who would have ever thought it.

Staf will also "protect" you from the Foundry NDAs. you develop with a "symbolic" version of the Cell Library, he runs the "Real" one and sends it to IMEC on your behalf. here's Staf's "symbolic" Cell Library, it's based on FreePDK45 https://gitlab.com/Chips4Makers/c4m-pdk-freepdk45/-/releases

Coriolis2 - http://coriolis.lip6.fr/ - is entirely Libre-Licensed. it's fully automated, you don't have to do any "hand-editing", it has unit tests (so you have demos you can look at and also check you installed everything right). we have some automated setup scripts for it if you're interested: https://git.libre-soc.org/?p=dev-env-setup.git;a=blob;f=cori...

LIP6 have a Silicon-proven ENTIRELY Libre Cell Library called nsxlib, if you really want to go that route. it's Silicon-proven in 360nm and 180nm.

Also, LIP6 have a relationship with a small town in Japan, they have 2 micron fab which is used for "training" of employees of the town. submission for that is entirely free. i know this exists but have not used it, and don't know more details, but i can probably put you in touch with Sorbonne University if you're serious.

and if you really really want to do "at home" stuff, Libre-Silicon is developing a 2in wafer fab, using Ultra-Violet DLPs and high-accuracy stepper motors, that you'll be able to buy and operate from your garage or lab. think "3D printing", i think they're aiming for 2000 nm or something (20 micron)? really big, but proves the concept.

lkcl··on Libre-SoC 180nm Power ISA v3.0 ASIC Submitted to IMEC MPW
interestingly, Libre-SOC and NLnet's funding pre-dates the google-sponsored Skywater 130nm process. also, because it's funded by NLnet we're not dependent on google, don't have to pass "conditions", and in particular were not forced to use OpenLane and were not limited to 48 pins controlled by a "Management Engine".

Staf actually developed actual IOpad Cells (from scratch), actual Standard Cells and a 4k SRAM block: we did not use the NDA'd TSMC Cell Libraries, here.

if we had used Skywater 130nm we would have been forced to ditch LIP6.fr (i cannot express enough how hard Jean-Paul Chaput has worked on coriolis2 for the past 18 months), we would not have been able to test the IOpads that Staf developed... yeah.

bottom line is we used a complete independent VLSI toolchain - fully automated - that has nothing to do with the USA or DARPA Military funding - and was developed with European expertise.

lkcl··on Servers as they should be – shipping early 2022
"Attests the software version is the version that is valid and shipped by the Oxide Computer Company"

this makes "Oxide Computer Company" the primary target and point of vulnerability in multiple ways.

1) rogue employees (state-sponsored, corporate espionage) could replace the software. customers could do nothing about it, and might not even be told.

2) sale of the company by the VCs or a Corporate take-over gives no guarantee that what is safe now will be safe in future, no matter what the VCs or the company says right now.

3) whatever expertise "Oxide Computer Company" thinks they have, they're the single-point-of-failure. the larger the number of customers, the less likely that a given vulnerability will be immediately fixed and distributed out.

this is just some of the possibilities. sorry to say that there's so many things wrong with this idea it's really hard to hold back and not say anything.

now, if the full source code right to the bedrock is available, and the CUSTOMER is given FULL CONTROL, THEN we do not have a problem.

by "full control", that includes:

* all DRM keys including TPM signing private keys * all peripheral initialisation source code (including DDR4 firmware, PCIe firmware and USB3 firmware) * BMC (Boot Management Console) source code * BIOS source code * Operating system source code * full source code for all tools and toolchains for the above to avoid vendor lock-in and the possibility of the toolchain itself introducing rogue code.

this is one hell of a list and it's almost impossible to fulfil with today's "NDA'd proprietary firmware 3rd party licensing" mindset. the only company in this secure server space to my knowledge that's achieved this is Raptor Engineering with the TALOS-II, when running with the Kestrel BMC replacement, on the Lattice ECP5 FPGA.

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
this is what we're doing in Libre-SOC, and using standard python software engineering practices (writing unit tests at every step of the way) to do it. some of those unit tests are actually Formal Correctness Proofs (asserts, but for hardware).

i was stunned to find that in the hardware world, test-driven development is not standard practice: people simply haven't been trained that way. they write thousands to tens of thousands of lines of unverified code and finally do an integration test.

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
unfortunately the data still has to be in a full 32BPP framebuffer (in order for 2D or 3D Window Managers to write to it).

all the graphics software will write at that full 32BPP rate: it is unfortunately not reasonable to expect all Graphics Software to perform data compression.

yes there will be a few opportunities for compressed streaming, and funnily enough we were just thinking "what if the GPU were to do the same trick as the Amiga used to do, 40 years ago, by following the scan-lines?"

in the case of video playback you might reasonably expect that an area of the screen is "reserved" (not written to) but that when the Video-Output HDL gets to that pixel instead reads directly from a completely different device: one that is formatted already in HDR/SDR or YUV.

effectively this is a modern "sprite" engine.

... but, again, we have significantly diverged from the original "Just Make It Real Simple Why Don't You Just", yeh? :)

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
yes. HDCP has "infected" HDMI, eDP, USB-C and so on.

this can entirely be avoided with:

* TFP410a (RGBTTL to DVI/HDMI)

* SN75LVDS83b (RGBTTL to LVDS)

* SSD2828 (RGBTTL to MIPI)

* various other converter ICs

4k btw is MENTAL bandwidth. multiply 3840 by 2160, then by 4, then by 60: this gives the number of bytes per second required of the internal memory bus.

turns out to be 2 gigabytes per second, doesn't it?

now check the datasheets on what affordable FPGA boards can do, what memory ICs they have, and whether they can cope with that level of bandwidth.

you'll find that there aren't any.

i do find it ironic that "incremental steps" are recommended, "to get something running", but lack of knowledge of the difficulty surrounding DRM and in ramping up to high speed leads to "disbelief". ah well :)

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
Jeff's evaluation of GPLGPU is fascinating: https://jbush001.github.io/2016/07/24/gplgpu-walkthrough.htm...

you are absolutely correct in that everything has moved on from "Fixed Function" of SGI, and how GPLGPU works (worked) - btw it's NOT GPL-licensed: Frank sadly made his own license, "GPL words but with non-commercial tacked onto the end" which ... er... isn't GPL... sigh - but everything commercially has now moved on to Shader Engines.

that basically means Vulkan.

however you may be fascinated to know, from Jeff's evaluation, that there are still startling similarities in basic functionality in not-GPL GPLGPU and in modern designs targetted at Shader Engines.

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
allo jeff nice to see you're around :) thank you so much for the time you spend guiding me through nyuzi. also for explaining the value of the metric "pixels / clock" as a measure for iteratively being able to focus on the highest bang-per-buck areas to make incremental improvements, progressing from full-software to high-performance 3D.

have you seen Tom Forsyth's fascinating and funny talk about how Larrabee turned into AVX512 after 15 years?

https://player.vimeo.com/video/450406346 https://news.ycombinator.com/item?id=15993848

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
this is unfortunately not true (that the RISC-V ISA was designed not to require currently-valid patents). people may believe that to be the case, but it's not. from 3rd hand i've heard that IBM has absolutely tons of patents that RISC-V infringes. whether IBM decide to take action on that is another matter. they're a bit of a heavyweight, so there would have to be substantial harm to their business for the "800 lb gorilla" effect to kick in.
lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
if it were done, say, as a Libre/Open processor, say, with the backing of NLnet (a Charitable Foundation), where the "Bad PR ju-ju" for trying it on was simply not worth the effort

if it were done. say, as a Libre/Open processor, say, with the backing of NLnet (a Charitable Foundation), where NLnet has access to over 450 Law Professors more than willing to protect "Libre/Open" projects from patent trolls by running crowd-funded patent-busting efforts

if it were done as a Libre/Open Hybrid Processor, based on extending an ISA such as ooo, I dunno, maybe OpenPOWER, which has the backing of IBM with a patent portfolio spanning several decades, who would be very upset if tiny companies like NVidia or AMD tried it on against a Charitably-funded project.

that would be a very interesting situation, wouldn't it? i wonder if there's a project around that's trying this as a strategy? hmmm, hey, you know what? there is! it's called http://libre-soc.org

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
this is easy to chuck together in a few days, literally, from pre-existing components found on the internet.

* litex (choose any one of the available cores)

* richard herveille's excellent rgb_ttl / VGA HDL https://github.com/RoaLogic/vga_lcd

* some sort of "sprite" graphics would do https://hackaday.com/2014/08/15/sprite-graphics-accelerator-...

the real question is: would anyone bother to give you the money to make such a project, and the question before that is: can you tell a sufficiently compelling story to get customers - real customers with money - to write you a Letter of Intent that you can show to investors?

if the answer to either of those questions is "no" then, with many apologies for pointing this out, it's a waste of your time unless you happen to have some other reason for doing the work - basically one with zero expectation up-front of turning it into a successful commercial product.

now, here's the thing: even if you were successful in that effort, it's so trivial (Richard Herveille's RGB/TTL HDL sits as a peripheral on the Wishbone Bus) that it's like... why are you doing this again?

the real effort is the 3D part - Vulkan compliance, Texture Opcodes, Vulkan Image format conversion opcodes (YUV2RGB, 8888 to 1555 etc. etc.), SIN/COS/ATAN2, Dot Product, Cross Product, Vector Normalisation, Z-Buffers and so on.

lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
unfortunately, if you make modifications and you want them to be "upstreamed" (using libre/open project terminology as an alonogy) you cannot do that without participating in the RISC-V Foundation. you can implement APPROVED (Authorized) parts of the RISC-V specification. you cannot arbitrarily go changing it and still call it "RISC-V", that's a Trademark violation.
lkcl··on Nyuzi – An Experimental Open-Source FPGA GPGPU Processor
only if the patent holder does not create an "improvement" on the old one. then the older (referenced) patent is extended. Bosch have done this specifically so that they can hold on to the original CAN Bus patent.
lkcl··on The Libre-SOC Hybrid 3D CPU [pdf]
not at all. we've two companies, one focussed on full transparency for Banking and other high-security environments, willing to sign Letters of Intent. we're looking for three more as it will make investment pretty straightforward.
Page 1 of 3Next →