HNHacker News
TopNewBestAskShowJobs

cpldcpu

660 karma · joined January 23, 2022

github.com/cpldcpu
submissionscomments
cpldcpu··on Commercial perovskite solar modules at SNEC 2024 trade show
>now they cost 8¢ a watt, down by half since last year's prices, prices which had been held steady for several years only by a price-fixing cartel.

Sorry, this is most likely not true and an unbased claim.

>this is precisely what you would expect in a niche technology as it becomes mainstream: the costs drop dramatically as you climb down the learning curve. one mechanism in that process is that the less efficient producers are driven out of the market because they can only sell at a loss

Yes, this is a common relationship in manufacturing known as wright's law. https://en.wikipedia.org/wiki/Wright%27s_Law

The cost/prices trends for PV are well documented, e.g. see slide 50 here: https://www.ise.fraunhofer.de/content/dam/ise/de/documents/p...

There was no recent transition into "mainstream". The scaling has been taking place for decades already and nothign was "held steady" for 12 years.

If you look at price development during the last months (https://www.pvxchange.com/Price-Index) you will notice that the prices drastically dropped by more than 50% during the last year. This is a clear sign of oversupply, regardless of discussions regarding tariffs etc.

To explain this kind of price drop based on the history scaling laws (24.4% per doubling) would require the total production capacity (and installation) to have more than quadrupled over the last year.

cpldcpu··on Commercial perovskite solar modules at SNEC 2024 trade show
Most likely the fab sells it at a loss, given the current situation in the industry:

https://www.reuters.com/business/energy/china-solar-industry...

cpldcpu··on Scientists develop fatigue-free ferroelectric material
10e6 cycles is nothing. A CPU at 10MHz writing at the same memory location would create that stress within 100ms.

Note sure if this is a misinformed article or some information is missing.

cpldcpu··on ASML Aims for Hyper-NA EUV, Shrinking Chip Limits
that was 20 years ago
cpldcpu··on 21.2× faster than llama.cpp? plus 40% memory usage reduction
It's just continued pretraining to "heal" the damage caused by switching the activation functions and enforcing sparsity.

Apparently they managed to recover original performance on standardized tests after continuing pretraining with the 150B tokens. There may be some more specialized knowledge lost that was not covered by their dataset.

cpldcpu··on Scalable MatMul-Free Language Modeling
>This reaches demoscene levels of crazy/impressive!

The exp/log trick to multiply with addition does indeed look very familiar. I know that a number of demos used it in the 90ies to simplify matrix multiplications for 3d graphics.

Bill Dally, the Nvidia Chief Scientist also seems to be a big fan of log8 representations: https://youtu.be/gofI47kfD28?t=1953

It seems it did not make it into the hardware though, yet...

cpldcpu··on Scalable MatMul-Free Language Modeling
Good point, I went right to the bitnet code. I will correct my original post.
cpldcpu··on Scalable MatMul-Free Language Modeling
The quantization approach is basically identical to the 1.58bit LLM paper:

https://arxiv.org/abs/2402.17764

The main addition of the new paper seems to be the implementation of optimized and fused kernels using triton, as seen here:

https://github.com/ridgerchu/matmulfreellm/blob/master/mmfre...

This is quite useful, as this should make training this type of LLMs much more efficient.

So this is a ternary weight LLM using quantization aware training (QAT). The activations are quantized to 8 bits. The matmal is still there, but it is multiplying the 8 bit activations by one bit values.

Quantization aware training with low bit weights seems to lead to reduced overfitting by an intrensic tendency to regularize. However, also the model capacity should be reduced compared to a model with the same number of weights and a higher number of bits per weights. It's quite possible that this only becomes apparent after the models have been trained with a significant number of tokens, as LLMs seem to be quite sparse.

Edit: In addition to the QAT they also changed the model architecture to use a linear transformer to reduce reliance on multiplications in the attention mechanism. Thanks to logicchains for pointing this out.

cpldcpu··on Zebra remains on the loose in Washington state, officials close trailheads
>Several people stopped to help corral the animals, including a rodeo clown and horse trainers.

ok

cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
Maybe something simpler, like a haar wavelet, would also work? Or DFT using Görtzel?
cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
On the CH32V003, a load should be two cycles if the code is executed from SRAM, there are additional wait states for load from flash. The V2A does only cache a single 32 bit instruction word, so there is basically no cache.

This publication seems to describe more details on the arduino implementation:

https://arxiv.org/abs/2105.02953

It appears that the code is even using floats in some implementations, which have to be emulated. So I'd wager that both on algorithmic level (QAT-NN) and implementation level there are some discrepancies that lead to better performance on the CH32V003.

cpldcpu··on Apple Vision Pro Successor Not Expected Until End of 2026
https://www.statista.com/statistics/677096/vr-headsets-world...

This seems to be inconsistent with these numbers

cpldcpu··on Apple Vision Pro Successor Not Expected Until End of 2026
>It was 23 billion dollar market in 2023.

How is that possible? Assuming an ASP of $333 USD (which is too high), this would be >70mio sold headsets?

cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
The latter one. The network is trained in full precision (this is required for the gradient calculation), but the weights are nudged towards the quantized values.
cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
The different is in using quantization aware training, where the quantization of the weights is already simulated during the training. This helps to restructure the network in a way where it can optimally store information in the allotted number of bits per weight.

When the NN is quantized only after training, a lot of information is lost, or you have to use less aggressive quantization that will have a lot of redundancy.

cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
Indeed using a mouse sensor for data input would be quite interesting. Mayber another option would just be a row of phototransistors.
cpldcpu··on Implementing Neural Networks on a "10-cent" RISC-V MCU
Great project! I used MNIST because it is easy to work with as a dataset. Audio classification would be quite interesting as a follow up, but I assume one would need some kind of transform to deal with the data in an easier way.
cpldcpu··on Clear – Open-Source FPGA ASIC
nah, it will easily fit. not sure about routing ressources though.

=== CPU8BIT2 ===

   Number of wires:                319
   Number of wire bits:           3029
   Number of public wires:          38
   Number of public wire bits:     209
   Number of memories:               0
   Number of memory bits:            0
   Number of processes:              0
   Number of cells:                 79
     $_DFF_PP0_                     24
     $lut                           55
cpldcpu··on Clear – Open-Source FPGA ASIC
I think it's a great initiative. However, marrying it to Caravel means that it basically a prototype and far away from a standalone product.

Would be great to see a path to a real OSHW ASIC. That would require either a substantially larger crowd funding initiative or somehow setting up an ecosystem with fixed die and package options for OSHW IC products.

I don't think the costs would be as high as everyone thinks, but the effort and time investment would be significant.

cpldcpu··on Clear – Open-Source FPGA ASIC
This CPU should fit :)

https://github.com/cpldcpu/MCPU

← PreviousPage 5 of 5