HNHacker News
TopNewBestAskShowJobs

cpldcpu

660 karma · joined January 23, 2022

github.com/cpldcpu
submissionscomments
cpldcpu··on 555 Timer Circuits
Well, you can also build microprocessors out of them:

https://hackaday.io/project/182915-555enabled-microprocessor

cpldcpu··on Un Ministral, Des Ministraux
The tokens are immediately transformed into embeddings (very large vectors), so the 17 bit values are not used for any computation.
cpldcpu··on Efficient high-resolution image synthesis with linear diffusion transformer
>We introduce a new Autoencoder (AE) that aggressively increases the scaling factor to 32. Compared with AE-F8, our AE-F32 outputs 16× fewer latent tokens,

Basically they compress/decompress the images more, which means they need less computation during generation. But on the flip side this should mean less variability.

Isn't this more of a design trade-off than an optimization?

cpldcpu··on Addition is all you need for energy-efficient language models
Bill Dally from nvidia introduced a log representation that basically allows to replace a multiplication with an add, without loss of accuracy (in contract to proposal above)

https://youtu.be/gofI47kfD28?t=2248

cpldcpu··on Addition Is All You Need for Energy-Efficient Language Models
I have to disagree. Nvidia spent a lot of effort on researching improved numerical representations. You can see a summary in this talk:

https://www.youtube.com/watch?v=gofI47kfD28

A lot of their work was published but went by unnoticed. But in fact the majority of their performance increase in new architecture is resulting from this work.

Reading between the lines, it seems that they came to the conclusion that a 4 bit representation with a group exponent ("FP4") is the most efficient representation of weights for inference. Reducing the number of bits in weights has the biggest impact on LLMs inference, since they are mostly memory bound. At these low bit numbers, the impact of using multiplication or other approaches is not really significiant anymore.

(multiplying a 4 bit wight with a larger activation is effectively 4 additions, barely more than what the paper proposes)

cpldcpu··on Addition is all you need for energy-efficient language models
It puzzles me that there does not seem to be a proper derivation and discussion of the error term in the paper. It's all treated indirectly way inference results.
cpldcpu··on Robert Dennard, DRAM Pioneer, has died
Well, it's basically the technical implementation of Moore's law, since Moore's law is just an empirical observation. (And maybe also a self-fulfilling prophecy)
cpldcpu··on Ever: Exact Volumetric Ellipsoid Rendering for Real-Time View Synthesis
what are you referencing?
cpldcpu··on Unexpected Consequences of McDonald's Touchscreen Kiosks
So we will also see an increate in obesity?
cpldcpu··on Fine-Tuning LLMs to 1.58bit
the performance is still a bit degraded though.
cpldcpu··on Chain of Thought empowers transformers to solve inherently serial problems
Can any of these tools do anything that the Github copilot cannot do? (Apart from using other models?). I tried Continue.dev and cursor.ai, but it was not immediately obvious to me. Maybe I am missing something workflow specific?
cpldcpu··on RISC-V CPU arrives on a tablet starting at $149
There probably are not too many at this point.

What would be awesome is to have an open and standardized NN accelerator to go with RISC-V, but that is a dream.

cpldcpu··on $50 2GB Raspberry Pi 5 comes with a lower price and a tweaked, cheaper CPU
What does ESP32 have to do with RPI?

The equivalent of to the ESP32 would be the Rasperry Pi Pico 1 / 2 / W. They start at $4, which is a fair price.

cpldcpu··on Flux better than Stable Diffusion
>whoever runs this site has been engaging in

You are suggesting that BFL is using "organic marketing" to push their product??

It may be worth mentioning that the BFL team actually consists of the people who invented Latent Diffusion (At CompViS, a university lab), then developed Stable Diffusion at Stability.ai and now they are pushing the state-of-the-art with Flux.1.

The attention is well deserved and Flux.1 is definitely the top model right now.

edit: had a look at r/AiArt. This seems to be a place where people post their "Bing Image Creator" output. Maybe that's not where the enthusiasts are. Try r/StableDiffusion

cpldcpu··on Hazard3: 3-stage RV32IMACZb* processor with debug
They mentioned somewhere that they could not designate the actual area increase since it is part of the synthesized digital area. The area increase to include the RISC-V cores was negligible, apparently.
cpldcpu··on Hazard3: 3-stage RV32IMACZb* processor with debug
It's also an interesting juxtaposition, because it directly allows to benchmark the architectures in the same system environment.

In the RP2350 it is possible to either use the RISC-V cores, the CM33 cores or even use one of each.

cpldcpu··on Zen 5's 2-ahead branch predictor: how a 30 year old idea allows for new tricks
My understanding is that they do not predict the target of the next branch but of the one after the next (2-ahead). This is probably much harder than next-branch prediction but does allows to initiate code fetch much earlier to feed even deeper pipelines.
cpldcpu··on Meta to release largest Llama 3 model on July 23 [405B]
At least they put all the relevant information into the title, so that it is not necessary to actually read the article.
cpldcpu··on Sonnet 3.5: Please write an asteroids game
Claude-3.5-Sonnet is SCHNITZEL mit BRATKARTOFFELN.

Waiting for 3.5-Opus or OpenAUs response...

cpldcpu··on Claude 3.5 Sonnet Reproduces BIG-Bench Canary String
Searching for the string on google yields hundreds of hits. Likely that it appears in recent webscrape-data.
cpldcpu··on ChatGPT is better at generating code for problems written before 2021
That is why the ML/AI community is usually shortcutting this process by publishing preprints on Arxiv.

A publication based on an LLM that was state-of-the art only until March of 2023 cannot be justified by long review times.

Edit: To be fair, it seems their preprint was first submitted in August 2023 and the IEEE article that is based on the paper was a bit slow...

cpldcpu··on ChatGPT is better at generating code for problems written before 2021
>Thus, in this study, we take the state-of-the-art ChatGPT (the default version of GPT-3.5), the recent popular product, as the representative of LLMs for evaluation.

Are they really publishing a paper based on GPT3.5 in July 2024? I am not sure these results are relevant in any way today.

Edit: Just for reference. The best model for coding today (according to most benchmarks) is Claude-3.5-Sonnet which is freely accessible. Also GPT-4o is freely accessible and is still vastly better than GPT-3.5.

The lm sys arena coding leaderboard (https://chat.lmsys.org/?leaderboard) lists sonnet-3.5 and gpt-4o jointly on #1 and GPT-3.5-Turbo on #35. You can freely download and run LLMs locally on your machine that are significantly better than GPT-3.5, for example Mistral Codestral.

There is really no reason to accept any results on GPT3.5 for relevant today. This is as if you were complaining that a computer from the 00ies is not running <recent operating system> well.

cpldcpu··on How to avoid flying on a Boeing 787 this summer
Extrapolating from the 737max to the 787 is ridiculous. Still prefer A380 and A350 though.
cpldcpu··on What happened to the artificial-intelligence revolution?
There is still impressive progress in large language models themselves every week. (Just to mention some of the past month: Claude-3.5-Sonnet, Gemma2, Nemetron and many more)

What is, in general, strange is that noone has really figured out how to do the plumbing. We have the LLM and they can perform almost any text related task, summarize search results or provide complex code snippets based on limited specification.

But somehow, the integration into work flows remains cumbersome.

- Search somehow seems to be burdened by the inability of the search providers to process entire webpages. Google, despite their search advantage, seems to only be able to process the search snippets instead of summarizing the entire content of websites. Most likely a copyright issue...

- Github Copilot is still basically autocomplete or a chat interface where I manually have to copy&paste results. Prompting for changes across multiple files is not really solved. (I know there is cursor, but my experience was quite mixed).

- All the hailed agents seem to create a lot of fluff but little actual code beyond what I would get with zero shot prompting. (Just tried a new tool today, which consumed $2.00 in API credits on the first task and left me with a broken codebase).

- Nothing that properly addresses slide generation yet?

Anthropics new workflow with Artifacts and Projects seems very promising and is a great leap forward. But it cannot natively process diffs or work with multiple soruce file and is therefore limited in total codelength.

As other people in this thread already remarked, maybe this is early stage technology that is pushed to commercialization too soon.

cpldcpu··on From the Transistor to the Web Browser, a rough outline for a 12 week course
>Building an FPGA board

This seems to be a bit odd? This is already a more tedious hardware project to debug, but when it is about learning the basics, building a much simpler circuit would provide more insight.

It's also a bit questionable why building hardware should be part of a full stack digitial systems course? It's very good knowledge for certain, but seems like a sidetrack for me.

cpldcpu··on SoftBank stock hits first record high in 24yrs – Arm and AI helped it get there
Anyways, I think ARM is on to something with their Edge AI IP, which enables inference in microcontrollers. So far, there is no Nvidia of microcontroller ML. If ARM managed to build a moat with proper software, they can go very far...
cpldcpu··on The Illustrated Transformer (2018)
whoa, that's awesome.
cpldcpu··on Nyquest NY8A051H – 1.5 cent microcontroller: weekend die-shot
Yeah, you can find them in all kinds of low-cost remote controls:

https://cdn.hackaday.io/images/9838991700773145118.file-1700...

cpldcpu··on Nyquest NY8A051H – 1.5 cent microcontroller: weekend die-shot
Use a neural network to detect numbers? (Well, this runs on a slightly more capable MCU, though)

https://cpldcpu.wordpress.com/2024/05/02/machine-learning-mn...

cpldcpu··on Commercial perovskite solar modules at SNEC 2024 trade show
>the price rises for the next two years were pre-announced at davos in january 02019 by eric luo, president of gcl, a top-ten chinese solar-panel company; this is a sort of announcement that cannot happen in a competitive market https://www.reuters.com/article/us-davos-meeting-solar-gcl-i...

From the article: "Solar panel prices tumbled around 30 percent last year after China, the world's largest producer, cut subsidies to shrink its bloated solar industry, pushing smaller manufacturers to the brink of collapse."

This is exactly the opposite of price fixing. They subsidized the market (consumption) and created a bloated industry. Removing the subsidies lead to overcapacity and hence a drop in prices below the rate of a healthy market.

This is exactly what is happening again this year. The difference is that we are seeing more than 50% price drop.

If you look at the fraunhofer PDF, you will notice that chinese companies own 95% of the market now. The way they were allowed to scale was by targeted chinese subsidies on PV projects that were not accessible to outside tenders.

← PreviousPage 4 of 5Next →