https://hackaday.io/project/182915-555enabled-microprocessor
660 karma · joined January 23, 2022
https://hackaday.io/project/182915-555enabled-microprocessor
Basically they compress/decompress the images more, which means they need less computation during generation. But on the flip side this should mean less variability.
Isn't this more of a design trade-off than an optimization?
https://www.youtube.com/watch?v=gofI47kfD28
A lot of their work was published but went by unnoticed. But in fact the majority of their performance increase in new architecture is resulting from this work.
Reading between the lines, it seems that they came to the conclusion that a 4 bit representation with a group exponent ("FP4") is the most efficient representation of weights for inference. Reducing the number of bits in weights has the biggest impact on LLMs inference, since they are mostly memory bound. At these low bit numbers, the impact of using multiplication or other approaches is not really significiant anymore.
(multiplying a 4 bit wight with a larger activation is effectively 4 additions, barely more than what the paper proposes)
What would be awesome is to have an open and standardized NN accelerator to go with RISC-V, but that is a dream.
The equivalent of to the ESP32 would be the Rasperry Pi Pico 1 / 2 / W. They start at $4, which is a fair price.
You are suggesting that BFL is using "organic marketing" to push their product??
It may be worth mentioning that the BFL team actually consists of the people who invented Latent Diffusion (At CompViS, a university lab), then developed Stable Diffusion at Stability.ai and now they are pushing the state-of-the-art with Flux.1.
The attention is well deserved and Flux.1 is definitely the top model right now.
edit: had a look at r/AiArt. This seems to be a place where people post their "Bing Image Creator" output. Maybe that's not where the enthusiasts are. Try r/StableDiffusion
In the RP2350 it is possible to either use the RISC-V cores, the CM33 cores or even use one of each.
Waiting for 3.5-Opus or OpenAUs response...
A publication based on an LLM that was state-of-the art only until March of 2023 cannot be justified by long review times.
Edit: To be fair, it seems their preprint was first submitted in August 2023 and the IEEE article that is based on the paper was a bit slow...
Are they really publishing a paper based on GPT3.5 in July 2024? I am not sure these results are relevant in any way today.
Edit: Just for reference. The best model for coding today (according to most benchmarks) is Claude-3.5-Sonnet which is freely accessible. Also GPT-4o is freely accessible and is still vastly better than GPT-3.5.
The lm sys arena coding leaderboard (https://chat.lmsys.org/?leaderboard) lists sonnet-3.5 and gpt-4o jointly on #1 and GPT-3.5-Turbo on #35. You can freely download and run LLMs locally on your machine that are significantly better than GPT-3.5, for example Mistral Codestral.
There is really no reason to accept any results on GPT3.5 for relevant today. This is as if you were complaining that a computer from the 00ies is not running <recent operating system> well.
What is, in general, strange is that noone has really figured out how to do the plumbing. We have the LLM and they can perform almost any text related task, summarize search results or provide complex code snippets based on limited specification.
But somehow, the integration into work flows remains cumbersome.
- Search somehow seems to be burdened by the inability of the search providers to process entire webpages. Google, despite their search advantage, seems to only be able to process the search snippets instead of summarizing the entire content of websites. Most likely a copyright issue...
- Github Copilot is still basically autocomplete or a chat interface where I manually have to copy&paste results. Prompting for changes across multiple files is not really solved. (I know there is cursor, but my experience was quite mixed).
- All the hailed agents seem to create a lot of fluff but little actual code beyond what I would get with zero shot prompting. (Just tried a new tool today, which consumed $2.00 in API credits on the first task and left me with a broken codebase).
- Nothing that properly addresses slide generation yet?
Anthropics new workflow with Artifacts and Projects seems very promising and is a great leap forward. But it cannot natively process diffs or work with multiple soruce file and is therefore limited in total codelength.
As other people in this thread already remarked, maybe this is early stage technology that is pushed to commercialization too soon.
This seems to be a bit odd? This is already a more tedious hardware project to debug, but when it is about learning the basics, building a much simpler circuit would provide more insight.
It's also a bit questionable why building hardware should be part of a full stack digitial systems course? It's very good knowledge for certain, but seems like a sidetrack for me.
https://cdn.hackaday.io/images/9838991700773145118.file-1700...
https://cpldcpu.wordpress.com/2024/05/02/machine-learning-mn...
From the article: "Solar panel prices tumbled around 30 percent last year after China, the world's largest producer, cut subsidies to shrink its bloated solar industry, pushing smaller manufacturers to the brink of collapse."
This is exactly the opposite of price fixing. They subsidized the market (consumption) and created a bloated industry. Removing the subsidies lead to overcapacity and hence a drop in prices below the rate of a healthy market.
This is exactly what is happening again this year. The difference is that we are seeing more than 50% price drop.
If you look at the fraunhofer PDF, you will notice that chinese companies own 95% of the market now. The way they were allowed to scale was by targeted chinese subsidies on PV projects that were not accessible to outside tenders.