HNHacker News
TopNewBestAskShowJobs

serialx

657 karma · joined March 9, 2011

[ my public key: https://keybase.io/serialx; my proof: https://keybase.io/serialx/sigs/oEU8jNpQOfoVNZZedCVSRW4mZWKfASS_FRlSJvi_tII ]
submissionscomments
serialx··on Zml-smi: universal monitoring tool for GPUs, TPUs and NPUs
Look into all-smi https://github.com/lablup/all-smi It supports all GPUs thinkable including Apple Silicon and many AI accelerator cards.
serialx··on How attention sinks keep language models stable
Yeah, attention sinks were applied to gpt-oss
serialx··on TinyZero
Ah sorry, you might be right. I meant "sparse reward" as a reward system that is mostly 0 but occasionally 1. Your "sparse reward" means only providing reward at the end of each output.
serialx··on TinyZero
I don't think it's only using sparse rewards because of the format rewards. The training recipe is pretty comprehensive and involves multiple stages.[1] The paper mentions that when only using the RL technique, the output is often not suitable for reading. (Language mixing, etc) That feels like a AlphaZero moment for LLMs?

[1]: https://www.reddit.com/r/LocalLLaMA/comments/1i8rujw/notes_o...

serialx··on TinyZero: Reproduction of DeepSeek R1 Zero in countdown and multiplication tasks
So to my understanding, this work reproduces DeepSeek R1's reinforcement learning mechanism in a very small language model.

The AI gets "rewards" (like points) for doing two things correctly:

Accuracy : Getting the right answer. For example, math answers must be in a specific format (e.g., inside a box) so a computer can easily check them. For coding problems, test cases verify if the code works.

Format : Using the <think> and <answer> tags properly. This forces the AI to organize its responses clearly.

So in this case, the training program can extract the model's answer by parsing <answer> tag. We can eval the answer and evaluate if it's correct or not. If it's correct give reward, else: no reward.

Create N such answers from a single question, create N reward array. This is enough for the RL algorithm to guide the model to be more smart.

serialx··on OpenWRT One Released: First Router Designed Specifically for OpenWrt
Change the currency to USD
serialx··on The Copper Plate Must Die
What are the compute requirements for solving efficient energy grid transmission? Is there a efficient algorithm that is able to solve this?
serialx··on Model Context Protocol
Is there any plans to add Well-known URI[1] as a standard? It would be awesome if we can add services just by inputting domain names of the services.

[1]: https://en.wikipedia.org/wiki/Well-known_URI

serialx··on Show HN: Pocache, preemptive optimistic caching for Go
PSA: You can also use singleflight[1] to solve the problem. This prevents the thundering herd problem. Pocache is an interesting/alternative way to solve thundering herd indeed!

[1]: https://pkg.go.dev/golang.org/x/sync/singleflight

serialx··on GPUs Go Brrr
Actually, llama.cpp running on Apple silicon uses GPU(Metal Compute Shader) to inference LLM models. Token generation is also very memory bandwidth bottlenecked. On high end Apple silicon it's about 400MB/s to 800MB/s, comparable to NVIDIA RTX 4090, which has memory bandwidth of 1000MB/s. Not to mention that Apple silicon has unified memory architecture and has high memory models (128GB, up to 192GB), which is necessary to run large LLMs like Llama 3 70B, which roughly takes 40~75GB of RAM to work reasonably.
serialx··on GPUs Go Brrr
Well, iPhone already does that with photos. :)
serialx··on VkFFT: Vulkan/CUDA/Hip/OpenCL/Level Zero/Metal Fast Fourier Transform Library
Now we just need VkDNN
serialx··on Show HN: My solar-powered, ePaper digital photo frame
Interesting! Maybe using small cells like these might make the frame and cell blend together?

https://a.aliexpress.com/_oDCStIL

serialx··on OpenLLaMA: An Open Reproduction of LLaMA
No since it’s stateful in the sense that inferencing is dependent on the past generated tokens.
serialx··on Achieving 100Gbps intrusion prevention on a single server
This is just awesome. I wish there are open and arduino like dev env for FPGAs.
serialx··on Tesla – Introducing V3 Supercharging
The difference reminds me of the energy efficiency a EV has.
serialx··on FindChips – Get instant insight into any electronic component
They also sell cut tape. So you can order from 1 to 100 parts. There are however some parts that have MOQ of ~50, but those are so cheap that 100 of them would be less than 10 USD.
serialx··on CNC milling with open source software
If you want small servos, just use steppers with drivers like TMC2130.

Anything bigger, many people seem to like ClearPath servos. Price to performance seem to be pretty good.

serialx··on Retropiler: Java 8 Standard Library Backport for Android
This is fixed now. The README now state GPL + Classpath Excemption
serialx··on Outrageously Large Neural Networks: Up to 137B Parameters
Looking at paper like this, I can't help to think about PDP. Will we be able to confirm Parallel distributed processing (PDP) theory in the near future?
serialx··on Ask HN: Anyone using Cloudflare for DNS only?
I'm using Cloudflare as a free DNS.
serialx··on Server Side TLS
Worth mentioning https://cipherli.st/ too. But I think more warning about HSTS is needed, since misconfiguring HSTS will cause the domain to be inaccessible for long periods.
serialx··on Google’s QUIC protocol: moving the web from TCP to UDP
QUIC uses TCP friendly backoff algorithm called CUBIC. Linux TCP implementation currently uses CUBIC too. One difference between TCP and QUIC is that the parameter beta of CUBIC is different.

One QUIC connection is equivalent to two TCP connections in that regard. So QUIC will only backoff half amount compared to TCP. In the design docs, they mention it's okay since one QUIC connection is equivalent to multiple TCP connections that a browser makes.

serialx··on Log Structured Merge Trees
It would be more interesting to see more modern approaches with KV storages like forestdb[1]. Couchbase already replaced their storage engine with forestdb[2].

[1]: https://www.computer.org/csdl/trans/tc/preprint/07110563.pdf [2]: https://github.com/couchbase/forestdb

serialx··on A Better Default Colormap for Matplotlib
This, changes everything.

The theme seaborn uses is actually a direct clone from ggplot2 from R

serialx··on NVIDIA Announces the GeForce GTX 1000 Series
Guys have a look at SSLShader[1]

[1]: http://shader.kaist.edu/sslshader/

serialx··on QUIC experiments [pdf]
There's QUIC support in Wireshark. Use the beta version. :)
serialx··on Cache-friendly binary search
Or you can use binary search with slightly off-center binary or quaternary (four-way) searches to provide more consistent lookup times:

http://www.pvk.ca/Blog/2012/07/30/binary-search-is-a-patholo...

serialx··on How We Designed for Performance and Scale
If you write the code within the `gen_server` guidelines, state migration is supported by `code_change`:

http://www.erlang.org/doc/man/gen_server.html#Module:code_ch...

For example:

http://stackoverflow.com/questions/1840717/achieving-code-sw...

BTW, You can even support downgrade. :)

serialx··on Korean Fan Death
> Every fan in South Korea has an automatic shut-off feature - you can’t even buy a fan that will run all night. Impossible to purchase.

This statement is simply not true. I live in South Korea and I always run it all night in summer.

Page 1 of 2Next →