HNHacker News
TopNewBestAskShowJobs

ElectricalUnion

1,185 karma · joined May 27, 2021

submissionscomments
ElectricalUnion··on Pi.dev: You Said No MCP
The "general adoption agent" for Earendil is their other product Lefos, that is based on Pi and uses email as the interface.
ElectricalUnion··on UTF-8000: Unlimited UTF-8
What you meant by "single effective character" is grapheme clusters. This whole discussion is about variable sized code points.
ElectricalUnion··on How GLM built its own inference infrastructure
Jevons paradox, technological improvements that increase the efficiency of a resource's use lead to a rise in total consumption of that resource.
ElectricalUnion··on How GLM built its own inference infrastructure
Isn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train or improve other LLMs?
ElectricalUnion··on GLM Built Its Own Inference Infrastructure
The only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.
ElectricalUnion··on A single firm is behind OpenAI, Anthropic, and Meta hacking scandals
Isn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train other LLMs?
ElectricalUnion··on So you want to use OpenRouter?
Can I block providers (for model, not in general) that have sightly cheaper input/output tokens but have more that 10x the average cache cost?

Can I block providers (for a model, not in general) that set up cache write cost when the mode is free cache writes?

ElectricalUnion··on Simple Is Not Small
> But they are.

Did you remember to:

- check for the other spawned process exit code?

- waitpid for all process in the other process chain?

- propagate/handle signals, like for example SIGINT/SIGTSTP/SIGPIPE/SIGHUP forward and back signals?

- change stdio buffering mode?

- remember to count how many bytes actually were written by write, and blocking if not, before clobbing the 64kb of the pipe buffer size with another write?

- flush, then close all file descriptors left behind by the pipes when it ends?

It's for reasons like that, that I don't trust anything non-trivial, not-shell to use pipelines correctly.

ElectricalUnion··on bzip3
Duckdb supports loading and saving to zstd for all it's base loading/saving formats csv/tsv/json/jsonlines, but, for good or bad, those are solid compression.

Under most r/w workloads, using parquet/lance/vortex/native-duckdb, with their built-in columnar compression will result in more performance AND space savings. Non-solid compression. Then, the query engine can push down your query predicate to a column row group level, instead of forcing it to decompress the entire dataset to operate.

Practical example: duckdb has syntax - https://duckdb.org/docs/lts/data/multiple_files/overview - to glob multiple files at once, but that really only works if you're applying push down query predicates instead of re-decompressing your entire data set per SELECT. I would say for most dataset, even 20%+ size is worth not having to decompress (or even download!) the entire dataset, to figure out if something fits the predicate.

After all, if you have to download and decompress the dataset back again to operate, then the "space savings" are gone.

ElectricalUnion··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Using macos on low memory regimes will make it use disk-based swap.

For example, a Macbook Neo (so in theory, something with around 4GiB of free RAM lying around) might eat around 900GB of writes a day while not doing much at all, because it's basically on low on RAM and swapping all the time.

ElectricalUnion··on Internet centralization and the original sin of NAT
> If not for NAT, we'd all need a firewall

A NAT implementation could broadcast any "WAN" side incoming packets to all link local clients (aka: put everyone in the DMZ). The only thing preventing that is a stateful firewall.

ElectricalUnion··on Transfer files over an Ethernet patch cable
If you're using specifically zstd, on the sender side, instead of tweaking the whole tunnel once, you can use --adapt to dynamically adjust to i/o conditions.
ElectricalUnion··on Samsung's Processing-in-Memory (PIM)
> PIM would have a full blown IO controller and cache subsystem to fetch remote operands

Don't we already have those in mainstream computing in the form of dedicated silicon in DMA controllers? Programmed input–output performance is often low throughput, high jitter and uses a lot of CPU.

ElectricalUnion··on Air Conditioning Is Not a Luxury, It Is a Necessity
Around and above 40°C, fans _are_ worse that useless - they may actively be killing you instead.

> Use electric fans only when temperatures are below 40˚C / 104˚F. In temperatures above 40˚C / 104˚F, fans will heat the body.

- https://www.who.int/news-room/fact-sheets/detail/climate-cha...

ElectricalUnion··on Why your local LLM feels dumber than it is
Qwen3 is not Qwen3.8?

> commited on May 21, 2025, over 1 year ago

Fairly old update to the README.md of (instead of Qwen3.8), should have raised some flags?

ElectricalUnion··on Qwen 3.8 27B
AFAIK the only well-known LLM with no "conventional" multimodal encoders required for audio and images is Gemma 4 12B.
ElectricalUnion··on Single log line is 49KB+ (ext4) / 110KB+ (btrfs) of systemd-journald disk writes
No duckdb (or parquet). If you want to avoid writes and write amplification, you really want to avoid re-writing all 122880 rows of a row group every time a single insert happens.
ElectricalUnion··on What I learned by putting GitHub Copilot behind a MitM proxy
Infisical, or Bitwarden Secret Manager? Those two look like perfectly reasonable if the llm is just careless (but still, nothing prevents the LLM from intentionally cat'ing /proc/self/environ or from running /usr/bin/env or set or similar)
ElectricalUnion··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
You can use any message you want, but the model was tested to react reasonably well to the specific token sequence of "\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n</think>.\n\n" (from a Alibaba paper, struggling to find it now)

Edit: arXiv:2505.09388 Qwen3 Technical Report

ElectricalUnion··on LLMs reward expertise
What "democratized computing" was cheaper computing and VisiCalc, not WIMP and GUI?

But again, VisiCalc is intentionally limited, it's not a all-powerful environment, on purpose. It's all about intentional limitations, making computation easier to reason about.

ElectricalUnion··on LLMs reward expertise
Almost all web chatbot providers have code sandboxes that they will run (limited) tooling for you. If you ask for it, it will run deterministic linters, formatters, format converters, tests for you. Older versions of Claude would for example happily try to reverse engineer entire artifacts for you, once you provide a URL.
ElectricalUnion··on LLMs reward expertise
> the terminal is a scary place

Isn't the the powerful, unlimited, unopinionated blank LLM text input waiting for your instructions eerily similar to a scary terminal?

WIMP and GUI paradigms are the exact the opposite: intentional dis-empowering, by design restrictions, enumeration of your few possible options. Those feel more constrained therefore safer.

ElectricalUnion··on Developers are attached to tools because tools encode trust
Problem is that bash is too sharp to handle to smart and gullible clankers without a sandbox - That I think everyone should be using anyways, for everything, even things not related to clankers - Android and Qubes are right. The app/vm, and whatever it tries, should not be considered trusted by default.
ElectricalUnion··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Aren't those speeds for the first few tokens, that, because of no context for attention to attend, are much faster to compute that the others? I expect the actual token speed to nosedive sharply as you get more context utilization.

What about the system prompt you're using? General purpose harness like Claude Code will insert their happy 22k tokens, even before your first useful token is processed. That might make it a no-go even _before_ you can even start, as the maximum context for this seems pretty limited (to make it fast)

And those LLMs are all "thinking", that is, rather that "one-shoting" the answer, they generate a lot of internal use reasoning tokens before starting to generate useful, visible response tokens. You can easily get to 30k tokens when your initial prompt is vague ambiguous garbage (as are naive transcriptions) as your LLM will "But wait, the user might have meant X, let me think more about this" lots of times.

No thinking (therefore much worse answers) will be a requirement.

ElectricalUnion··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Then have it summarise the meeting transcription over the entire week for the report next week.

ftfy.

ElectricalUnion··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
By the time the tokens start coming out 30h later you might need to use your laptop again...
ElectricalUnion··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Ok, this one will take just 30h (compared to that other project that would take 6.25 days) to start writing output tokens after you say hi in Claude Code.
ElectricalUnion··on Kimi-K3 on HuggingFace
Won't the answer (even for a pretty basic message like "hi") at SSD speeds take like a _entire week_ to _start showing useful output?_ (attempting to do 22k average claude code system prompt + 32k thinking tokens thru 0.1t/s throughput)
ElectricalUnion··on Do you hate XML? (2010)
The funny thing is that JSON parsing is usually kinda unsafe in it's main target language JavaScript, and usually safe in other languages, because of the `__proto__` prototype pollution.
ElectricalUnion··on What ORMs have taught me: just learn SQL (2014)
> I don't think providing 90% of the structure you need is a failed abstraction.

It is, when the "10%" is the actual hot queries that your system will use the most?

Code right now "is so cheap". You can provide your favourite LLM with your database schema, and some domain comments, and ask it a query to fetch/update data, and it will generate somewhat sane queries for you. You can then inspect those queries yourself, send them to another LLM or human to review and, when they look OK, ship it.

And when it comes time to debug it, you have, you know, an actual query, not some pseudo-query in a custom DSL. No need to implement runtime telemetry just to try to figure out if the ORM actually made the query you thought it was supposed to do.

Page 1 of 28Next →