HNHacker News
TopNewBestAskShowJobs

benob

425 karma · joined December 6, 2017

submissionscomments
benob··on Early rogue AI agent activity and attempts to hack found on urlquery.net
Couldn't find the reference but I remember some time ago a first generation automated gun killing the audience at an army show. Was the gun maker convicted of manslauther?

--edit-- Was a bit older than I remembered: https://slashdot.org/story/07/10/18/1847231/robotic-cannon-l...

benob··on Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
https://github.com/earthspecies/NatureLM-audio
benob··on iPhone Duo
Lots of people in non-tech jobs are dropping their laptop/desktop for their phone for all work-related tasks. They love having more screen real estate, plus their company pays for the premium.
benob··on Stealing Reasoning Traces from Proprietary LLM APIs
A natural next step is to use the reasoning traces to jailbreak the stronger models (https://arxiv.org/pdf/2603.12277)
benob··on As a Windows user, it's a surreal way to install a program
Maybe the correct UX could be to list the locations of settings data, their size and ask the user whether they want to leave them, put them aside in a dedicated folder ("the attic", "the basement" or whatever), or remove them
benob··on The session you cannot take with you
What matters for this injection strategy to work is to follow quite closely the style of the reasoning. It's particularly effective if you copy reasoning from the same context. If you cannot see the reasoning, you cannot duplicate it's style.

That said, including instances of the attack in training is already a good countermeasure.

benob··on The session you cannot take with you
People will keep hiding reasoning because it allows prompt injection https://arxiv.org/pdf/2603.12277, in addition to facilitating distillation (you don't pay the full cost of RL)
benob··on Google fixed more Chrome bugs in June than over the past two years, thanks to AI
The next step is to build software without bugs
benob··on How many of the 170k English words do you know?
Longest definition and semi-columns are strong biases for right answer. Also, my run contained a lot of adjectives for which it is pretty obvious that noun definitions do not match.
benob··on Apple reveals new AI architecture built around Google Gemini models
It may be a clever move. By using the same models as android (contractually?), they can compete on the user experience which they typically handle better than android phone providers.
benob··on The LLM warnings Google fired Timnit Gebru over have all come true
And papers on bias amplification in ML predate LLMs. I remember this specific one which was a spotlight paper at EMNLP:

Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints, Zhao et al.

https://arxiv.org/abs/1707.09457

benob··on Removing the modem and GPS from my 2024 RAV4 hybrid
Does changing the date fix it?
benob··on Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
Deployed it to a huggingface space: https://huggingface.co/spaces/benoitfavre/needle-playground

You can check the very simple docker file there.

benob··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Here is llama-bench on the same M4:

  | model                    |       size |     params | backend    | threads |            test |                  t/s |
  | ------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |
  | qwen35 27B Q4_K_M        |  15.65 GiB |    26.90 B | BLAS,MTL   |       4 |           pp512 |         61.31 ± 0.79 |
  | qwen35 27B Q4_K_M        |  15.65 GiB |    26.90 B | BLAS,MTL   |       4 |           tg128 |          5.52 ± 0.08 |
  | qwen35moe 35B.A3B Q3_K_M |  15.45 GiB |    34.66 B | BLAS,MTL   |       4 |           pp512 |        385.54 ± 2.70 |
  | qwen35moe 35B.A3B Q3_K_M |  15.45 GiB |    34.66 B | BLAS,MTL   |       4 |           tg128 |         26.75 ± 0.02 |
So ~60 for prefill and ~5 for output on 27B and about 5x on 35B-A3B.
benob··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
I get ~5 tokens/s on an M4 with 32G of RAM, using:

  llama-server \
   -hf unsloth/Qwen3.6-27B-GGUF:Q4_K_M \
   --no-mmproj \
   --fit on \
   -np 1 \
   -c 65536 \
   --cache-ram 4096 -ctxcp 2 \
   --jinja \
   --temp 0.6 \
   --top-p 0.95 \
   --top-k 20 \
   --min-p 0.0 \
   --presence-penalty 0.0 \
   --repeat-penalty 1.0 \
   --reasoning on \
   --chat-template-kwargs '{"preserve_thinking": true}'
35B-A3B model is at ~25 t/s. For comparison, on an A100 (~RTX 3090 with more memory) they fare respectively at 41 t/s and 97 t/s.

I haven't tested the 27B model yet, but 35B-A3B often gets off rails after 15k-20k tokens of context. You can have it to do basic things reliably, but certainly not at the level of "frontier" models.

benob··on Making RAM at Home [video]
I miss the comment tagging system: insightful, informative, interesting, funny. It would make sense for hn.
benob··on Every plane you see in the sky – you can now follow it from the cockpit in 3D
Space station tracking: https://flight-viz.com/cockpit.html?lat=40.64&lon=-73.78&alt...
benob··on Simplest Hash Functions
I just realized that a hash function is nothing less than the output of a deterministic random number generator xored with some data
benob··on Exploiting the most prominent AI agent benchmarks
No, the failure is the human written prompt
benob··on What if the browser built the UI for you?
The author emphasizes accessibility and coherence as a benefit but another interesting one is composability which does not emerge naturally in the world of UI. Create a UI for a pair of websites like a command line for grep and wc. LLMs already provide that but under the natural language interaction primitive. UI could allow for branded experiences, ad delivery and whatnot in ways that natural language doesn't.
benob··on EmDash – a spiritual successor to WordPress that solves plugin security
"That allows us to license the open source project under the more permissive MIT license."
benob··on Google's 200M-parameter time-series foundation model with 16k context
I would say:

- decomposition: discover a more general form of Fourrier transform to untangle the underlying factors

- memorization: some patterns are recurrent in many domains such as power low

- multitask: exploit cross-domain connections such as weather vs electricity

benob··on Ollama is now powered by MLX on Apple Silicon in preview
Ollama is a user-friendly UI for LLM inference. It is powered by llama.cpp (or a fork of it) which is more power-user oriented and requires command-line wrangling. GGML is the math library behind llama.cpp and GGUF is the associated file format used for storing LLM weights.
benob··on TurboQuant: Redefining AI efficiency with extreme compression
Maybe they quantized a bit too much the model parameters...
benob··on TurboQuant: Redefining AI efficiency with extreme compression
This is the worst lay-people explanation of an AI component I have seen in a long time. It doesn't even seem AI generated.
benob··on Arm AGI CPU
This reminds me of Intel talking about faster web browsing with the new Pentium
benob··on Prompt Injecting Contributing.md
The real question is when will you resort to bots for rejecting low-quality PRs, and when will contributing bots generate prompt injections to fool your bots into merging their PRs?
benob··on Pretraining Language Models via Neural Cellular Automata
Reminds me of "Universal pre-training by iterated random computation" https://arxiv.org/pdf/2506.20057, with bit less formal approach.

I wonder if there is a closed-form solution for those kinds of initialization methods (call them pre-training if you wish). A solution that would allow attention heads to detect a variety of diverse patterns, yet more structured than random init.

benob··on Zig – Type Resolution Redesign and Language Changes
Time to start zig++
benob··on AI and the Ship of Theseus
It's funny that real value is now in test suites. Or maybe it's always been...
Page 1 of 7Next →