HNHacker News
TopNewBestAskShowJobs

underlines

783 karma · joined August 22, 2012

submissionscomments
underlines··on Vote on which of Hacker News' challenges for AI have been met
i filed my swiss taxes for 2025 in 2026 (april) by dumping everything (local tax law, tax guide, my and my wife's documents, bank statements, income statements, etc.) into a folder and asking claude to fill it out. i had nothing to fix. submitted it.
underlines··on Fable 5.1 Solves the Cyphral Distich, a 370-year-old cipher
This falls into the obscure ("not important") but not difficult part of problems, that has an extreme long tail, compared to exceptionally difficult and important.

indeed as the author mentioned, LLMs can greatly help in areas where there is a long tail of not so important, easy to solve problems, that humans just don't have the time or priority to focus on. But combine this long tail of problems that can be solved: accumulated this might still be very beneficial as a sum of things.

underlines··on Astra and Fable still hack on simple variants of alignment evals from 2025
who tf uses prompting to "pretty please don't cheat on this"? the best practices for ages (in terms of ai) is to separate the eval from the test code/agent.

another best practices every single solution using LLMs/agents should implement is "never trust the llm".

underlines··on Show HN: My Claude quota ran out in 10 minutes, so I made a tool to find out why
You can't ask Claude if your quota ran out. You have to wait for the reset...
underlines··on Thomson Reuters Launches Its Own Frontier Model
Please challenge me and explain why this is a break through worth the headline they use for their own work.

How I understand it, without really reading into it:

- Thomson Reuters did not "create" a frontier model, they gave some money, maybe a bit of their data to Imperial College London, and took Alibaba's Qwen 3.6 35B A3B Model for a basic fine tune.

- "They" (some undergrads at Imperial College London) used an existing ablation framework to undo some of the topic alignment of the original model.

- "They" fine tuned on some domain knowledge trying to preserve general knowledge - here maybe, just maybe some data came from Thomson Reuters.

As a result: one of 100's of Qwen 3.6 35B A3B sparse model fine tunes, just for publicity, to write overstating headlines like "X created their own frontier model"

underlines··on uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook
Hyperpersonalized code to its rescue. With the qwen3.8-27b release tomorrow and the advent of vibe coding and the improvements into proper agentic engineering we are moving towards hyperpersonalized software. This means we don't need a centralized maintainer for smaller tools. the facebook ad blocking cat and mouse game is something I cna very well imagine to just be something my computer would solve in real time, updating the uBlock filter/userscript to adapt to any fb changes in near real-time.
underlines··on Qwen3.8-2.4T
https://modelscope.cn/models/Qwen/Qwen3.8-27B/summary Countdown at 28h now.

Sadly they seem to not be releasing a sparse 35b A3b or anything inbetween "too large to host for mortals" and "fits into a consumer rtx". Probably not to eat away their profits on their API serving. 120b - 300b is a dead space right now, very few good releases in that size range. (I know there are, but the big labs aren't releasing stuff here)

underlines··on uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook
Hyperpersonalized code to its rescue. With the qwen3.8-27b release tomorrow and the advent of vibe coding and the improvements into proper agentic engineering we are moving towards hyperpersonalized software.

This means we don't need a centralized maintainer for smaller tools. the facebook ad blocking cat and mouse game is something I cna very well imagine to just be something my computer would solve in real time, updating the uBlock filter/userscript to adapt to any fb changes in near real-time.

underlines··on Mario Meets Pareto
Previously on HN: https://news.ycombinator.com/item?id=39936246
underlines··on Dear Microsoft, enough is enough
Controversial thought: Browsers will become a niche and fall into obscurity like IRC nowadays, based on what I observed working in South East Asia, where people don't even know what browsers are and "the internet" are walled social networks and apps.
underlines··on Leak reveals Google's Aluminium OS with a 16-minute video
I don't get it either. i have a pixel 10 and it looks exactly like desktop mode, when connected to a docking station...
underlines··on The locals don't know
I am Swiss and like most 30-40yo from my generation, Hiking in the splendid nature, lakes and mountains, visiting touristy places and the overall scenery was something uncool, for old people. We were forced to do that in school. Nobody in their right mind would do it. :)

I then moved abroad to Bangkok, working an office job. Although BKK is great for consumerism and convenience, especially with cheap labor available for almost anything, you can get quite lazy. The bad traffic, non-pedestrian friendly (non existent) city planning and little nature left also makes it a bit cumbersome to find nature nearby. This made me appreciate nature, hiking and nice scenery. (Of course Thailand has lots of beautiful nature and scenery, but not so much of an active outdoor scene)

Coming back to Switzerland after 6 years, I became the biggest tourist, going hiking every weekend, spending time at our tourist destinations, but also all the second tier ("unseen") places only locals know. I tried so much stuff that in the past I thought is tourist stuff, and most of it is simply great.

I also became much more understanding, open and helpful to expats, foreigners and tourists in my country.

underlines··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
well, your own, unleaked ones, representing your real workloads.

if you can't afford to do that, look at a lot of them, eg. on artificialanalysis.com they merge multiple benchmarks across weighted categories and build an Intelligence Score, Coding Score and Agentic score.

underlines··on Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
depends on format, compute type, quantization and kv cache size.
underlines··on Average is all you need
i fail to understand how text2sql on quite simple data sources is anything to write home about 3 years after it came onto the market? can someone elaborate?

where it gets interesting is when you have a custom system that your LLM surely never saw (custom ERP) that has 50 sometime cryptic tables, unclear look up tables and unexplained flags.

something no text2sql solution solved for us.

we built a second mcp that lets the agent look up business logic (generated from source code) and then does better queries. that i think is something i never read in a blog post about a text2sql solution.

underlines··on “Car Wash” test with 53 models
I maintain a private evaluation set of what many call "misguided attention" questions.

In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions.

A lot of viral "frontier model fails THIS simple question" examples are essentially carefully constructed token sequences designed to bias the statistical prior toward an intuitively wrong answer. Small wording changes can flip results entirely.

If you systematically expand the prompt space around such questions—adding or removing minor contextual cues you'll typically find symmetrical variants where the same models both succeed and fail. That suggests sensitivity to framing and distributional priors (adding unnecessary info, removing clear info, add ambiguity, ...), not necessarily absence of reasoning capability.

underlines··on The peculiar case of Japanese web design (2022)
The reposts go on and on, while Japanese web design stays the same

https://news.ycombinator.com/item?id=25148942

https://news.ycombinator.com/item?id=6718067

https://news.ycombinator.com/item?id=16254569

https://news.ycombinator.com/item?id=30523955

https://news.ycombinator.com/item?id=33745146

https://news.ycombinator.com/item?id=35209424

https://news.ycombinator.com/item?id=16272392

https://news.ycombinator.com/item?id=8975735

underlines··on HeyWhatsThat
The (pun ahead) peak of this method imho is implemented in "PeakFinder", afaik uses a low res nation wide (switzerland) height map, after initial gps fix it downloads local high res height map, calculates peak contours based of current location AND height and overlays that grid including the peak names onto the camera feed using the gyro and compass.

It's quite easy to build accurate geo-related applications in Switzerland due to the excellent work of the government office "Swiss Topo" that maps every tree, every house, every road in the whole country. Trees in cities have metadata such as: year planted, type etc. :)

Johnny Harris, the map aficionado mentioned Swiss maps and Swiss Topo's dedication multiple times in his videos.

underlines··on Show HN: Ghidra MCP Server – 110 tools for AI-assisted reverse engineering
Tool stuffing degrades LLM tool use quality. 100+ tools is crazy. We probably need a tool that does relevant tool retreaval and reranking lol
underlines··on AGENTS.md outperforms skills in our agent evals
Oh got, this scales bad and bloats your context window!

Just create an MCP server that does embedding retrieval or agentic retrieval with a sub agent on your framework docs.

Finally add an instruction to AGENT.md to look up stuff using that MCP.

underlines··on Copyparty, the FOSS file server [video]
I also look for a sophisticated self hosted, open source transcoding solution as a web app, but in the mean time, the complete opposite: no bells and whistles, no config, no control except size: https://github.com/JMS1717/8mb.local

or do you mean a web based file manager / video gallery with transcoding capabilities?

underlines··on Production RAG: what I learned from processing 5M+ documents
rag will be pronounced differently ad again and again. it has its use cases. we moved to agentic search having rag as a tool while other retrieval strategies we added use real time search in the sources. often skipping ingested and chunked soueces. large changes next windows allow for putting almost whole documents into one request.
underlines··on Why Self-Host?
Even though I work as an IT Professional, I was almost always the only person not self hosting anything at home and not having a NAS.

I jumped the hoop and bought a Ugreen nas with 4 bays where the first thing I did was installing TrueNAS CE onto it and then use ChatGPT with highly customized prompts and the right context (my current docker-compose files).

Without much previous knowledge of docker, networking etc. except what I remembered from my IT vocational education from 15 years ago, I now have:

- Dockerized Apps

- App-Stacks in their own App-Network

- Apps that expose web UI not via ports, but via Traefik + Docker labels

- Only Traefik 443 ports reachable from WAN, plus optional port forwarding for non-http services

- Optional Cloudflare Tunnel

- Automatic Traefik TLS termination for LAN and WAN for my domain

- Split-DNS to get hostnames routed properly on LAN and WAN

- CrowdSec for all exposed containers

- Optional MFA via Cloudflare for exposed services

- Local DHCP/DNS via Technitium

- Automatic ZFS snapshots and remote backups

- Separation between ephemeral App data (DBs, Logs) on SSD and large files on HDD

underlines··on SonyShell – An effort to “SSH into my Sony DSLR”
Yes, aware of that, and nothing recent works with it, the last progress sadly was years ago.

I guess DMCA/Sony Lawyers and the relatively low market share for expensive cameras is the main reason why a PlayStation, an iPhone or a Nintendo Jailbreak is more appealing to reverse engineers than a Sony Camera Jailbreak.

underlines··on SonyShell – An effort to “SSH into my Sony DSLR”
I had the NX1 with all the premium lenses and some photos still seem to be better than what my Sony A7-M4 shoots. But no 10bit 4:2:2 for video and no real flat profile was a bummer. I loved the persistent mod though. Sold all NX1 gear years ago, moved to a Sony A7-M3 and then A7-M4. Full Frame has some great benefits.
underlines··on SonyShell – An effort to “SSH into my Sony DSLR”
Reading the title, I thought: finally someone rooted/jailbroke sony cameras.

On Canon you can run Magic Lantern, an extensive mod that adds many features to Canon cameras.

Even Samsung N1 had SD Card loadable mods before they moved away from the camera market.

Rooting sony seems impossible, I never saw someone Working on it Since their Fullframe lineup launched.

underlines··on SonyShell – An effort to “SSH into my Sony DSLR”
that's a feature of github when renaming repos
underlines··on Jan – Ollama alternative with local UI
Jan here too, and I work with LLMs full time and I'm a speaker about these topics. Annoying how many times people ask me if Jan.ai is me lol
underlines··on Ollama's new app
Heads up, there’s a fair bit of pushback (justified or not) on r/LocalLLaMA about Ollama’s tactics:

    Vendor lock-in: AFAIK it now uses a proprietary llama.cpp fork and builts its own registry on ollama.com in a kind of docker way (I heard docker ppl are actually behind ollama) and it's a bit difficult to reuse model binaries with other inference engines due to their use of hashed filenames on disk etc.

    Closed-source tweaks: Many llama.cpp improvements haven’t been upstreamed or credited, raising GPL concerns. They since switched to their own inference backend.

    Mixed performance: Same models often run slower or give worse outputs than plain llama.cpp. Tradeoff for convenience - I know.

    Opaque model naming: Rebrands or filters community models without transparency, biggest fail was calling the smaller Deepseek-R1 distills just "Deepseek-R1" adding to a massive confusion on social media and from "AI Content Creators", that you can run "THE" DeepSeek-R1 on any potato.

    Difficult to change Context Window default: Using Ollama as a backend, it is difficult to change default context window size on the fly, leading to hallucinations and endless circles on output, especially for Agents / Thinking models.
---

If you want better, (in some cases more open) alternatives:

    llama.cpp: Battle-tested C++ engine with minimal deps and faster with many optimizations

    ik_llama.cpp: High-perf fork, even faster than default llama.cpp

    llama-swap: YAML-driven model swapping for your endpoint.

    LM Studio: GUI for any GGUF model—no proprietary formats with all llama.cpp optimizations available in a GUI

    Open WebUI: Front-end that plugs into llama.cpp, ollama, MPT, etc.
underlines··on GEPA: Reflective prompt evolution can outperform reinforcement learning
Looking at a Problem from various perspectives, even posing ideas, is exactly what reasoning models seem to simulate in their thinking CoT to explore the solution space with optimizations like MCMC etc.
Page 1 of 9Next →