783 karma · joined August 22, 2012
indeed as the author mentioned, LLMs can greatly help in areas where there is a long tail of not so important, easy to solve problems, that humans just don't have the time or priority to focus on. But combine this long tail of problems that can be solved: accumulated this might still be very beneficial as a sum of things.
another best practices every single solution using LLMs/agents should implement is "never trust the llm".
How I understand it, without really reading into it:
- Thomson Reuters did not "create" a frontier model, they gave some money, maybe a bit of their data to Imperial College London, and took Alibaba's Qwen 3.6 35B A3B Model for a basic fine tune.
- "They" (some undergrads at Imperial College London) used an existing ablation framework to undo some of the topic alignment of the original model.
- "They" fine tuned on some domain knowledge trying to preserve general knowledge - here maybe, just maybe some data came from Thomson Reuters.
As a result: one of 100's of Qwen 3.6 35B A3B sparse model fine tunes, just for publicity, to write overstating headlines like "X created their own frontier model"
Sadly they seem to not be releasing a sparse 35b A3b or anything inbetween "too large to host for mortals" and "fits into a consumer rtx". Probably not to eat away their profits on their API serving. 120b - 300b is a dead space right now, very few good releases in that size range. (I know there are, but the big labs aren't releasing stuff here)
This means we don't need a centralized maintainer for smaller tools. the facebook ad blocking cat and mouse game is something I cna very well imagine to just be something my computer would solve in real time, updating the uBlock filter/userscript to adapt to any fb changes in near real-time.
I then moved abroad to Bangkok, working an office job. Although BKK is great for consumerism and convenience, especially with cheap labor available for almost anything, you can get quite lazy. The bad traffic, non-pedestrian friendly (non existent) city planning and little nature left also makes it a bit cumbersome to find nature nearby. This made me appreciate nature, hiking and nice scenery. (Of course Thailand has lots of beautiful nature and scenery, but not so much of an active outdoor scene)
Coming back to Switzerland after 6 years, I became the biggest tourist, going hiking every weekend, spending time at our tourist destinations, but also all the second tier ("unseen") places only locals know. I tried so much stuff that in the past I thought is tourist stuff, and most of it is simply great.
I also became much more understanding, open and helpful to expats, foreigners and tourists in my country.
if you can't afford to do that, look at a lot of them, eg. on artificialanalysis.com they merge multiple benchmarks across weighted categories and build an Intelligence Score, Coding Score and Agentic score.
where it gets interesting is when you have a custom system that your LLM surely never saw (custom ERP) that has 50 sometime cryptic tables, unclear look up tables and unexplained flags.
something no text2sql solution solved for us.
we built a second mcp that lets the agent look up business logic (generated from source code) and then does better queries. that i think is something i never read in a blog post about a text2sql solution.
In many of these cases, the issue isnt failed logical reasoning. Its ambiguity, underspecified context, or missing constraints that allow multiple valid interpretations. Models often fail not because they can’t reason, but because the prompt leaves semantic gaps that humans silently fill with shared assumptions.
A lot of viral "frontier model fails THIS simple question" examples are essentially carefully constructed token sequences designed to bias the statistical prior toward an intuitively wrong answer. Small wording changes can flip results entirely.
If you systematically expand the prompt space around such questions—adding or removing minor contextual cues you'll typically find symmetrical variants where the same models both succeed and fail. That suggests sensitivity to framing and distributional priors (adding unnecessary info, removing clear info, add ambiguity, ...), not necessarily absence of reasoning capability.
https://news.ycombinator.com/item?id=25148942
https://news.ycombinator.com/item?id=6718067
https://news.ycombinator.com/item?id=16254569
https://news.ycombinator.com/item?id=30523955
https://news.ycombinator.com/item?id=33745146
https://news.ycombinator.com/item?id=35209424
It's quite easy to build accurate geo-related applications in Switzerland due to the excellent work of the government office "Swiss Topo" that maps every tree, every house, every road in the whole country. Trees in cities have metadata such as: year planted, type etc. :)
Johnny Harris, the map aficionado mentioned Swiss maps and Swiss Topo's dedication multiple times in his videos.
Just create an MCP server that does embedding retrieval or agentic retrieval with a sub agent on your framework docs.
Finally add an instruction to AGENT.md to look up stuff using that MCP.
or do you mean a web based file manager / video gallery with transcoding capabilities?
I jumped the hoop and bought a Ugreen nas with 4 bays where the first thing I did was installing TrueNAS CE onto it and then use ChatGPT with highly customized prompts and the right context (my current docker-compose files).
Without much previous knowledge of docker, networking etc. except what I remembered from my IT vocational education from 15 years ago, I now have:
- Dockerized Apps
- App-Stacks in their own App-Network
- Apps that expose web UI not via ports, but via Traefik + Docker labels
- Only Traefik 443 ports reachable from WAN, plus optional port forwarding for non-http services
- Optional Cloudflare Tunnel
- Automatic Traefik TLS termination for LAN and WAN for my domain
- Split-DNS to get hostnames routed properly on LAN and WAN
- CrowdSec for all exposed containers
- Optional MFA via Cloudflare for exposed services
- Local DHCP/DNS via Technitium
- Automatic ZFS snapshots and remote backups
- Separation between ephemeral App data (DBs, Logs) on SSD and large files on HDD
I guess DMCA/Sony Lawyers and the relatively low market share for expensive cameras is the main reason why a PlayStation, an iPhone or a Nintendo Jailbreak is more appealing to reverse engineers than a Sony Camera Jailbreak.
On Canon you can run Magic Lantern, an extensive mod that adds many features to Canon cameras.
Even Samsung N1 had SD Card loadable mods before they moved away from the camera market.
Rooting sony seems impossible, I never saw someone Working on it Since their Fullframe lineup launched.
Vendor lock-in: AFAIK it now uses a proprietary llama.cpp fork and builts its own registry on ollama.com in a kind of docker way (I heard docker ppl are actually behind ollama) and it's a bit difficult to reuse model binaries with other inference engines due to their use of hashed filenames on disk etc.
Closed-source tweaks: Many llama.cpp improvements haven’t been upstreamed or credited, raising GPL concerns. They since switched to their own inference backend.
Mixed performance: Same models often run slower or give worse outputs than plain llama.cpp. Tradeoff for convenience - I know.
Opaque model naming: Rebrands or filters community models without transparency, biggest fail was calling the smaller Deepseek-R1 distills just "Deepseek-R1" adding to a massive confusion on social media and from "AI Content Creators", that you can run "THE" DeepSeek-R1 on any potato.
Difficult to change Context Window default: Using Ollama as a backend, it is difficult to change default context window size on the fly, leading to hallucinations and endless circles on output, especially for Agents / Thinking models.
---If you want better, (in some cases more open) alternatives:
llama.cpp: Battle-tested C++ engine with minimal deps and faster with many optimizations
ik_llama.cpp: High-perf fork, even faster than default llama.cpp
llama-swap: YAML-driven model swapping for your endpoint.
LM Studio: GUI for any GGUF model—no proprietary formats with all llama.cpp optimizations available in a GUI
Open WebUI: Front-end that plugs into llama.cpp, ollama, MPT, etc.