HNHacker News
TopNewBestAskShowJobs

lukax

916 karma · joined January 11, 2012

submissionscomments
lukax··on Gemini 4 Argon
Yes, through OpenRouter.
lukax··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
OpenCode checks model name and registers the appropriate tools.

const usePatch = model.modelID.includes("gpt-") && !model.modelID.includes("oss") && !model.modelID.includes("gpt-4")

Pi uses its own tools, like Armin wrote in the linked article.

lukax··on HarnessTax: How Much Does the Harness Matter for Coding Agents?
What matters more is that you use the tools that the target model was fine-tuned on.

E.g. for editing files with Claude models you should use Edit(file_path, old_string, new_string, replace_all) but with GPT models you should use apply_patch_call(patch) (where patch is a custom patch string with custom grammar).

It appears newer models are better at narive harness tool calls and worse at custom tools that look similar to default tools.

https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/

lukax··on The DeepMind Institute
https://news.ycombinator.com/submitted?id=vertigoruntime

This looks like a new account from a PR agency focusing on AI labs.

lukax··on DeepSeek-v4-flash-vision-exp
This is only a problem with OpenAI Chat Completions API. With OpenAI Responses API and Anthropic Messages API a tool output can be text, image or file.
lukax··on DeepSeek v4 Price Increase
DeepSeek-V4-Flash

Input (Cache Miss)

Previous: $0.14

New - Off-Peak: $0.22 (approx. 1.57x)

New - Peak: $0.44 (approx. 3.14x)

Input (Cache Hit)

Previous: $0.0028

New - Off-Peak: $0.007 (2.5x)

New - Peak: $0.014 (5x)

Output

Previous: $0.28

New - Off-Peak: $0.66 (approx. 2.35x)

New - Peak: $1.32 (approx. 4.71x)

---

DeepSeek-V4-Pro

Input (Cache Miss)

Previous: $0.435

New - Off-Peak: $0.66 (approx. 1.51x)

New - Peak: $1.32 (approx. 3.03x)

Input (Cache Hit)

Previous: $0.003625

New - Off-Peak: $0.022 (approx. 6.07x)

New - Peak: $0.044 (approx. 12.14x)

Output

Previous: $0.87

New - Off-Peak: $1.98 (approx. 2.27x)

New - Peak: $3.96 (approx. 4.55x)

lukax··on Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Yes. They are.

https://news.ycombinator.com/item?id=46389934

lukax··on ChatGPT claims rogue AI attacked more companies
They somehow forgot to mention that Hugging Face tried to use frontier models to analyze the attack but all models rejected. They had to use GLM 5.2 deployed locally.

https://huggingface.co/blog/security-incident-july-2026

lukax··on Using Go for Mobile Apps
Also for WASM? It's just not really there compared to Rust's wasm bindgen.
lukax··on Using Go for Mobile Apps
I can highly recommend Rust for this. The same Rust engine powering web app (WASM + React), iOS (SwiftUI), Android (Kotlin Compose) and desktop (Tauri).

https://github.com/koofr/vault

lukax··on Benchmarking coding agents on Databricks' multi-million line codebase
Could it be that users of Pi are more senior and know better how to prompt and that's why the pass rate is higher?
lukax··on Benchmarking coding agents on Databricks' multi-million line codebase
Yes, this is a known issue. A significant amount of Edit tool calls fails in Pi witg newer models.

https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/

lukax··on GLM 5.2 and the coming AI margin collapse
Well, Microsoft just started offering Kimi K2.7 through Copilot hosted on Azure.

https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-av...

Cursor Composer 2 and 2.5 are also fine tunes of Kimi K2.5

It looks like politics don't matter when it comes to economics.

lukax··on Netflix Simplified Batch Compute with Kueue
It's refreshing to see a tech article that isn't about AI. It feels like 5 years ago.
lukax··on NUMA: Cores, memory, and the distance between them
NUMA can cause really crappy performance. We deployed a Go based LLM gateway in Kubernetes deployed on a server with hundreds of CPU cores. We didn't explicitly set GOMAXPROCS so Go runtime scheduled goroutines over different CPUs and it constantly used 200% CPU and GC was causing latency spikes. Then we set GOMAXPROCS 8 and all performance issues went away. Until recently Kubernetes didn't work well with NUMA.
lukax··on Nearly half of LG smart TV apps contain residential proxy SDKs
Well, that's how data for training LLMs is scraped.
lukax··on Claude Fable 5
Zero data retention policies.
lukax··on Theseus: Translating Win32 to WASM
There is also Retrotick.

https://retrotick.com/

It simulates x86 (win32 and win16) and implements Windows APIs in javascript and renders window frames with DOM and contents with canvas (e.g. GDI translates to browser canvas operations). A lot of programs run already but a lot of APIs are not yet implemented.

I successfullt spent a few days extending it to run a Click & Create based game from my childhood.

lukax··on Xiaomi MiMo-v2.5 Series API Permanent Price Reduction Up to 99%
Huawei Ascend AI Accellerators. DeepSeek V4 model architecture was optimized for Chinese hardware.
lukax··on What color is your function? (2015)
That's just not true. Let's say you have a form validation library with a public api that supports custom validators Validate(name string, value string) bool. Then you decide that your validator now needs to make an HTTP request. This request needs context so that tracing is propagated and needs to return (bool, error) so that error is propagated up instead of silently ignoring it or logging it and returning false. This is coloring. You can use context.Background the same way you can use blocking in other languages. It just doesn't feel right and it breaks things.
lukax··on Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
I wonder what would happen if they used Kimi 2.5 directly instead of Cursor Composer 2.5. Composer is a fine tune of Kimi. Probably they didn't want to test "Chinese" models.
lukax··on GitHub is having issues now
It looks like migration to Azure is not going very well

https://news.ycombinator.com/item?id=45517173

lukax··on GitHub is having issues now
They are migrating from their own datacenters to Azure
lukax··on It's OK to compare floating-points for equality
See the implementation of Python's math.isclose

https://github.com/python/cpython/blob/d61fcf834d197f0113a6a...

lukax··on It's OK to compare floating-points for equality
You generally want both relative and absolute tolerances. Relative handles scale, absolute handles values near zero (raw EPSILON isn’t a universal threshold per IEEE 754).

The usual pattern is abs(a - b) <= max(rel_tol * max(abs(a), abs(b)), abs_tol) to avoid both large-value and near-zero pitfalls.

lukax··on Someone Bought 30 WordPress Plugins and Planted a Backdoor in All of Them
Do you really need to roll your own NIO HTTP server? You could just use Jetty with virtual threads (still uses NIO under the hood though) and enjoy the synchronous code style (same as Go)
lukax··on Someone bought 30 WordPress plugins and planted a backdoor in all of them
Rust wasm ecosystem also needs a lot of crates to do anything useful, a lot of them unmaintained.
lukax··on Vite Vulnerable to Arbitrary File Read via Vite Dev Server WebSocket
Combine that with CVE-2025-24010 and any website was able to read any file on developers' computers.

https://github.com/advisories/GHSA-vg6x-rcgg-rjx6

lukax··on WSL Manager
Looks nice but still a bit sad that Flutter is used instead of something native given that they don't need the app to be cross-platform.

Well, even Microsoft uses React Native for a lot of Windows-only apps.

lukax··on Show HN: I built a sub-500ms latency voice agent from scratch
Sorry, I commented too soon. Did you also try Soniox? Why did you decide to use Deepgram's Flux (English only)?
Page 1 of 4Next →