HNHacker News
TopNewBestAskShowJobs

lukax

916 karma · joined January 11, 2012

submissionscomments
lukax··on Show HN: I built a sub-500ms latency voice agent from scratch
Or you could use Soniox Real-time (supports 60 languages) which natively supports endpoint detection - the model is trained to figure out when a user's turn ended. This always works better than VAD.

https://soniox.com/docs/stt/rt/endpoint-detection

Soniox also wins the independent benchmarks done by Daily, the company behind Pipecat.

https://www.daily.co/blog/benchmarking-stt-for-voice-agents/

You can try a demo on the home page:

https://soniox.com/

Disclaimer: I used to work for Soniox

Edit: I commented too soon. I only saw VAD and immediately thought of Soniox which was the first service to implement real time endpoint detection last year.

lukax··on Web Components: The Framework-Free Renaissance
Wow, XSS just waiting to happen.

  <h3>${this.getAttribute('title')}</h3>
lukax··on Audio is the one area small labs are winning
Never any mention of Soniox and they are on the Pareto frontier[1]

https://www.daily.co/blog/benchmarking-stt-for-voice-agents/

lukax··on Soniox: Real-time transcription in 60 languages
Also see how it compares to other providers:

https://soniox.com/compare

lukax··on Nano-vLLM: How a vLLM-style inference engine works
Not really in the PagedAttention kernels. Paged attention was integrated into FlashAttention so that FlashAttention kernels can be used both for prefill and decoding with paged KV. The only paged attention specific kernels are for copying KV blocks (device to device, device to host and host to device). At least for FA2 and FA3, vLLM maintained a fork of FA with paged attention patches.
lukax··on Another user's pCloud setup is visible in my pcloud drive
pCloud has been leaking files between users

https://www.reddit.com/r/pcloud/comments/1qhpr4k/vault_got_a...

https://www.reddit.com/r/pcloud/comments/1qhibbe/pcloud_sudd...

https://www.reddit.com/r/pcloud/comments/1qhxuco/followup_to...

https://www.reddit.com/r/pcloud/comments/1qhibbe/comment/o18...

lukax··on Vibecoding #2
Maybe AWS ParallelCluster which is a managed SLURM on AWS.
lukax··on CVEs affecting the Svelte ecosystem
It's not that simple to safely parse HTTP request form. Just look at Go security releases related to form parsing (a new fix released just today).

https://groups.google.com/g/golang-announce/search?q=form

5 fixes in 2 years related to HTTP form (url-encoded and multipart).

- Go 1.20.1 / 1.19.6: Multipart form parsing could consume excessive memory and disk (unbounded memory accounting and unlimited temp files)

- Go 1.20.3 / 1.19.8: Multipart form parsing could cause CPU and memory DoS due to undercounted memory usage and excessive allocations

- Go 1.20.3 / 1.19.8: HTTP and MIME header parsing could allocate far more memory than required from small inputs

- Go 1.22.1 / 1.21.8: Request.ParseMultipartForm did not properly limit memory usage when reading very long form lines, enabling memory exhaustion.

- Go 1.25.6 / 1.24.12: Request.ParseForm (URL-encoded forms) could allocate excessive memory when given very large numbers of key-value pairs.

Probably every HTTP server implementation in every language has similar vulnerabilities. And these are logic errors, not even memory safety bugs.

lukax··on Locating a Photo of a Vehicle in 30 Seconds with GeoSpy
You can buy a totaled car for cheap and use its VIN.
lukax··on VSCode rebrands as "The open source AI code editor"
I guess there's a lot of pressure from Cursor and Google's Antigravity. Also with Zed you can bring your own API key which VS Code didn't support for a long time.
lukax··on I program on the subway
17 years ago I went to a summer vacation with my family (still a teenager). That meant 10 days without any internet connectivity. I just got my first laptop and I was allowed to take it with me. I was reverse engineering MSN Messenger's user to user and profile picture exchange protocol from TCP dumps. MSN Messenger did not use any encryption. Before I went to the vacation I recorded a bunch of sessions with Wireshark (maybe it was still Ethereal back then). Then for 10 days I was just trying to figure out from the dumps how the binary protocol worked and was writing the code without any way to test it. When I came back I just had to fix some minor bugs and it worked. Fun times.
lukax··on Build Android apps using Rust and Iced
I've done business logic sharing where the engine was written in Rust, WASM for web with React for UI, uniffi-rs for Android and iOS with Kotlin Compose for Android and SwiftUI for iOS, Tauri for desktop.

There were no good examples for how to do this but once it was set up it worked extremely well.

It uses tokio for Android/iOS/desktop and even embeds a web server for fake API for end to end testing (even on mobile)

https://github.com/koofr/vault

lukax··on A2UI: An Open Spec for Agent-Generated User Interfaces (Google)
They made this transport agnostic so it's compatible with A2A and AG-UI.
lukax··on LLM Year in Review
Google is doing that with A2UI. LLM will be able to decide how to present info to the user.
lukax··on GPT-5.2-Codex
Will it also remove the whole D:\?
lukax··on Ask HN: How can I get better at using AI for programming?
Spokenly on macOS with Soniox model.
lukax··on Structured outputs on the Claude Developer Platform
In OpenAI and a lot of open source inference engines this is done using llguidance.

https://github.com/guidance-ai/llguidance

Llguidance implements constrained decoding. It means that for each output token sequence you know which fixed set of tokens are allowed for decoding the next token. You prepare token masks so that in the decoding step you limit which tokens can be sampled.

So if you expect a JSON object the first token can only be whitespace or token '{'. This can be more complex because the tokenizers usually allow byte pair encoding which means they can represent any UTF-8 sequence. So if your current tokens are '{"enabled": ' and your output JSON schema requires 'enabled' field to be a boolean, the allowed tokens mask can only contain whitespace tokens, tokens 'true', 'false', 't' UTF-8 BPE token or 'f' UTF-8 BPE token ('true' and 'false' are usually a single token because they are so common)

JSON schema must first be converted into a grammar then into token masks. This takes some time to be computed and takes quite a lot of space (you need to precompute token masks) so this is usually cached for performance.

lukax··on When Tesla's FSD works well, it gets credit. When it doesn't, you get blamed
Kind of similar to Agentic coding. The code works? Yay, well done, AI. It't doesn't? You didn't prompt it correctly.
lukax··on Typst 0.14
And the hayro library is standalone and can easily be used outside of Typst. It only uses CPU and is pure Rust so it can also be used with WebAsembly. Link to demo below.

https://github.com/LaurenzV/hayro

https://laurenzv.github.io/hayro/

lukax··on ChatGPT Launches 'Company Knowledge'
> It’s powered by a version of GPT‑5 that’s trained to look across multiple sources to give more comprehensive and accurate answers.

So another GPT-5 fine-tune. Codex also uses a custom GPT-5 fine-tune.

Does fine-tuning make sense now? Or do you have to be OpenAI to fine-tune the models with a mix of existing data and new behaviours?

lukax··on OpenAI ChatKit
And CopilotKit needs a JS backend for a proxy that holds the state for the chat which is quite a pain if you want to scalably self-host. Maybe ChatKit will be less lock-in than CopilotKit.
lukax··on OpenAI ChatKit
This looks very similar to CopilotKit and AG-UI. With CopilotKit you can have some tools implemented client-side and others server-side using AG-UI.

https://www.copilotkit.ai/

With ChatKit you can proxy everything through your ChatKit API "proxy" and implement the "agentic" stuff in the proxy.

Even the React APIs look very similar.

Looks like a vendor lock-in and a way to stop the AG-UI ecosystem

lukax··on Tile Language: DSL for High-Performance GPU/CPU/Accelerators Kernels
Example kernels for DeepSeek-V3.2-Exp

https://github.com/tile-ai/tilelang/tree/main/examples/deeps...

lukax··on The AI coding trap
It's really accurate and supports 60+ languages
lukax··on The AI coding trap
Have you tried Soniox? It's really not expensive ($0.12/h, $200 free credits when you sign up) and really accurate.

https://soniox.com/

You can use it with Spokenly (free app, bring your own Soniox API key) on macOS and iOS (virtual voice keyboard)

https://spokenly.app/

Disclaimer: I've worked for Soniox

lukax··on I built Foyer: a Rust hybrid cache that slashes S3 latency
Maybe you have an ad-blocker that just hides the popup but does not restore scrolling (scrolling is usually prevented when popups are visible)
lukax··on After Babel Fish: The promise of cheap translations at the speed of the Web
Soniox offers real-time speech-to-text with real-time translation between 60+ languages (mostly to/from English with some additional pairs between more popular languages). It operates with minimal amount of context and produces the translated text as soon as possible.

https://soniox.com/

Disclaimer: I used to work for Soniox

lukax··on Gluon: a GPU programming language based on the same compiler stack as Triton
Is this Triton's reply to NVIDIA's tilus[1]. Tilus is suposed to be lower level (e.g. you have control over registers). NVIDIA really does not want the CUDA ecosystem to move to Triton as Triton also supports AMD and other accelerators. So with Gluon you get access to lower level features and you can stay within Triton ecosystem.

[1] https://github.com/NVIDIA/tilus

lukax··on If my kids excel, will they move away?
word
lukax··on VibeVoice: A Frontier Open-Source Text-to-Speech Model
Have you tried Soniox for speech recognition? It supports Croatian. Or are you just looking for self-hosted open-source models? Soniox is very cheap ($0.1/h for async, $0.12/h for real-time) and you get $200 free credits on signup.

https://soniox.com/

Disclaimer: I used to work for Soniox

← PreviousPage 2 of 4Next →