HNHacker News
TopNewBestAskShowJobs

michalwarda

323 karma · joined February 26, 2021

submissionscomments
michalwarda··on Everyone is comparing GitHub stars. But what about issues?
I've noticed how the opencode team has around 100 issues opening every day and I wondered how it is compared to other big repos in history. So I've created https://github-history.com to look at it.

Hope you'll like it!

michalwarda··on We built 270 realistic chats so agents don't freeze while tools run
AI shouldn’t block the conversation just because a tool is busy. To evaluate that behavior properly, we needed good data—so we made it.

AsyncTool is a Hugging Face dataset of 270 high‑quality, multi‑turn (and I mean up to 60 turns) conversations where the assistant keeps talking while tools work in the background. Each case is different, grounded in real JSON‑Schema tool definitions, and the tool calls/results are consistent and make sense with no fabricated states or magical shortcuts.

What’s inside - 18 scenario templates × 15 renders = 270 conversations. - Conversations run 10–30 “in‑world” minutes with filler chat, retries, status checks, and out‑of‑order returns. - Every row includes messages, tools, and meta so you can replay transcripts, inspect schemas, and trace provenance. - Protocol features: <tool_ack /> placeholders, -FINAL handoffs, mixed sync/async chains, transient failures, and fatal‑error surfacing. - License: Apache‑2.0.

We’re exploring how agents can ack now, answer later - waiting for the right signal (last relevant tool result vs. last user question) while staying natural and helpful. This dataset gives you supervised signals to: - finetune assistants that acknowledge async work without hallucinating tool states, - build guardrails/regression tests for routers juggling retries and reordered responses, - evaluate “answered at the right time” behavior.

We’re also publishing the generator so you can reproduce or extend everything locally. If you’re building tool‑using agents - or just tired of UIs that freeze—this should help you train, test, and iterate faster.

Built with Torque → https://usetorque.dev/

michalwarda··on React for Datasets
We’ve spent a while trying to generate a very specific, complex, multi‑turn conversation dataset. Most tools pushed us toward glue scripts and one‑off pipelines that were hard to review or reuse. Torque is our attempt to make this boring and predictable: a small, MIT‑licensed framework that treats dataset generation like building a UI. You define small, declarative “components” and compose them into pipelines. The goal is clear code and repeatable runs, not another heavy DSL. It’s early and open source. We’d love feedback on the API design, examples you’d like to see, and rough edges we should fix first. Docs and code: https://github.com/qforge-dev/torque
michalwarda··on Show HN: Coyote – Wildly Real-Time AI
It can also set reminders and do stuff when you're not looking and talks to you by itself rather than always being triggered by the user message.
michalwarda··on Show HN: Coyote – Wildly Real-Time AI
Fully understand the WhatsApp part. Do you use any other communicators? Discord, Signal or something else? We are looking for more "interfaces".

When it comes to async that's exactly what we are trying to "solve". Right now models are built in a way that they expect tool results right after tool calls.

We are building datasets and a model with asynchronous interface as it's core. https://huggingface.co/qforge/Qwen3-14B-AT you can read more about it here.

For tooling attached to it right now we are using Pipedream integrations and are planning to move to an open source public solution configurable by users so u can set whatever you want.

So imagine you have a single chat interface that steers a fleet of other agents in the background for code. But not by handing off the memory but navigating the tasks in an async way.

michalwarda··on Show HN: Coyote – Wildly Real-Time AI
Right now it can: - handle real tasks in the background — emails, calendar stuff, research, finding info, organizing data - chat naturally without feeling like you're talking to a bot - remember context and keep conversations flowing - work with integrations (gmail, calendar, docs, maps, etc.) so it can actually do stuff, not just talk about it - multi-task — you can ask it multiple things and it can handle them in parallel also if u mispronounce anything it can update the existing stuff that is happening.
michalwarda··on Show HN: Coyote – Wildly Real-Time AI
It's like you're talking to a real person. No stop buttons, no waiting. Interrupt, add details, or change direction anytime. Just like a natural conversation.

Context here means any additional information.

michalwarda··on I made Poke.com email me its system prompt lol
I was wondering how bitchy the prompt was. Turns out other pretty weird stuff was there.
michalwarda··on Show HN: MCP powered cross platform Superwhisper
It supports Local and Cloud models for Transcription and we're super close in releasing LLM integration with multiple providers including local ones.
michalwarda··on MCP powered cross platform Siri
We've released a new version of qSpeak v0.1.47.

Our vision on the future of computer use through voice assistance. Available everywhere in your system to do whatever you'd like including advanced agent based tool usage.

We're having a completely free beta now so hope you'll like it and test it out!

michalwarda··on Show HN: QSpeak – An alternative for WisprFlow supporting local LLMs and Linux
Yea, it's not been easy. We're working on supporting more and more distros.
michalwarda··on Show HN: QSpeak – An alternative for WisprFlow supporting local LLMs and Linux
Hey, together with my colleagues, we've created qSpeak.app

qSpeak is an alternative to tools like SuperWhisper or WisprFlow but works on all platforms including Linux.

Also we're working on integrating LLMs more deeply into it to include more sophisticated interactions like multi step conversations (essentially assistants) and in the near future MCP integration.

The app is currently completely free so please try it out!

michalwarda··on I compared my daughter against SOTA models on math puzzles
Very cool post! I wonder how much will it affect the psychology of next generations.
michalwarda··on [dead]
Based on current builds data you can see that even in early access the build distribution looks better than in PoE 1.
michalwarda··on Unleashing the Power of Knowledge Graphs – BuildEL Release v0.3
Hey everyone. We've just released version 0.3 of BuildEL, an Elixir based fully open source AI orchestrator used also for RAG retrieval. Hope you'll like it :). You can read more about it here: https://buildel.ai/blog/buildel-0_3 This release brings embeddings based relation graphs which is a huge eye opener when it comes to RAG development using vector databases. It was a big challenge to do it but thanks to language abstractions it turned out to be a very small amount of code :) It uses advanced math algorithms to narrow multidimensional embedding vectors to 2d. If you have any question you can ask, and I'll do my best to respond :)
michalwarda··on Solving the out-of-context chunk problem for RAG
I guess it's because people are not using tools enough yet. In my tests giving LLM access to tools for retrieval works much better then trying to guess what the RAG would need to answer. ie. LLM decides if it has all of the necessary information to answer the question. If not, let it search for it. If it still fails than let it search more :D
michalwarda··on Solving the out-of-context chunk problem for RAG
This works until relevant information is colocated. Sometimes though, for example in financial documents, important parts reference each other through keywords etc. That's why you can always try and retrieve not only positionally related chunks but also semantically related ones.

Go for chunk n, n - m, n + p and n' where n' are closest chunks to n semantically.

Moreover you can give this traversal possibility to your LLM to use itself as a tool or w/e whenever it is missing crucial information to answer the question. Thanks to that you don't always retrieve thousands of tokens even when not needed.

michalwarda··on Buildel 0.2 release – no-code AI orchestrator
Hey, me and my team have been working further on our Open Source tool called Buildel. It's an AI orchestrator with built in functionalities to quickly create your own bots, automations and advanced AI workflows. All of that without much vendor lockin because of standardized APIs and fully documented and accessible codebase. Would love for everyone to check it out at https://buildel.ai/blog/buildel-0_2

In this release we've added a new design, new workflow editor, new interfaces, tools and much more!

michalwarda··on [dead]
Together with my wife we've created a game in 24h that was solved by 3000 people within 7 days of it's creation. You can read how we host it it on a Raspberry Pi from our home.
michalwarda··on Do you know programming languages? progle()
Pick a language, submit, and based on the attributes that show up you pick another one. Then knowing the previous green ones you find the one that you are looking for. Similar to https://www.nytimes.com/games/wordle/index.html but instead of words you send languages.
michalwarda··on Do you know programming languages? progle()
Me and my wife just created a new wordle like game where you try and find a programming language based on hints. Hope you like it! If you have any suggestions etc. be sure to send them :)
michalwarda··on Bun vs Node Benchmark - no one cares about speed as much as your CI does
Whole bun test runner is implemented in zig + the super quick startup of JSC
michalwarda··on Bun vs Node Benchmark - no one cares about speed as much as your CI does
Fixed. Thanks for info :).

When it comes to demo and bun then i suggest focusing on deno until bun hits stable release. It's still rough around the edges.

michalwarda··on Bun vs Node Benchmark - no one cares about speed as much as your CI does
It definitely shows how much potential improvement there is for test suites in TS world.
michalwarda··on Bun vs Node Benchmark - no one cares about speed as much as your CI does
Bun is fast. But honestly I don't care about the fact that it can handle more requests. What matters for me is how much it affects my dev experience. See how Bun performs in a non-typical benchmark.
michalwarda··on How to contribute to a project you have no idea about
A short post about how I've contributed to Bun, a huge project that is written in a language I don't understand, and does stuff that I've never worked on. And what work framework I use to work on projects like that.
michalwarda··on Self hosting in 2023
Sorry to annoy you ^^. I'm not a native speaker and I guess it's my wording in the post. What I meant is "setting it up for every new project I would be working on might be cumbersome".

Honestly I thought only friends and family would read it and I didn't spend enough time polishing the nuances in the writing. I didn't expect it to blow up so much :D.

But going back to the point I fully understand your frustration about the overcomplicated approach. Though for my defense I'm planning on hosting many more apps using this setup. In fact I already am but didn't want to overcomplicate the post.

I left it for part 2. Subscribe RSS for followup xD. I feel like a YouTuber after saying that.

michalwarda··on Self hosting in 2023
Yeah. Maybe I phrased it badly but I meant like more complicated apps than a blog :D
michalwarda··on Self hosting in 2023
Possible on both.
michalwarda··on Self hosting in 2023
I had a huge paragraph in the article about decentralization and how important it's in my mind for the future of internet but I've scrapped it because it felt like I was a blockchain guy even without mentioning it. So I'll make it simpler.

Honestly I just feel awesome seeing those blinking lights in my room and thinking it's sending packets to other people.

Page 1 of 2Next →