HNHacker News
TopNewBestAskShowJobs

songrenchu

56 karma · joined October 8, 2014

submissionscomments
songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
We added 4 things to remediate lowest common denominator. The first is task interface, it is portable cross harnesses. Second is capability discovery. Each implemenation needs to specify capabilities it doesn't support. Third is config that captures harness differences. Last is extensions like metadata and extra fields.
songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
Yes, these two are good use cases. With no harness lock-in, a fallback and context handoff can solve LLM vendor service unreliable. Loop engineering can also be done for achieving long horizon goals a single harness loop struggles to close
songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
Thank you for the discussion! This really nudges us to think how to make the broader white collar agent use case message more clear
songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
You are right, for coding scenario, I also stick with one (CC in my case, really got disappointed at codex during gpt-5.4 time and never came back since then)

We need HarnessRouter when we need to package the harness agent as part of the product backend to serve the end users. In that scenario, the harness needs specific instructions, MCP tools, skills pre-configured, so it can reliably receive requests from upstream product components and deliver result to downstream product components.

We put 4 demo agent products for white collar working scenarios in our starter kit: PPT agent, Spreadsheet agent, Bi Dashboard agent, Video editing agent. Each of them is backed by a different harness setup. Video editing is most sophisticated so it's CC + Opus 5. The other 3 are more simpler use cases so default setup in the kit is set to Hermes + DeepSeek V4 Pro.

Take the PPT agent use case, for sure you can hook the same tools and skills to local Claude Code or Codex, but it only works for yourself using it locally. If you are building a AI PPT product (like Gamma), you need to host the harness setup somewhere in the cloud together with other product code. That's when you can use HarnessRouter as the PPT generation/manipulation component of the product, with the chosen harness baked in. For sure you can build the same harness wrapper plumbing as we did in HarnessRouter to make the same stack work, but using HarnessRouter the development time is shorten as we have already get the nitty gritty engineering details covered

songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
The router sits between application layer and the harness layer. HarnessRouter is the spec translation layer that translates the unified interface into each harness's own api format. Each harness is treating somehow like a blackbox, and they talk to the models as is. For routing across harnesses, think about it as an aggregator, like OpenRouter. The application layer have multiple use cases and each function could backed by a different harness. We do have smart routing feature on our roadmap to support use cases of harness fallback, cost optimization, etc
songrenchu··on Show HN: HarnessRouter: Unified interface for agent harnesses
Think of it as OpenRouter, but for different agent harnesses, not models. Instead of sticking to any one harness, you can route to and use any of them through a unified API
songrenchu··on A ChatGPT-like assistant but private for developers
This is interesting, with e2e private chat, there are much more possibilities being unlocked without worrying about censorship
songrenchu··on Ask HN: What is the current status of RAG-as-a-service tools out there?
Take a look at https://epsilla.com/ ?
songrenchu··on Show HN: Create Email Templates in Seconds Not Hours (v0.dev for Emails)
Great, does it support creating monthly newsletter like emails?
songrenchu··on Show HN: Continuous-eval – Granular evaluation of GenAI pipelines
Nice! Componentization of gen AI workflow is a new trend, and we do need an evaluation framework like this
songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
Thank you for pointing it out. We just removed it from README
songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
Thank you for the insightful topic! By reading the question itself drive me think a lot.

For the database perspective, instead of dividing the table schema into 3 parts: id, metadata, embedding, we designed in a way closer to SQL, treat vector as another data type, and let user to define any number of fields in a table. ID is just an annotation of a field (composite key might be overkilling for now). There will be another debate on whether schemaful or schemaless is the right approach, we can leave it here for now

With this foundation, we already covers 1 and 2. And in our roadmap we also plan to cover 3, with multi-modal data type support. We think the real big advantage of embedding is on unstructured data (documents, images, video, audio, etc), and storing the embedding of multi-modal data and connect them through semantic relevance will open up big opportunities. And this fits with the table and fields-based design for introducing cross table embedding index on connecting different shape data.

And from the multi-modal data perspective comes the problem where do we store those data? One way is we provide a generic binary data type that let users put anything. Another way which most enterprise will do is integrate us with a larger data warehouse/data lake system. And this opens up the requirement for us on supporting data streaming in/out with kafka connector, spark connector, etc.

And totally agree that SQLite works so well in huge amount of scenarios, now there is DuckDB. We also see some other players like LanceDB taking this approach to be Vector DB space's SQLite. We are also pretty close to announce our Python in-process package support, so docker / a separate server is not a must have anymore.

For inference, this is a broader direction for us for now. We are open to explore this space and see if the serverless architecture on cloud can provide extra efficiency benefit to the market

songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
For now we didn't put a limit on the dimension of the vectors, so the machine can fit as much as #vector * #dimension * sizeof(float) into memory. For now we just support dense vector, and in the future we will work on sparse vector support for much higher dimension. I think you are referring to the "Curse of dimensionality" problem. Here is my thoughts: in a graph-based index such as SpeedANN, or HNSW, each vector is treated as a node in the graph, and the index is a nearest neighbor graph. Different from spatial partition-based indices, the topology quality of the nearest neighbor graph is independent from the dimensionality of the vectors. Our benchmark is on 960 dimension vector, but we will do more experiments in sparse vectors in the future
songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
You are right. We designed our storage in a segment-based way, with configurable segment size, so it can horizontally scale in the future cross multiple workers in one machine, and cross multiple machine cluster. And the search will become a two-stage search: find top K in each segment, then a global merger (can also be horizontally scaled) to merge the results from all segments
songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
Thank you for sharing! DiskANN was published in 2019 and SpeedANN in 2022. DiskANN is specialized in disk based ANNS solution, and it's focus on the scenario where the vectors don't fit into memory. SpeedANN is in-memory solution and specialize in low latency query, which is the scenario we want to tackle for now. We can further extend our engine to support DiskANN and other index algorithms based on our customer's requirements Thanks for pointing out the benchmark figure, we just fixed it to have consistent colors
songrenchu··on Show HN: Epsilla – Open-source vector database with low query latency
You are right, there are numerous vector databases in the market. Most of them (including us) are still pretty early and a lot of enterprise readiness features to build. Including role based / privilege based access control, authN/Z integration, data versioning and backup / restore, fault tolerance, data streaming in/out, etc. We have first hand experience on the enterprise level product development and sales in our previous job at a series D graph database startup, and we will apply our learnings there to make Epsilla enterprise ready in the next few months