HNHacker News
TopNewBestAskShowJobs

a24venka

66 karma · joined March 30, 2023

submissionscomments
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Really appreciate the detailed feedback.

1) Agreed on the show-tell-show framing: Chat is a natural starting point for people frustrated with linear interfaces, but you're right that the a-ha is the orchestration and outputs. We'll keep this in mind as we build out our gallery and demos going forward.

2) Right now our primary users are knowledge workers and researchers who need to do complex, multi-step work. The benchmarks help establish credibility, but we're building out more use case demos and a gallery to make it tangible for a broader audience. On the human-in-the-loop point, the agents do pause and come back to the user, and there is scope for iterations as well, but we haven't highlighted this well enough. We'll do a better job of showing that going forward.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Thanks for the feedback! On the chat interface point, we actually think chat is still a great way to get into the product, but we leverage the canvas behind the scenes to let the agents do better work and give users the ability to audit and visualize what's happening. Good note on the demo topic though, we have some broader, non-AI demos coming soon on our website and newsletter.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
This is really great to hear, thank you! Have fun with the prototype, let us know how it goes.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Spot on. The persistence layer is a huge part of what makes the canvas work.

For failures, we handle it at multiple levels: first, standard retries and fallbacks to alternate models/providers. If that fails, the agents look for alternate approaches to accomplish the same task (e.g. falling back to web search instead of browser use).

For completeness, you can also manually re-run or edit individual blocks if they fail (though the agents may or may not consider this depending on where they are in their flow).

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Great question. The core of Spine is coordinating multiple specialized agents across multiple models, using the canvas to store and pass context selectively so each agent works with exactly what it needs.

On the eval side, we ran Spine Swarm against GAIA Level 3 and Google DeepMind's DeepSearchQA and hit #1 on both.Full writeup: https://blog.getspine.ai/spine-swarm-hits-1-on-gaia-level-3-...

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
This is great to hear, thank you! Would love to hear your thoughts once you see your final report and explore some of those other opportunities.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Great framing. You're right that context fragility is a big part of it. The canvas helps because each block maintains its own context explicitly, and connected blocks pass context between blocks without polluting the agents' context windows.

On conflict resolution, the synthesizer block can see all upstream outputs, so it has full visibility into any divergence. It does surface contradictions to the user, though this is something we're constantly improving.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Agreed. The term has been overloaded lately. We also refer to it as a visual workspace which perhaps captures it a bit better.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Agreed. We will make sure this comes through in our website.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Really appreciate the detailed feedback and questions! And yes, we'll take the website criticism as a compliment :)

Good callout on the canvas navigation, we'll look into middle mouse button support.

To answer your questions: 1) GitHub integration is on our roadmap. Right now you can export outputs manually but we want to make this seamless. 2) All your canvases are saved and you can search them by name in your dashboard. We're also working on a dedicated section for deliverables across canvases. 3) Yes to both! You can manually add or edit blocks, or kick off new agent runs that build on existing work. 4) You can currently only share public links of your canvas to others (but you can make it private at any point). We are testing out a teams feature which allows you to share canvases with members on your team securely. Beyond that, we are working on adding roles and email-based sharing controls which is in our roadmap. 5) Claude Code in a block is a really interesting idea. We don't support that today but we're thinking about computer use and coding workflows. 6)BYOK (bring your own keys) is something we've heard interest in and are considering. Self-hosting isn't available right now, though we do support private deployments for enterprise customers if that's ever relevant. 7) Love the 'preferred web browser' framing. Right now you can search canvases but searchable artifacts across canvases is definitely where we want to head.

Thanks for giving it a real spin, this kind of feedback is incredibly valuable.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Thanks! Great question. We see canvases as living workspaces, you can revisit, iterate on, and build on them over time.

But the deliverables (docs, slides, code) are first-class outputs you can export and use independently. So it works both ways depending on the workflow.

Kavla looks cool, canvas-based SQL is a great use case for this kind of thinking!

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
This is good feedback and definitely something we are improving.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
The daily refresh isn't a cap on usage, it's additional credits you get each day (resets to 1,500 nightly regardless of use).

You can use your full 30k balance in a single run if needed. The daily refresh just tops you back up over time so you're not waiting for a monthly reset.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Definite miss on our part, we're working on making the product experience more visible upfront on our landing page.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Good call and noted. We're working on making the product experience more visible upfront.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
This is a fair point, we are exploring progressive disclosure on the canvas to better utilize the space and make the key artifacts more readily visible. We do have other panels (the chat, task and deliverable) that have alternate views of what the agent did and the key deliverables.

Beyond human auditability, the canvas helps the agents do a better job by generating in parallel, exploring branches and passing context to each other in a structured way.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Credits are consumed by the blocks that get generated, not by the agents themselves. Some blocks are cheaper than others. A simple prompt or image block is a single model call, while browser use or deliverable blocks like documents and spreadsheets run models in a loop and cost more. Blocks also cost more when they have more blocks connected to them (more input tokens).

In the demo video I shared, the task cost about ~7,000 credits since it ran around 10 BrowserUse blocks and produced multiple deliverables.

If you want to fix a specific block (or set of blocks), you can select them and the chat will scope itself to primarily work on those. In that case fewer blocks run, so it's cheaper.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Fair point, we should be more upfront about the sign-up step. Given that tasks are long-running and token-intensive, we do need an auth barrier to protect against abuse, but we can definitely do a better job signaling that before you hit the canvas.
a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
We've only partially explored this so far, but it's a great suggestion.

The canvas architecture naturally supports this kind of loop since agents can already read and build on each other's outputs — so the plumbing is there, it's more about building the right orchestration on top. Definitely something we're exploring.

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
That is definitely a valid way of using Spine as well. You can just work in the chat and consume the deliverables similar to how you would in other tools.

The canvas helps when you want to trace back why an output wasn't what you expected, or if you're curious to dig deeper.

Even beyond auditability, the canvas also helps agents do better work: they can generate in parallel, explore branches, and pass context to each other in a structured way (especially useful for longer-running tasks).

a24venka··on Launch HN: Spine Swarm (YC S23) – AI agents that collaborate on a visual canvas
Thanks for the feedback! Definitely agree that we could do more with the marketing site. We're working on a gallery page to showcase some demos.
a24venka··on Replacing file systems with a canvas produced SOTA agent performance
Hey HN — Akshay & Ashwin here, co-founders of Spine AI (YC S23).

We've been rethinking how AI agents work together. Instead of a single model in a chat loop or agents reading/writing to a file system, we built a visual canvas where multiple agents collaborate across connected blocks — and it turns out this architecture significantly outperforms both single and multi-agent systems on hard tasks.

The approach has three parts:

1. Canvas-based workspace — Agents operate on an infinite canvas of intelligent blocks (web browsing, prompts, tables, memos) that connect and pass context to each other. Instead of a flat file system, agents get a structured, non-linear environment that mirrors how complex problems actually decompose.

2. Tiered multi-agent orchestration — An orchestrating agent decomposes tasks, delegates to specialized persona agents (researcher, analyst, reviewer), and manages dependencies. Agents validate each other's work before passing it downstream, catching errors before they compound across long chains.

3. Dynamic multi-model ensembling — Rather than one model for everything, we select from 300+ models per subtask. When confidence is low, we pull in additional models and treat disagreement as a signal for deeper scrutiny — like classical ML ensembling, but at the agent level.

The results: 61.5% on GAIA Level 3 (vs Manus 57.7%, OpenAI Deep Research 47.6%) and 87.6% on DeepSearchQA (vs Perplexity 79.5%, Gemini Deep Research 66.1%). Same frontier models available to everyone — the difference is architecture.

Because everything runs on the canvas, we could audit our agents' work step by step. That's how we caught what appear to be mislabeled questions in the GAIA dataset itself — we link to sample canvases in the post so you can see the reasoning traces.

Spine Swarms is open to try at www.getspine.ai. Happy to go deep on any of the architecture.

a24venka··on Your job is to deliver code you have proven to work
There is a heavy emphasis on testing the code as the way to provide guarantees that it works. While this is a helpful tool, I often find that the best engineers are ones who take a more first principles approach to the code and can reason about why the solution is comprehensive (covers all edge cases) and clean (easy for humans and LLMs to build on).

It often takes discipline to think and completely map out solutions before you build. This is where experience and knowing common patterns can also help.

When you have the experience of having manually written or read a lot of code it helps at the very least quickly understand what the LLMs are writing and reason about it later even if not at the beginning.