HNHacker News
TopNewBestAskShowJobs

bleonard

425 karma · joined December 2, 2010

CEO and cofounder @ Grouparoo Previously: CTO and technical cofounder @ TaskRabbit

[ my public key: https://keybase.io/bleonard; my proof: https://keybase.io/bleonard/sigs/A_ilc1h8TmtRx3gMAUkpDEdRZIcQ9W-aNqb3yTKQQ3Q ]

submissionscomments
bleonard··on Claude Opus 5.5
The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%

So longer threads get cheaper and one-shots stay the same price.

bleonard··on How GLM built its own inference infrastructure
We have been running a lot of agentic benchmarks with the various loops and tool calls on longer threads - we routinely see 90%+

Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%

[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon

bleonard··on Pion, an agent designed to run any company autonomously
I've been spending all my time recently thinking about LLM costs. From that, I am currently of the mind that the most interesting question is if they can get it to work in the first place. After that, like with "performance" in the past, I believe there are ways to optimize things. We see the code harnesses doing this out of necessity, for example.

Some of the tools we use: https://www.induction.ai/docs/context-management

bleonard··on Show HN: 30k IKEA items in flat text
A blast from the past. When Taskrabbit was acquired by IKEA, I built several tools that went through the whole catalog via various crawling approaches. One tool was to estimate how long it would be to put each item together for an initial training set.
bleonard··on Kafka is Fast – I'll use Postgres
I am excited about the Rails defaults where background and cache and sockets are all database driven. For normal-sized projects that still need those things, it's a huge win in simplicity.
bleonard··on Ask HN: Who is hiring? (March 2025)
Ario | Onsite in Palo Alto, CA | Full-time | heyario.com

I recently started at Ario where we are building an app for parents that uses AI to make things easier. We are targeting saving them one hour every day. Previously, I co-founded Taskrabbit. That was somewhat similar but now it's time to get LLMs in on the action! It's in the app store, but still early in the game, so we are looking for engineers to build it out. We are using Python and React Native.

To learn more: brian at heyario.com

bleonard··on Airbyte 1.0 – Marketplace, AI Assist, Gen AI Support and Enterprise GA
We have to work on the scaling. There is lots of web scraping and LLM calls that we need to make sure works under load. I'm sure there will be quality improvements as well.
bleonard··on Airbyte 1.0 – Marketplace, AI Assist, Gen AI Support and Enterprise GA
AMA: I was on the team that built the AI Assist functionality. Happy to chat about it. I became obsessed with documentation sites for a while there.
bleonard··on Pivotal Tracker will shut down
Just came to say, I still think this is the best balance between the many factors of running a dev team. I keep trying to recreate it in every tool I use.
bleonard··on ELTP: Extending ELT for Modern AI and Analytics
Airbyte acquired our Reverse ETL company, Grouparoo, 1.5 years ago. There is so much to solve making just the Extract and Load work well and so much value that comes from that, we have been busy there. I'm excited to circle back to publishing next year.

I like how the article notes that the stuff we were talking about with Reverse ETL (mostly activating your data in SaaS systems like Salesforce, Zendesk, etc) is one important part of Publishing. But we are also seeing traditional use cases like file uploads and new fancy stuff like vector databases.

bleonard··on IKEA’s knowledge graph and why it has three layers
Maybe some day this will be available in JSON.

I built a system for TaskRabbit that scraped all the IKEA products from a variety of sources and ran algorithms to determine their category and predict how long they would take to be assembled. Then there was a Mechanical Turk sort of system for human input. When combined with real-world feedback from the Taskers, it was pretty good.

For better or worse, I've personally been through the entire catalog multiple times.

bleonard··on Airbyte acquires Grouparoo to accelerate Data Movement
From the Grouparoo team, we are so excited to be working with Airbyte to be able to move data anywhere you want it to go.
bleonard··on Launch HN: Castled Data (YC W22) – Open-Source Reverse ETL
Congrats to the Castled team on the launch.

At Grouparoo, this is a primary use case. We have a UI that engineers use locally. This helps gets things right. It outputs a JSON configuration that is checked in. When that is deployed, it does all the syncing.

bleonard··on Launch HN: Hightouch (YC S19) – Sync data from data warehouses to SaaS tools
It's hard to leave the comparisons dangling, for sure. But I'll defer for now. Congrats on the launch :-)
bleonard··on Launch HN: Hightouch (YC S19) – Sync data from data warehouses to SaaS tools
There are probably some nuances one level down. Things our users have told us they can do in these areas that, to my knowledge, Hightouch doesn't do:

* Combine data from different sources to define a model. We'v seen using Postgres as a source of truth and supplementing with Snowflake data, for example.

* Add tags to contacts in mailchimp, zendesk or make lists of them in customer.io, Pardot, etc based on segmentation. I believe Hightouch Audiences is more like a filter.

* Full workflow with branches, PRs, test suite in a repo. I saw Hightouch added git syncing to a known branch yesterday and it looks cool, but it's not the full workflow yet.

I'm certainly trying to keep it in the friendly-competition area, especially on this thread :-)

bleonard··on Launch HN: Hightouch (YC S19) – Sync data from data warehouses to SaaS tools
Congrats on the launch! Hightouch looks great and this need is real. Things seem to be going well, so I don't think I'm taking too much away by mentioning that we have been been working on Grouparoo, an open source alternative that solves similar pain points.

A few differences: git developer workflow focused (branches, CI, PRs, etc), ability to self host, segmentation in destinations (tagging people in mailchimp based on rules, for example)

https://www.grouparoo.com

bleonard··on Ask HN: What do you think about ETL tools and Data warehousing technology?
We just finished up the Open Source Data Stack conference, which is all about this topic. Feel free to check out the reply.

Specifically, open source approaches to the modern data stack where the trend is picking the right tools for the job that revolve around the warehouse central data store.

The pieces discussed were around getting data in (Snowplow events, Meltano ELT), transforming it (dbt), reporting (Superset), getting it back into tools (Grouparoo Reverse ETL), and orchestrating things (Dagster).

https://www.opensourcedatastack.com

bleonard··on Development Workflow for Reverse ETL
We are super excited to share our new developer tooling that makes these integrations so easy. Let me know if you have any questions
bleonard··on Slidev – Presentation Slides for Developers
Here was my attempt several years ago that used reveal.js

https://github.com/bleonard/jekyll-slides

I found it useful to have notes as well that showed up when presenting on the other screen.

bleonard··on Ask HN: Is there a way to efficiently subscribe to an SQL query for changes?
It's certainly not production situation, but I quickly made this [1] a while back to help with "watching" certain queries for debugging. It uses a polling approach.

[1] https://github.com/grouparoo/db_watch

bleonard··on Grouparoo: Declarative Data Sync
It's open source and free. We're still figuring out the paid offering. Add yes, it is hard :-)

My experience is that a "contact us" button for a startup is just as likely to be and MVP test as it is to be some nefarious trickery.

bleonard··on Grouparoo: Declarative Data Sync
We can use any query to bring in and are considering various ways to process it after that. That's in the near term roadmap.

The main kind of "processing" that's done now is using all these properties to calculate cohort membership (High Value users) so that all these tools can use it: Zendesk to route tickets, Marketo to trigger a campaign, even the product to change their dashboard.

bleonard··on Grouparoo: Declarative Data Sync
There are usual suspects for sources. Databases (MySQL, Postgres) and data warehouse (Snowflake, BigQuery, RedShift).

Though data out is more common, we can also bring in data from any given SaaS tool as well. For example, we have a Mailchimp _source_ that will pull in people as they signup through their form.

There is a plugin model and a few Typescript interfaces to implement to be either a source or a destination.

bleonard··on Grouparoo: Declarative Data Sync
We were on Hacker News a few months back. Since then, developers have been using the UI to set up automated data movement from their databases to Mailchimp, Marketo, Salesforce, and more.

But we also heard that they wanted it more like their normal development workflow. So we now make it even easier to sync data to cloud-based tools via declarative data models and integrations. With this, you manage data sync just like you would any other part of your stack and Grouparoo takes care of getting the right data to the places you want.

We’re excited and around to answer any questions. - comments here - Slack: https://www.grouparoo.com/chat - Email: brian at grouparoo.com

bleonard··on Launch HN: Polytomic (YC W20) – Get internal data to your business teams
This post might be of interest to you: https://medium.com/memory-leak/reverse-etl-a-primer-4e6694dc...

Astasia went through the nascent "reverse ETL" world, including Polytomic, Census, and our open source solution, Grouparoo.

bleonard··on Launch HN: Airbyte (YC W20) – Open-Source ELT (Fivetran/Stitch Alternative)
To continue to compliment this stack, feel free to check out Grouparoo, a "reverse ETL" tool that is also self-hosted. https://www.grouparoo.com
bleonard··on Airbyte: Simple and extensible open-source EL(T)
Just for fun and completeness, there's also taking action on this data outside of your warehouse by making your external tools smarter - in some cases, even writing _back_ to the things you EL'd into it. For example, writing product usage to Salesforce. These are the kinds of things we are focusing on at Grouparoo, also an open source project.
bleonard··on Show HN: Candymail – Email Automation for Node.js
Check out Grouparoo [1] for your syncing. We are working on Intercom now.

[1] https://www.grouparoo.com

bleonard··on Rails 6.1
It's cool to see "DHH" still in these changelogs.
bleonard··on Ask HN: How are you performing email flows on your early stage product?
I would recommend separating the service between transactional and marketing. Otherwise, there is always something going on with opt-out.

So use Sendgrid or Mailgun for transactional. Hit their API to send a mail. There are pros/cons around keeping the templates in there, but I would lean that way.

For drip campaigns, I would sync the product database to a tool like Mailchimp or the others ones that have been mentioned. We made an open source way to do that syncing, called Grouparoo.

Page 1 of 3Next →