HNHacker News
TopNewBestAskShowJobs

dangoldin

7,247 karma · joined March 25, 2008

Currently leading engineering team @ TripleLift (http://triplelift.com/). Formerly I cofounded Makers Alley (http://makersalley.com) and Pressi (formerly Glossi) (http://getpressi.com/dan). Before that I had a stint as a wall street quant, a management consultant, and a brief stint as a product manager.

I write at http://dangoldin.com.

I do a minimal amount of consulting on work with a focus on quantitative engineering, scaling systems, and architecture review. If you have an interesting project with some fun challanges definitely reach out!

Contact info: dangoldin gmail

submissionscomments
dangoldin··on IBM to acquire Confluent
I led the engineering team of a large adtech company (TripleLift - order of hundreds of billions of events/day) and we evolved from self hosting Kafka, to paying a vendor (Instacluster), to migrating to RedPanda.

RedPanda was a huge win for us. Confluent never made sense to us since we were always so cost conscious but the complexity/risk of managing a critical part of our infra was always something I worried about. RedPanda was able to handle both for us - cheaper than Kafka hosting vendors with significantly better performance. We were pretty early customers but was a huge win for us.

dangoldin··on Ask HN: Who is hiring? (December 2025)
Twing.AI | AI/Full-Stack Engineer & Forward Deployed Engineer | Remote (US-friendly timezones) or NYC | Full-time or Part-time

We’re a lean (4 people, 3 of which are engineers), bootstrapped team building a private-equity automation platform (deal analysis, document processing, chat, investment theme research). Shipping fast, learning daily, and looking for hands-on engineers who want outsized impact.

- Stack: React Router 7/React 19, shadcn/Tailwind v4, Node/TypeScript (strict), PostgreSQL + Prisma + pgvector, S3; AI workflows with embeddings & LLMs.

- Why join: small team with direct ownership of core product, zero bureaucracy, Greenfield decisions and ability to learn extremely quickly, you’ll ship to paying users weekly.

AI / Full-Stack Engineer

- Own features end-to-end: data ingestion → AI pipelines → UI. - Design and tune LLM/embedding flows; keep latency/cost in check. - Requirements: strong TypeScript/Node, modern React, SQL + data modeling; comfort evaluating/using AI models; bias to ship.

Forward Deployed Engineer

- Work directly with customers (PE analysts/partners) to turn messy workflows into productized features. - Rapidly prototype, measure impact, harden for production. - Requirements: full-stack chops, great product sense, calm under ambiguity, willing to jump between code, data, and customer calls.

Happy to also talk to the self-taught. We care more about ability than credentials and value slope over intercept.

How to apply: Email jobs@twing.ai with “HN – [Role]” in the subject, a short note on something you’ve shipped, and any links (GitHub/portfolio). No recruiters.

dangoldin··on AWS multiple services outage in us-east-1
I worked at an adtech company where we invested a bit in HA across AZ + regions. Lo and behold there was an AWS outage and we stayed up. Too bad our customers didn't and we still took the revenue hit.

Lesson here is that your approach will depend on your industry and peers. Every market will have their won philosophy and requirements here.

dangoldin··on Ask HN: How to Deal with a Bad Manager?
I'd reach out to your former manager that you got along with and ask for their advice. Seems you had a good rapport and they know the company dynamics.
dangoldin··on Dbt Labs acquires SDF Labs
I'm sure you've heard of SQLMesh but that seems like a potential fit. Or is it still too heavy handed?
dangoldin··on GitHub Git Operations Are Down
Yea - definitely. Just not ideal and something that needs to be built out, tested, etc.
dangoldin··on GitHub Git Operations Are Down
Problem is that often you also end up relying on GitHub for CI/CD so not as easy of a change. Imagine GH being down and you need to deploy a hotfix. How do you handle that? Especially, if you followed best practices and set up a system where all PRs need to go through code review.
dangoldin··on Same Query, Different Hash: Snowflake's query hashing blind spots
Yea - the idea is that Snowflake will generate these after a query runs in order to help you look at multiple runs of the same query. So imagine you run a query that's "select a from b where c = 1" and you want to find all examples of that query running. That's where "query_hash" comes in. But Snowflake also says well what if we let you be generic about the parameters - so "where c=1" and "where c=2" and "where c=300000" all have the same query_parameterized_hash.

That's the intent but turns out it's only doing a very simple hashing and not actually looking at the canonical version of the query. For example it won't treat aliases/renames as the same even though it should. This makes it harder to look at all queries that are in essence doing the same thing.

dangoldin··on Exploring How Cache Memory Works
Really cool stuff and a nice introduction but curious how much modern compilers do for you already. Especially if you shift to the JIT world - what ends up being the difference between code where people optimize for this vs write in a style optimized around code readability/reuse/etc.
dangoldin··on Bento: Open-source fork of the project formerly known as Benthos
Yea - I get that argument but these days it's just hard to do infra as true FOSS with the hyperscalers and current cloud economics. There is a community license and and the code is visible. Not saying it's ideal but Redpanda is further into the open source world than WarpStream.
dangoldin··on Bento: Open-source fork of the project formerly known as Benthos
FWIW - Redpanda open sources their core product - https://github.com/redpanda-data/ while WarpStream keeps their core product proprietary - https://github.com/warpstreamlabs
dangoldin··on Building an open data pipeline in 2024
You don’t need to. dbt/sqlmesh are competitive. I just like the model of sqlmesh over dbt but dbt is much more dominant.
dangoldin··on Building an open data pipeline in 2024
From egress + storage cost standpoint absolutely which ends up being a big factor for these large scale data systems.

There’s a prior discussion on HN about that post: https://news.ycombinator.com/item?id=38118577

And full disclosure but I’m author of both posts - just shifted my writing to be more focused on the company one.

dangoldin··on Building an open data pipeline in 2024
Author here. Basic idea is you want some way of defining metrics. So something like “revenue = sum(sales) - sum(discount)” or “retention = whatever” which need to be generated via SQL at query time vs built in to a table. Then you can have higher confidence multiple access paths all have the same definitions for the metrics.
dangoldin··on Building an open data pipeline in 2024
Yes but most data-heavy tasks are parallelizable. SQL itself is naturally parallelizable. There's a reason Apache RAPIDs, Voltron, Kinetica, Sqream, etc exist.

Full transparency I don't have huge amount of experience at working on this massive scale and to your point you need to understand the problem and constraints before you propose a solution.

dangoldin··on Building an open data pipeline in 2024
Author here and there's nuance here but as a rule of thumb data size is a decent enough proxy. Audience here isn't everyone and the goal was to give less experienced data engineers and folk a sense of modern data tools and a possible approach.

But what did you mean by "Read the first paragraph of the `Cost` section"?

dangoldin··on AWS acquires Talen's 960MW nuclear data center campus in Pennsylvania
That's when your batch jobs are running. Partially kidding but companies will wait until there's lower demand to take advantage of spot pricing.
dangoldin··on Show HN: PRQL in PostgreSQL
Not to change your direction but something I've been toying around is being able to support Algebraic types when defining tables. That way you can offload a lot of the error checking to the database engine's type system and keep application code simpler.
dangoldin··on From S3 to R2: An economic opportunity
Author here but some ideas I was thinking about: - An open source data pipeline built on top of R2. A way of keeping data on R2/S3 but then having execution handled in Workers/Lambda. Inspired by what https://www.boilingdata.com/ and https://www.bauplanlabs.com/ are doing. - Related to above but taking data that's stored in the various big data formats (Parquet, Iceberg, Hudi, etc) and generating many more combinations of the datasets and choose optimal ones based on the workload. You can do this with existing providers but I think the cost element just makes this easier to stomach. - Abstracting some of the AI/ML products out there and choosing best one for the job by keeping the data on R2 and then shipping it to the relevant providers (since data ingress to them is free) for specific tasks. -
dangoldin··on From S3 to R2: An economic opportunity
Author here - have you tried using R2? As others mentioned there's also Sippy (https://developers.cloudflare.com/r2/data-migration/sippy/) which makes this easy to try.
dangoldin··on From S3 to R2: An economic opportunity
Author here and it is true that costs within a region are free and if you do design your system appropriately you can take advantage of it but I've seen accidental cases where someone will try to access in another region and it's nice to not even have to worry about it. Even that can be handled with better tooling/processes but the bigger point is if you want to have your data be available across clouds to take advantage of the different capabilities. I used AI as an example but imagine you have all your data in S3 but want to use Azure due to the OpenAI partnership. It's that use case that's enabled by R2.
dangoldin··on From S3 to R2: An economic opportunity
Author here and really cool link to Sippy. I love the idea here since you're really migrating data as needed so the cost you incur is really a function of the workload. It's basically acting as a caching layer.
dangoldin··on Pytudes
Yes - he gives it credit at the bottom of the page.
dangoldin··on Vercel Is Down
Probably a function of what looks to be an AWS Lambda outage.
dangoldin··on Civilization 7 Is in Development
I read an interview a while back with a game developer who was asked why video games have historically had so much fighting and he responded with "it's easy to write." Take away is that as AI improves we will move to a world where conflict in video games will be very different - your example but also imagine being able to actually have to argue/convince characters in games using free form speech.
dangoldin··on Incident affecting Google Ad Manager
Back up as of 10:29PM ET.
dangoldin··on Incident affecting Google Ad Manager
From their page it's not even a "Service outage" but instead a "Service disruption." Wonder what an outage would look like.
dangoldin··on Click the Paw
From another post on the HN homepage: 'Too many employees, but few work': Pichai, Zuckerberg sound the alarm

https://www.business-standard.com/article/international/too-...

dangoldin··on Ask HN: Who is hiring? (September 2021)
TripleLift | Software Engineering | Remote & Local (USA / Canada)

We're looking for software engineers to help us scale our real time bidding exchange. We're an adtech company focused on innovative formats that power the open web and are currently running ~60 billion auctions a day. AdTech is definitely not for everyone but it gives a great foundational understanding for how the web actually works sets you up nicely for whatever comes next. It's also one of the few industries that allow you to jump into large scale and complex systems quickly.

Our official jobs are at https://triplelift.com/careers/ but even if you don't see anything listed just email me (dgoldin@triplelift.com) and I'll see what we can do.

The stack is Java (Netty) for the real time bidding system, the standard Python/Scala stack on the data engineering and ML side, and various services in JavaScript, TypeScript, and PHP.

dangoldin··on Ask HN: Who is hiring? (August 2021)
TripleLift | Software Engineering | Remote (USA / Canada)

We're looking for software engineers to help us scale our real time bidding exchange. We're an adtech company focused on innovative formats that power the open web and are currently running ~60 billion auctions a day. AdTech is definitely not for everyone but I believe it gives a great understanding for how the web works and gives you the ability to work on large scale and complex systems quickly.

Our official jobs are at https://triplelift.com/careers/ but even if you don't see anything listed just email me (dgoldin@triplelift.com) and I'll see what we can do.

The stack is Java (Netty) for the real time bidding system, the standard Python/Scala stack on the data engineering side, and various services in JavaScript, TypeScript, and PHP.

Page 1 of 34Next →