HNHacker News
TopNewBestAskShowJobs

tbrannan

17 karma · joined April 5, 2024

submissionscomments
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Great minds think alike!
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
because it is, but its still true lol
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Thanks!! means a lot coming from you. Best of luck at Supabase.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Really appreciate the input. I'll make sure to give you early access once we implement this, I'll keep you posted.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Noted, we've been focused on Postgres first but SQL Server keeps coming up. Appreciate the feedback.
tbrannan··on Show HN: Prism.Tools – Free and privacy-focused developer utilities
I haven’t come across anything else like this. It’s genuinely impressive.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
This is useful. What if you ran a CLI locally that extracts just the statistical profile from prod cardinality, relationship ratios, etc. and uploaded that? We'd never touch your database, you just hand us the metrics and we match the shape.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Not yet, but you're the second person in this thread to call out distribution control as a gap. It's on our radar now. Thanks for the feedback.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
That is a really good point one-to-many relationships blow up fast. The trunk table idea is interesting, would simplify how people reason about limits. Appreciate the feedback, genuinely helpful!
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Appreciate the Snaplet comparison, they were doing good work. You're right that realistic looking strings are the easy part. We're focused on relational integrity first (FKs, constraints, realistic cardinality), but business domain logic is the next layer. What kinds of rules would be useful for you? Things like weighted distributions, time-based patterns, conditional relationships?
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Different focus, ShadowTraffic is config-driven and optimized for streaming/Kafka workloads. We're schema-driven: point us at your DDL and we generate relational test data automatically. Less config, more just give me test data that fits my tables.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Thanks for the feedback! Honestly, we're still dialing in the tiers, what row limits would feel reasonable to you for your use case? Always helpful to hear what people actually need.
tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Thanks! At 1M rows, I think a few things matter:

Streaming: Can't hold it all in memory. Generate in chunks, write, release, repeat.

Format choice: Parquet with row groups is fast and compresses well. SQL needs batched inserts (~1000/statement). Direct DB writes via COPY skip serialization entirely is usually fastest.

FK relationships: The real bottleneck. Pre-generate parent PKs, hold in memory, reference for children. Gets tricky with complex graphs at scale.

Parallelization: Row generation is embarrassingly parallel, but writes are serial. Chunk-then-merge is on our radar but not shipped yet.

What does your stat product need, realistic distributions or pure volume/stress testing?

tbrannan··on Show HN: DDL to Data – Generate realistic test data from SQL schemas
Thanks! To clarify, the core engine isn't AI. It's deterministic pattern matching, so it runs in milliseconds with no token costs. There's an optional "Story Mode" that uses AI for narrative-coherent data (like "a churning SaaS with seasonal trends"), but the base product is just schema parsing + smart type inference.

The difference from Faker: you don't write any code. Paste your CREATE TABLE, get data back. Faker is a library you have to integrate, configure field-by-field, and maintain as your schema changes. Different use case — more like "I need a seeded database in 30 seconds" vs "I'm building a test suite."

Fair point on pricing though, still figuring that out. Appreciate the feedback.

tbrannan··on Show HN: Claude Reflect – Auto-turn Claude corrections into project config
I recently learned about the hooks and skills feature. This is cool