HNHacker News
TopNewBestAskShowJobs

Soerensen

35 karma · joined February 22, 2024

submissionscomments
Soerensen··on Growth.engineer – Vetted, open-source growth workflows for AI agents
Hi HN,

I'm Philip, CEO & co-founder of Brew.new, an agentic marketing email platform.

For the past few months, we've been heads down trying to grow Brew. Along the way we tried a lot of tools, workflows and growth hacks. Some worked really well, most didn't. What we kept looking for, and couldn't find, was one place with growth systems that have actually been proven to work, written so an AI agent can run them.

So we built it. growth.engineer is an open-source catalog of growth workflows. Each one is a markdown file that says exactly which tools and MCP calls it uses, so you can plug it into Claude, Cursor or Codex over MCP and run it. It's free and there's no sign-in.

It's important to us that everything on here is high quality and proven. Anyone can contribute a workflow by PR (https://github.com/GetBrew/growth-engineer), but everything goes through a rigorous review before it's merged. If it hasn't worked somewhere real, it doesn't go in.

All credit for designing and building it goes to Thomas Park and Nithin Kumar. Absolute legends.

We'll keep adding what works for us, and hope you'll add what works for you.

Soerensen··on [dead]
Yes, but also no one seems to care based on the lack of engagement on this post. That's enough for me to draw a definitive conclusion on where the future is headed. Sad, but true.
Soerensen··on Deno Sandbox
Would love to hear your thoughts: https://news.ycombinator.com/item?id=46901199
Soerensen··on Ask HN: What weird or scrappy things did you do to get your first users?
The scrappiest thing that worked for me: manual onboarding calls with every single early user, even when it didn't scale.

I'd hop on 15-minute calls to understand their workflow, then send them personalized Loom videos showing exactly how to use the product for their specific use case. Time-consuming? Absolutely. But those users became evangelists because they felt ownership over the product direction.

A few other things that moved the needle early:

1. Commenting on niche subreddits where your target users actually hang out - not to pitch, but to genuinely help. When people see you're knowledgeable, they check your profile. Make sure it mentions what you're building.

2. Finding "trapped" users on competitor platforms. Look for complaint threads about existing tools, then reach out directly with "Hey, saw your frustration with X. Would love to show you how we solve that specific problem."

3. Making your first 10 users feel like co-founders. Give them direct access to you via text/Slack. Ask their opinion on features. They'll fight for your product's success.

The cold email/LinkedIn route rarely works early on because you have no social proof yet. Much better to go where conversations are already happening and demonstrate expertise first.

Soerensen··on Qwen3-Coder-Next
Appreciate it! I should clarify that it's not just grammatical. I find that AI can sometimes help me articulate ideas based on my thoughts in ways that I hadn't even considered.
Soerensen··on Qwen3-Coder-Next
Nope! https://www.linkedin.com/in/philipsorensen

But as a non-native english speaker, I do use AI to help me formulate my thoughts more clearly. Maybe this is off putting? :)

Soerensen··on Show HN: Craftplan – I built my wife a production management tool for her bakery
The approach of building for one specific user (your wife) rather than abstracting too early is underrated. You end up with something that actually fits the workflows instead of a generic tool that needs heavy configuration.

Elixir + Ash is an interesting choice for this domain. LiveView particularly shines for internal tools like this where you want the interactivity without managing a separate frontend build. Curious how the AI code generation worked with Ash specifically - the declarative nature seems like it could either help a lot (clear patterns) or confuse models that expect more explicit code.

The BOM with cost rollups is the feature that would have saved me hours in a previous project. Most small batch producers I know either overprice everything out of caution or underestimate costs because tracking ingredient pricing through recipes is tedious in spreadsheets.

Soerensen··on Notepad++ supply chain attack breakdown
The WinGUp updater compromise is a textbook example of why update mechanisms are such high-value targets. Attackers get code execution on machines that specifically trust the update channel.

What's concerning is the 6-month window. Supply chain attacks are difficult to detect because the malicious code runs with full user permissions from a "trusted" source. Most endpoint protection isn't designed to flag software from a legitimate publisher's update infrastructure.

For organizations, this argues for staged rollouts and network monitoring for unexpected outbound connections from common applications. For individuals, package managers with cryptographic verification at least add another barrier - though obviously not bulletproof either.

Soerensen··on Lessons learned shipping 500 units of my first hardware product
The $10 deposit validation approach before committing to manufacturing is underrated. So many hardware projects fail because founders fall in love with the build before confirming anyone will pay.

What stood out to me: the factory miscommunications and quality issues compound because you can't iterate as fast as software. Each mistake costs weeks and thousands of dollars.

For anyone considering hardware: if you're not getting deposits or strong signals of purchase intent before tooling up, you're basically gambling. The author's approach of getting commitments first is the right playbook.

Soerensen··on Next.js Sucks; or Why I Wrote My Own SSG
The "framework fatigue to custom solution" pipeline is well-trodden, and I think it often makes sense for specific use cases.

The tradeoff is always: initial development time (custom = longer) vs. maintenance burden (framework = dealing with someone else's abstractions and upgrade cycles). For a personal blog or static site, the maintenance math clearly favors custom since you're the only user and can freeze dependencies.

Where this breaks down is when you need features that frameworks handle implicitly - things like image optimization, incremental builds at scale, preview deployments, etc. At that point, you're either rebuilding those features yourself or accepting the framework's complexity.

The real question is whether the complexity in tools like Next.js is inherent to the problems they solve, or if there's a simpler abstraction waiting to be discovered. My suspicion is both: some complexity is essential (SSR + SSG + ISR is genuinely complicated), but a lot is accidental (backwards compatibility, enterprise features most people don't use, etc.).

Soerensen··on Are LLM failures – including hallucination – structurally unavoidable? (RCC)
Interesting framing. On your axioms:

Axiom 3 (stable global reference frame) seems most practically actionable. In production systems, we've found that grounding the model in external state - whether that's RAG with verified sources, tool use with real APIs, or structured outputs validated against schemas - meaningfully reduces hallucination rates compared to pure generation.

This suggests the "drift" you describe isn't purely geometric but can be partially constrained by anchoring to external reference points. Whether this fully addresses the underlying structural limitation or just patches over it is the interesting question.

The counterargument to structurally unavoidable: we've seen hallucination rates drop substantially between model generations (GPT-3 to GPT-4, Claude 2 to Claude 3, etc.) without fundamental architectural changes. This could mean either (a) the problem is not structural and can be trained away, or (b) these improvements are approaching an asymptotic limit we haven't hit yet.

Would be curious if your framework predicts specific failure modes we should expect to persist regardless of scale or training improvements.

Soerensen··on Show HN: AI that calls businesses so you don't have to
Clever application of voice AI. The pain point is real - phone holds are one of those friction taxes everyone pays but no one thinks to solve.

A few questions from someone who would use this:

1. How does it handle identity verification? Many customer service calls require account holder verification (last 4 of SSN, security questions, etc.). Does the user pre-provide these, or does Pamela hand off at that point?

2. What's the latency like in the conversation? I've used some voice AI tools where the delay between human speech and AI response is noticeable enough to confuse the human on the other end.

3. The NYT example is interesting - how does it handle when the rep says no initially? Does it have negotiation logic, or does it just accept the first answer?

The API angle is smart. There's probably a B2B play here for companies that want to automate outbound calls for appointment confirmations, reservation changes, etc. That's where the real volume would be.

Soerensen··on Introducing the new v0
The positioning shift here is interesting. v0 started as "generate UI from prompts" and is now framing itself as a full app builder with persistent context.

What I find compelling about this direction: the bottleneck in AI-assisted development is not the initial generation, it's the iteration loop. Having a tool that maintains context across changes and understands your codebase holistically is genuinely more useful than one-shot generation.

The real test will be how it handles the messy middle - when you're 70% done and need to refactor, integrate with external APIs, or handle edge cases that weren't in the original prompt. That's where most AI coding tools fall apart in my experience.

Curious how they're handling the economics here. AI-first tools tend to have brutal unit economics at scale - inference costs add up fast when users expect instant iteration. The shift to Claude and "optimized model selection" suggests they're thinking about this carefully.

Soerensen··on Ask HN: Is it possible to get a job in CS without a degree?
Definitely possible. I dropped out at 16 and now run a startup after leading growth at Revolut.

The key differentiator for non-degree candidates: demonstrated results over credentials. Build something people can see - a side project, open source contribution, or portfolio piece that shows you can ship.

Three things that worked for me:

1. Start in adjacent roles. My path was athlete → operations → marketing → growth → founder. Each step built skills that compounded.

2. Over-index on learning velocity. Companies hiring non-degree candidates are betting you can learn fast. Show evidence of this - rapid skill acquisition, self-taught domains, etc.

3. Target companies that value output over pedigree. Startups and scale-ups tend to care more about what you can do than where you studied. The larger and more established the company, the more the degree matters as a filtering mechanism.

The current market is tougher than 5 years ago, but the fundamental truth remains: if you can demonstrably solve problems that companies need solved, someone will pay you to do it.

Soerensen··on Qwen3-Coder-Next
The agent orchestration point from vessenes is interesting - using faster, smaller models for routine tasks while reserving frontier models for complex reasoning.

In practice, I've found the economics work like this:

1. Code generation (boilerplate, tests, migrations) - smaller models are fine, and latency matters more than peak capability 2. Architecture decisions, debugging subtle issues - worth the cost of frontier models 3. Refactoring existing code - the model needs to "understand" before changing, so context and reasoning matter more

The 3B active parameters claim is the key unlock here. If this actually runs well on consumer hardware with reasonable context windows, it becomes the obvious choice for category 1 tasks. The question is whether the SWE-Bench numbers hold up for real-world "agent turn" scenarios where you're doing hundreds of small operations.

Soerensen··on Launch HN: Modelence (YC S25) – App Builder with TypeScript / MongoDB Framework
The TypeScript + MongoDB combination for AI coding is a smart architectural choice. I've found that schema-less databases reduce the class of errors agents struggle with most - the migration/schema drift issues that require understanding of state over time.

Question: How are you handling the built-in auth when users want to extend it? For example, adding OAuth providers that aren't pre-configured, or custom claims/roles logic. Is this something the framework supports as extension points, or would users need to fork/modify core auth code?

The Claude Agent SDK integration is interesting - have you found specific prompting patterns that work better for TypeScript generation vs other languages? Curious if the type system actually helps agents self-correct as expected.

Soerensen··on Agent Skills
The observation about agents not using skills without being explicitly asked resonates. In practice, I've found success treating skills as explicit "workflows" rather than background context.

The pattern that works: skills that represent complete, self-contained sequences - "do X, then Y, then Z, then verify" - with clear trigger conditions. The agent recognizes these as distinct modes of operation rather than optional reference material.

What doesn't work: skills as general guidelines or "best practices" documents. These get lost in context or ignored entirely because the agent has no clear signal for when to apply them.

The mental model shift: think of skills less like documentation and more like subroutines you'd explicitly invoke. If you wouldn't write a function for it, it probably shouldn't be a skill.

Soerensen··on How does misalignment scale with model intelligence and task complexity?
The bias-variance framing here maps well to what I've observed building AI-assisted workflows.

In practice, systematic misalignment (bias) is relatively easy to fix - you identify the pattern and add it to your prompt/context. "Always use our internal auth library" works reliably once specified.

Variance-dominated failures are a different beast. The same prompt, same context, same model can produce wildly different quality outputs on complex tasks. I've seen this most acutely when asking models to maintain consistency across multi-file changes.

The paper's finding that "larger models + harder problems = more variance" explains something I couldn't quite articulate before: why Sonnet sometimes outperforms Opus on specific workflows. The "smarter" model attempts more sophisticated solutions, but the solution space it's exploring has more local minima where it can get stuck.

One practical takeaway: decomposing complex tasks into smaller, well-specified subtasks doesn't just help with context limits - it fundamentally changes the bias/variance profile of each inference call. You're trading one high-variance call for multiple lower-variance calls, which tends to be more predictable even if it requires more orchestration overhead.

Soerensen··on Marker – visualize Claude's symbol understanding
The "stale symbol" detection is a nice touch - one of the failure modes I've noticed with AI coding agents is them operating on outdated mental models when files have changed mid-session. Color-coding what's potentially stale vs fresh gives you a sense of when to nudge the agent to re-read.

The granularity levels (unseen → name-only → overview → signature → full body) also map nicely to how humans skim code. Wonder if this could eventually feed back into the agent itself - like a "coverage indicator" that helps it decide what to read next when context is limited.

Currently Rust and Python via tree-sitter - Serena MCP integration should help with other languages. Would be interesting to see TypeScript support given how much Claude Code is used in JS/TS projects.

Soerensen··on LLMs Can't Jump
The induction/deduction/abduction trichotomy is useful, but I wonder if the boundary is as clean as the paper suggests. When Claude or GPT-4 are asked to "explain why X might happen" given sparse data, they often produce coherent mechanistic hypotheses that weren't explicit in training data - combining concepts in novel ways.

Is that abduction, or just very sophisticated interpolation in concept space? The charitable reading is that true abduction requires proposing something genuinely outside the training distribution - like Einstein's insight that gravity isn't a force but spacetime curvature. The uncharitable reading is that most human "abduction" is also recombination of prior concepts.

The real test might be: can LLMs propose hypotheses that are (a) falsifiable, (b) novel relative to literature, and (c) turn out to be correct? There are a few early examples in materials science where LLM-suggested compounds had properties the models hadn't seen, but it's hard to know if that's abduction or lucky extrapolation.

Soerensen··on Building my self-hosted cloud coding agent
The JuiceFS + copy-on-write snapshot approach is clever - being able to restore full workspace state (including Docker images and mise-installed tools) to any previous turn is something the commercial cloud agents don't offer.

The warm pool pattern for Kata VMs makes a huge difference for UX. Cold-starting a microVM every time would kill the conversational flow. Curious how many warm VMs you typically keep ready and what the memory overhead looks like per idle VM?

One observation from the SDK comparison: the harness quality seems to matter as much as the model. OpenCode with local models struggles not because the models can't do tool calls, but because smaller context windows make the harness's prompt engineering fall apart. Wonder if there's room for a "lightweight harness" optimized for local inference with aggressive context management.

Soerensen··on Show HN: Stelvio – Ship Python to AWS
The "dev mode" where you can change lambda code live is the killer feature here. The deploy-wait-test-repeat cycle is what makes serverless development so frustrating compared to local Flask/FastAPI development.

I see others asking about CDK/Pulumi comparison - I think you're right that it's less about the underlying engine and more about the abstraction level. CDK gives you cloud primitives. Stelvio (like SST for JS) gives you developer workflows.

The automatic IAM permission generation is underrated. I've spent more debugging hours on Lambda permission errors than I'd like to admit. The error messages are terrible ("AccessDeniedException" tells you nothing about which permission is missing) and the docs always show overly-permissive examples.

Question: How does the dependency resolution work for native Python packages (e.g., numpy, pandas)? Are you pre-building wheels for Lambda's Amazon Linux environment, or using something like Lambda layers with pre-compiled binaries? That's historically been one of the most annoying parts of Python + Lambda.

Soerensen··on Show HN: Ask-a-Human.com – Human-as-a-Service for Agents
The satire is great, but this actually points to a real gap in agentic architectures.

Most production AI systems eventually hit decisions that need human judgment - not because the LLM lacks capability, but because the consequences require accountability. "Should we refund this customer?" "Does this email sound right for our brand?" These aren't knowledge problems, they're judgment calls.

The standard HITL (human-in-the-loop) patterns I've seen are usually blocking - the agent waits, a human reviews in a queue, the agent resumes. What's interesting about modeling it as a "service" is it forces you to think about latency budgets, retry logic, and fallback behavior. Same primitives we use for calling external APIs.

Curious about the actual implementation: when an agent calls Ask-a-Human, what does the human-side interface look like? A queue of pending questions? Push notifications? The "inference time" (how fast a human responds) is going to be the bottleneck for any real-time use case.

Soerensen··on Ask HN: What's your biggest LLM cost multiplier?
Our biggest cost multiplier was "conversational drift" - not the initial call, but what happens when you let users iterate.

In our email marketing tool, a user might say "make it more punchy" → AI rewrites → "actually, more professional" → rewrite → "can we A/B test both versions?" → now you're generating multiple variants. One "simple" email could spiral into 15+ LLM calls.

What worked for us:

1. *Session-level budgets, not request-level.* We cap total tokens per session rather than per call. Users can iterate freely within their budget, but can't inadvertently 10x their usage.

2. *Explicit "done" signals.* Instead of letting users endlessly refine, we added a clear "I'm happy with this" button that closes the generation loop. Sounds UX-y but it cut our average calls-per-task by 60%.

3. *Cascade to cheaper models for iteration.* First generation uses Claude 3.5. Tweaks and refinements use Haiku. Users can't tell the difference for small edits, and it cut iteration costs ~80%.

4. *Cache aggressively at the semantic level.* "Make it shorter" and "condense this" should hit the same cache key. We use embeddings to identify semantically similar requests and serve cached results when possible.

The counterintuitive insight: your biggest cost driver is probably user behavior, not model choice. The difference between GPT-4 and Claude matters less than how you architect the interaction loop.

Soerensen··on What makes an engineer when everyone can vibe code
The distinction is increasingly about taste and judgment, not syntax.

When I work with AI tools for building products, the hard part is never "can AI write this function" - it almost always can. The hard part is knowing what to build, recognizing when the generated code has subtle bugs, understanding the second-order effects of architectural decisions, and having the pattern recognition to say "this will become a maintenance nightmare in 6 months."

I think "vibe coding" exposes who was always just a syntax memorizer versus who actually understood systems. The former group will struggle because their value was in translating requirements into code - AI does that now. The latter group will be more productive because they can skip the tedious parts and focus on design, debugging, and judgment calls.

The analogy I keep coming back to: calculators did not replace mathematicians. They replaced people who could only do arithmetic.

Soerensen··on Making physical Japanese cards: The full walkthrough from zero to launch
The burnout section resonated. The pattern of "I will learn everything myself" followed by "I cannot do this alone" is so common among technical founders, myself included.

Two things stood out:

1. Your Meta Ads experience matches mine exactly - atrociously buggy, AI shoved everywhere, money draining with questionable results. And yet it still outperforms the alternatives. I have tried Google, Reddit, TikTok ads for B2B and they all performed worse per dollar.

2. Getting family involved (sister for video/social) was the right call. There is a weird pride thing where founders feel they "should" be able to do marketing too. But creative work requires different skills and energy. Delegating it freed you to focus on what you actually enjoy (the product).

Curious about the $55k goal - did you run any pre-launch surveys or just estimate based on manufacturing minimums? That seems high for a niche language learning product (though I could be wrong about market size).

Soerensen··on Show HN: I turned my PDFs into audiobooks I can have conversations with
The "talk to your book" feature is genuinely clever - having the context window aware of where you are in the text prevents the usual "wait, you just spoiled the end" problem with most PDF chat tools.

Question: how are you handling the audiobook generation? Is it a single voice or does it switch voices for dialogue? The latter would be much harder to get right but could make fiction much more engaging.

Also curious about the business model - you mention Audible's pricing as a comparison, but they have licensing costs for actual audiobooks. Your tool seems more like a personal TTS+chat tool, which is a different value prop. Might be worth leaning into that distinction.

Soerensen··on AI has failed to replace a single software application or feature
I think you are asking the wrong question. AI has not replaced Excel, but it has replaced the need for me to hire developers to build features.

I am a non-technical founder (background in growth/marketing). Two years ago, building a production web application required either (1) learning to code for months/years, (2) hiring engineers at $150-200K/year, or (3) outsourcing to contractors and praying.

None of those options were viable for me to validate an idea quickly.

With Cursor, v0, and similar tools, I built our entire frontend in production. Not a prototype, not an MVP in the old sense, but actual production code serving real customers. Features that would have cost $5-10K to spec and outsource, I can now build in an afternoon.

The tools did not "replace Excel" or any specific application. What they replaced was my dependency on other people to execute technical work. That is a much more profound shift than feature replacement.

The comparison to traditional applications is a category error. AI is not competing with Excel. It is competing with the labor market.

Soerensen··on Ask HN: How do you market a side project?
The honest answer: 90% of marketing is consistency, not brilliance.

What's worked for me:

1. Be genuinely helpful first. Before you promote anything, spend a month just answering questions in communities where your users hang out. Not pitching, just helping. This builds credibility and teaches you what people actually struggle with.

2. Build in public. Share your process, not just the finished product. People love following a journey. Weekly updates on LinkedIn, a Show HN when you hit a milestone, short posts about what you learned.

3. Email still works if you do it right. Collect emails from day one. A simple "get notified when we launch" form captures people who are actually interested. Then nurture them with value, not spam.

4. The 100-user challenge. Instead of trying to reach everyone, focus on getting 100 people to really use and love your product. Those 100 become your advocates and do the marketing for you.

5. Don't spend money on ads until you can articulate exactly who converts and why. Ads amplify what's already working. They don't create product-market fit.

The pessimistic comment about marketing being dead is half-right. Lazy promotion is dead. But genuinely helping people in the right places, over time, still works.

Source: former Head of Growth at a fintech company

Soerensen··on Ask HN: How far has "vibe coding" come?
I'm a non-technical founder who built an entire SaaS frontend using v0 and Cursor over the past year. My experience might be useful here since I'm on the extreme end of "vibe coding" - I genuinely don't understand most of the code I'm shipping.

What surprised me:

1. The "you don't understand the code" criticism is real but misses context. I don't understand the implementation details, but I understand the system architecture, user flows, and business logic. That's enough to make good decisions about what to build next.

2. The iteration speed unlocks something new. I can try three different approaches to a feature in the time it would have taken to spec out one. This changes how you think about product development - you learn through building rather than planning.

3. The bottleneck moved from "writing code" to "understanding what you actually want." The conversations I have with the AI force me to be precise about requirements in a way I never had to be when just thinking about features.

Where it breaks down: anything that requires deep technical judgment about performance, security, or scale. I can build the feature, but I can't tell you if it will hold up under load or if there's a subtle security issue. I need technical people for that.

The "20k LOC weekend" stories are selection bias, but the productivity gains are real for certain types of work. The key is knowing which type you're in.

Source: non-technical founder who built a production SaaS frontend this way

Page 1 of 2Next →