HNHacker News
TopNewBestAskShowJobs

tmcneal

671 karma · joined September 4, 2010

Co-Founder @ Reflect (YC S20) - https://reflect.run

Email: todd at reflect.run

submissionscomments
tmcneal··on Effortless AI: No-Code Automation Using N8n Cloud and OpenAI Vision API
I'm not quite sure how n8n works after reading the docs, but our company Reflect provides an AI-driven approach to automation that may be similar, but for a more narrower use-case: automated end-to-end testing.

You can see a video example of how it works in our docs here (https://reflect.run/docs/recording-tests/testing-with-ai/), but the idea is that you describe the actions and assertions you want to take in plain-text prompts, and the AI interprets those prompts in real time and executes them against a running browser session. In practice, it's a lot like writing a manual test script and having it automatically execute. We use both GPT 3.5 and 4 and will be releasing Vision support once OpenAI has deemed the gpt-4-turbo-with-vision model ready for production use.

tmcneal··on Introducing Adept Experiments – use AI workflows to delegate repetitive tasks
Done! We added the ability to use it as a fixture. Documented here: https://github.com/zerostep-ai/zerostep#playwright-fixture
tmcneal··on Show HN: ZeroStep – AI actions and assertions for Playwright
Hey HN - excited to share this with you and hopefully get some feedback. ZeroStep is a JavaScript library that adds the power of AI to Playwright tests. ZeroStep’s ai() function lets developers test a web application using simple plain-text instructions, embedded directly in their Playwright tests.

The goals of this library are:

- Make tests easier to write. Writing an ai() step is equivalent to describing the action (“Fill out the form with realistic values”), assertion (“Verify there are no errors displayed on the page”), or extraction (“What is the shopping cart total?”) in plain English. The AI does the rest. There’s no need to write selectors, or add data-test-id attributes all over your app so that your tests have stable locators. Our website and GitHub repo show several examples of tricky testing scenarios that are made easy with ZeroStep.

- Make tests less frustrating to maintain. Unlike selectors, ZeroStep ai() steps are not tightly coupled to your app’s markup. The ZeroStep AI interprets your ai() steps at runtime which means even large scale changes in the app won’t break your tests, so long as your functional requirements remain unchanged. Selectors, in my opinion, are one of the worst leaky abstractions in software development. I can’t imagine how many dev hours have been lost due to them. We’re happy to not use any concept of selectors with this approach.

- Keep it simple - ai() steps have no predefined syntax (unlike Cucumber). You just need to be able to clearly describe what action, assertion, or extraction you want the AI to perform.

tmcneal··on Introducing Adept Experiments – use AI workflows to delegate repetitive tasks
Ah got it. So GPT is non-deterministic, but we somewhat handle that by having a caching layer in our AI. Basically if you make an ai() call, and we see that the page state is identical to a previous invocation of that exact AI prompt, then we will not consult the AI and install return you the cached result. We did this mainly to reduce costs and speed up execution of the 2nd-to-nth run of the same test, but it does make the AI a bit more deterministic.

There are some new features in GPT-4-Turbo that will let us handle determinism better, and we will be exploring that once GPT-4-Turbo is stable.

tmcneal··on Introducing Adept Experiments – use AI workflows to delegate repetitive tasks
Pricing is listed on https://zerostep.com - you get 1,000 ai() calls per month for free, and then the cheapest paid plan is 2,000 ai() calls per month for $20, 4,000 for $40, etc. So basically you pay a penny per ai() call.

In terms of reliability - we have a hard dependency on the OpenAI API, so that's what will affect reliability the most. We're using GPT-3.5 and GPT-4 models, which have been fairly reliable, but we'll bump to GPT-4-Turbo eventually. Right now GPT-4-Turbo is listed as "not suited for production use" in OpenAI's docs: https://platform.openai.com/docs/models

tmcneal··on Introducing Adept Experiments – use AI workflows to delegate repetitive tasks
For anyone looking to try this in an E2E testing context, we just released a library for Playwright called ZeroStep (https://zerostep.com/) that lets you script AI based actions, assertions, and extractions.

This is a working example that tests the core "book a meeting" workflow in Calendly:

    import { test, expect } from '@playwright/test'
    import { ai } from '@zerostep/playwright'

    test.describe('Calendly', () => {
      test('book the next available timeslot', async ({ page }) => {
        await page.goto('https://calendly.com/zerostep-test/test-calendly')

        await ai('Verify that a calendar is displayed', { page, test })
        await ai('Dismiss the privacy modal', { page, test })
        await ai('Click on the first available day of the month', { page, test })
        await ai('Click on the first available time in the sidebar', { page, test })
        await ai('Click the Next button', { page, test })
        await ai('Fill out the form with realistic values', { page, test })
        await ai('Submit the form', { page, test })

        const element = await page.getByText('You are scheduled')
        expect(element).toBeDefined()
      })
    })
tmcneal··on Show HN: Reflect – Create end-to-end tests using AI prompts
Hi HN,

Three years ago we launched Reflect on HN (https://news.ycombinator.com/item?id=23897626). We're back to show you some new AI-powered features that we believe are a big step forward in the evolution in automated end-to-end testing. Specifically, these features raise the level of abstraction for test creation and maintenance.

One of our new AI-powered features is something we call Prompt Steps. Normally in Reflect you create a test by recording your actions as you use your application, but with Prompt steps you define what you want tested by describing it in plain text, and Reflect executes those actions on your behalf. We're making this feature publicly available so that you can sign up for a free account and try it for yourself.

Our goal with Reflect is to make end-to-end tests fast to create and easy to maintain. A lot of teams face issues with end-to-end tests being flaky and just generally not providing a lot of value. We faced that ourselves at our last startup, and it was the impetus for us to create this product. Since our launch, we've improved the product by making tests execute much faster, reducing VM startup times, adding support for API testing, cross-browser testing etc, and doing a lot of things to reduce flakiness, including some novel stuff like automatically detecting and waiting on asynchronous actions like XHRs and fetches.

Although Reflect is used by developers, our primary user is non-technical - someone like a manual tester, or a business analyst at a large company. This means it's important for us to provide ways for these users to express what they want tested without requiring them to write code. We think LLMs can be used to solve some foundational problems these users experience when trying to do automated testing. By letting users express what they want tested in plain English, and having the automation automatically perform those actions, we can provide non-technical users with something very close to the expressivity of code in a workflow that feels very familiar to them.

In the testing world there's something called BDD, which stands for Behavior-Driven Development. It's an existing way to express automated tests in plain English. With BDD, a tester or business analyst typically defines how the system should function using an English-language DSL called "Gherkin", and then that specification is turned into an automated test later using a framework called Cucumber. There are two main issues that we've heard a lot when talking to users practicing BDD:

1. They find the Gherkin syntax to be overly restrictive. 2. Because you have to write a whole bunch of code in the DSL translation layer to get the automation to work, non-technical users who are writing the specs have to rely heavily on the developers writing the DSL translation layer. In addition, the developers working on the DSL layer would rather just write Selenium or Playwright code directly versus having to use English language as a go-between.

We think our approach solves for these two main issues. Reflect's prompt steps have no predefined DSL. You can write whatever you want, including something that could result in multiple actions (e.g. "Fill out all the form fields with realistic values"). Reflect takes this prompt, analyzes the current state of the DOM, and queries OpenAI to determine what action or set of actions to take to fulfill that instruction. This means that non-technical users who practice BDD can create automated tests without developers having to build any sort of framework under the covers.

Our other AI feature is something we call the 'AI Assistant'. This is meant to address shortcomings with the Selectors (also called Locators) that we generate automatically when you're using the record-and-playback features in Reflect. Selectors use the page structure and styling of the page to target an element, and we generate multiple selectors for each action you take in Reflect. This approach works most of the time, but sometimes there's just not enough information on the page to generate good selectors, or the underlying web application has changed significantly at the DOM-layer while being semantically equivalent to the user. Our "AI Assistant" feature works by falling back to querying the AI to determine what action to take when all the selectors on hand are no longer valid.

This uses the same approach as prompt steps, except that the "prompt" in this case is an auto-generated description of the action that we recorded (e.g. something like "Click on Login button", or "Input x into username field"). We're usually able to generate a good English-language description based on the data in the DOM, like the text associated with the element, but on the occasions that we can't, we'll also query OpenAI to have it generate a test step description for us. This means that Selectors effectively become a sort of caching layer for retrieving what element to operate on in for a given test step. They'll work most of the time, and element retrieval is fast. We believe that this approach will be resilient to even large changes to the page structure and styling, such as a major redesign of an application.

It's still early days for this technology. Right now our AI works by analyzing the state of the DOM, but we eventually want to move to a multi-modal approach so that we can capture visual signals that are not present in the DOM. It also has some limitations - for example right now it doesn't see inside iframes or Shadow DOM. We're working on addressing these limitations, but we think our coverage of use cases is wide enough that this is now ready for real-world use.

We're excited to launch this publicly, and would love to hear any feedback. Thanks for reading!

tmcneal··on Ask HN: Great text based games to play?
If you're a fan of Tolkien, check out MUME (Multi-user Middle Earth), a MUD that's been online since 1990: https://mume.org/

Its unofficial community site is http://elvenrunes.com/ which has hosted forums and player-submitted "logs" (text logs of PvP fights) for over 20 years.

tmcneal··on Ask HN: How do you start a startup in your 30s when you have wife/kids/mortgage?
I feel compelled to provide a counter-point to all of the "you can't" comments here.

It's possible. I started a company with my co-founder three years ago. We both were mid-thirties with kids when we started. We've since went through YC, raised a seed-round, and the company has been growing ever since.

What worked for me:

- I saved up money and drew from that while we were getting the company off the ground. My co-founder and I went close to a year without a salary. Having a spouse that works also helps.

- Having a relatively low cost-of-living helped. It never felt like we were reducing our quality of life, although this was in the middle of the pandemic so we weren't able to go on vacations etc even if we wanted to.

Building a company is a lot of work, but that's true regardless if it you're married with children or not. I think parents have the advantage of being forced to be efficient with their time. It did require a lot of work after-hours coding to get things off the ground, but I probably would've been coding on something anyway.

tmcneal··on Oldest and Fatherless: The Terrible Secret of Tom Bombadil (2011)
This reminds me of the fan theory that Tom Bombadil is secretly the Witch-King of Angmar: http://www.flyingmoose.org/tolksarc/theories/bombadil.htm
tmcneal··on Ask HN: Best dev tool pitches of all time?
Is it possible to personalize your pitches to individual users? At our startup [1] we try to get straight to point when pitching the product and demo something that is as close as possible to how the person we're talking to would actually use the product.

For example, here's a video I just recorded a few minutes ago for someone that I've been talking to via email: https://www.loom.com/share/01fd4a6963a04258908f7b12e2afaa3a

One advantage we have is that it only takes a few minutes to show the product, and it works on any publicly available site so with a little research it's pretty easy to show something that's pretty close to how they'd use the product themselves.

[1] https://reflect.run

tmcneal··on Show HN: tdm – Terraform for Test Data
tdm is an open-source library for generating seed data for your QA and staging environments.

I'm the co-founder of Reflect (https://reflect.run) which is a no-code tool for creating automated regression tests. We've realized that an accurate tool for building and running tests is necessary, but not always sufficient when it comes to being successful with automated regression testing. If your tests can't make assumptions about the state of your application, it doesn't matter how accurate your testing tool is; your tests are going to be flaky.

Every end-to-end test makes some implicit assumptions about the state of the application. For example, if you were testing an e-commerce store, you'd create a test that clicks through the site, adds a product to the cart, enters a dummy payment method, and validates that an order has been placed. There's going to be lots of baked-in assumptions in this test about the state of the application. If that product goes out of stock next month, or its product name changes, or its size / color attributes are different, then the test will probably fail. We built tdm to help software teams manage the underlying data that their end-to-end tests depend on. It's meant to run in your test environment just before your test suite executes, and it gets your application in a consistent state so that the implicit assumptions in your tests are correct.

tdm operates like a Terraform for test data; you describe the state that your data should be in, and tdm takes care of putting your data into that state. Rather than accessing your database directly, tdm interfaces with your APIs. This means that the same approach to managing your first-party data can also be used to manage test data in third-party APIs. If the API has an OpenAPI spec, we can auto-gen Typescript bindings to make integration easier.

Test data is defined as fixtures that are checked into source code. These fixtures look like JSON but are actually Typescript objects, which means your data gets compile-time checks, and you get the structural-typing goodness of TS to wrangle your data as you see fit.

Similar to Terraform, you can run tdm in a dry-run mode to first check what changes will be applied, and then run a secondary command to apply those changes. With this "diffing" approach, any data that's generated by the tests themselves gets cleared out for the next run.

I'd love to get any feedback on this approach! Hopefully this is something that can be useful regardless of what you're using to build E2E tests.

tmcneal··on Playwright: Automate Chromium, WebKit and Firefox
I wonder if the infrastructure that's driving the browser tests is underpowered. If the browser process is dying silently or CPU is getting maxed out, it could manifest in what you're describing where it happens intermittently and there's very little to go on. I'm assuming you're running these tests in Chrome... You could check the Chrome debug logs to see if anything is being spit out there.
tmcneal··on Playwright: Automate Chromium, WebKit and Firefox
Yes definitely, there's lots of products in the QA space trying to tackle the problem you're describing. I'm a co-founder of a no-code product in the space (https://reflect.run). Being no-code has the advantage of enabling all QA testers to build test automation, regardless of coding experience.
tmcneal··on Playwright: Automate Chromium, WebKit and Firefox
I'm not sure if any of these are pertinent to your tests, but these are the issues I see most often that cause flaky tests:

- Hard-coded waits in your code, like "Thread.sleep(1000)". A better alternative is to replace hard-coded waits with something that waits for an element or value to appear on the page. i.e. click on a button and wait for a 'Success' message to appear. Puppeteer and Playwright both have good constructs for doing this.

- Needless complexity in the tests. Conditionals in particular are a code-smell and indicate there's something needlessly complex about the test.

- No test data management strategy. The more assumptions you can make about the state of your application, the simpler your tests become. Ideally tests are running in an environment that nothing else is touching, and you're seeding data into that environment before tests run. I personally don't believe in mocking data in regression tests since that quickly becomes hard to manage.

We spend a lot of time thinking about these issues at my company and wrote a guide that covers other common regression testing issues in more detail here: https://reflect.run/regression-testing-guide/

tmcneal··on YC’s $500k Standard Deal
This also incentivizes founders to raise at a higher cap to minimize the dilution from the $375k. I imagine it'll be harder for investors to negotiate a lower cap or get a discount.
tmcneal··on DuckDuckGo's daily search queries surpassed 100M, a 47% increase
They're making more than $100mm a year via ads and have been profitable since 2014 according to this Techcrunch article: https://techcrunch.com/2021/06/16/on-a-growth-tear-duckduckg...

They use Bing's ad network (or at least they used to) and don't expose any privacy details about their users, so advertisers can't target you based on demographic or psychographic information.

tmcneal··on DuckDuckGo's daily search queries surpassed 100M, a 47% increase
DuckDuckGo publishes daily traffic metrics themselves here: https://duckduckgo.com/traffic
tmcneal··on Show HN: I made a Chrome extension that can automate any website
Congrats on the launch DK! It's awesome to see how polished you've made this.
tmcneal··on Chromium: Permit blocking of view-source: with URLBlocklist
For context, this is is not allowing websites to prevent View Source for all users. Rather it is a setting that administrators can set on Google Chrome installations within their company/school/etc.

That said, I highly doubt that Google would take this feature request seriously if this were not associated with another Google property. For example, I can't imagine that if Typeform users opened a similar bug ticket that it would see any traction. To me this highlights a danger I had not considered before around the monopoly power of Google - inane requests that would benefit properties within the Google ecosystem get undue consideration despite it harming the broader ecosystem.

tmcneal··on Beg Bounties
There seems to be a cottage industry of folks scraping places like ProductHunt and hitting them up with these emails. We posted to ProductHunt twice and got multiple "beg bounties" along with someone claiming they could get us to #1 product of the day, for a nominal fee of course.
tmcneal··on Ask HN: Which NoCode platforms are fine?
That's fair - if you compare us to raw compute we are going to be well over an order of magnitude more expensive, even with the volume discounts we offer at our "contact us for pricing" tier.

Usually when folks are considering us, they're comparing us to the other alternatives which generally are either (1) have folks on staff do each regression test pass manually, (2) hire additional devs/automation engineers to build a code-based automation suite, or (3) have existing developers spend part of their time building and maintaining a code-based automation suite. When you look at the cost of those alternatives we tend to be pretty competitive.

Where we're not a good fit pricing-wise is when someone wants to run very simple tests at a high volume, like running synthetic testing / monitoring scenarios in production, or for load testing.

tmcneal··on Ask HN: Which NoCode platforms are fine?
Wow, thank you!
tmcneal··on Ask HN: Which NoCode platforms are fine?
We're a no-code platform for testing, if that counts. :) Should work pretty well for testing applications built with no-code tools since the same person who builds the app can also write tests for it.

Would love to get your feedback if you do try us: https://reflect.run/

tmcneal··on Ask HN: Which NoCode platforms are fine?
Co-founder of Reflect here: Thanks for the mention!

I think a lot of people think of "no-code" tools as something exclusively for non-developers; things like Webflow and Bubble for creating apps. I don't think that's ever really been the case, Zapier being a great example of that, but there's been a lot more tools coming to market in the past couple years that actually work well as a way to replace tedious code-based workflows. i.e. Use Retool instead of building UIs for internal apps, use Reflect to build automated tests instead of Selenium, etc.

tmcneal··on Launch HN: Rainforest QA (YC S12) – No-Code UI Test Automation
Hey there - I'm one of the co-founders of Reflect so just to give my perspective:

The workflow for creating tests in Reflect is pretty similar to Rainforest: we both expose a "cloud browser" that loads up your webapp and you interact with that to create your tests. The biggest difference workflow-wise is that Reflect records all your actions automatically, whereas with Rainforest you often need to both specify what step you're going to take, and then actually perform that action in the browser itself. Recording everything automatically is technically harder to pull off since it's forced us to ensure we accurately record every step you take, but we think it makes for a better workflow since you can create tests faster, and there's less chance of inaccuracies that cause tests to not be repeatable.

I would quibble with the statement that you need to be far more technical to use Reflect - we're a no-code product after all. :) We have plenty of folks who aren't developers using our product. But the good thing is that both products have free tiers, so users can always give us both a try for free and decide for themselves.

Edit: Also their statement about Reflect running in headless mode is incorrect. Our test grid is a cluster of VMs: we spin up a Docker container for each test run, and each Docker container is running the test steps using a normal non-headless browser.

tmcneal··on How big tech runs tech projects and the curious absence of Scrum
Github Issues, Linear, or ClickUp
tmcneal··on We killed our end-to-end test suite
Our product is an end-to-end testing tool, so it's always interesting to see what issues companies hit with E2E tests and how they solve them. What's interesting about Nubank's experience is that after deleting their E2E suite, they realized that replacing them with integration tests wasn't providing enough value. There's a lot value in E2E tests, but so many orgs end up taking the wrong approach and ending up with a slow, flaky test suite.

We wrote a guide [1] for building automated test suites based on our experience working with and talking to software orgs. Teams who get value out of E2E tests generally do the following things right:

1. They keep tests as small as possible. This makes maintenance easier and forces a separation-of-concerns in the tests.

2. They factor the tests so they can run in parallel. This, plus shorter tests, is the best way to mitigate the slowness issue brought up in the article.

3. They have a good strategy for test data management. It looks like Nubank had test data represented as fixtures, but then somehow manual testing in that same environment was clobbering test data and causing false failures. A better strategy for managing test data could have solved for this. Or maybe even just running the automated tests in an isolated environment.

[1: https://reflect.run/regression-testing-guide/]

tmcneal··on While posting to Tumblr, E and W keys just stopped working
I remember visiting tumblr.com once a long time ago and they were actually serving the index.php file as a download rather than rendering the homepage. It was indeed a several thousand long file of spaghetti code.
tmcneal··on How we use web components
It's surprising how poorly supported Shadow DOM is across the major UI testing frameworks. To access a Shadow DOM element you'll typically need to write a JS shim that uses the DOM APIs to get at the element, and then do other tricks to work around things like closed vs. open shadow roots and handle how event propagation works with shadow DOM.

Shameless plug, but we're building a no-code testing framework (https://reflect.run) that does all of this for you automatically. Works great with Shadow DOM (both open and closed) as well as more esoteric things - for example did you know Salesforce's Lightning framework overrides how Shadow DOM works? :D

Page 1 of 4Next →