HNHacker News
TopNewBestAskShowJobs

anerli

225 karma · joined June 4, 2023

submissionscomments
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Hey, curious about your use cases for a chrome extension, care to share more?

To answer your question - BAML is as DSL that helps to define prompts, organize context, and to get better performance on structured output from the LLM. In theory you should be able to map over similar logic to other clients.

anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Only one that's worth using ;)
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
I think the difficulty with this approach is (1) you want a good "lookup" mechanism - given a task, how do you know what cache should be loaded? you can do a simple string lookup based on the task content, but when the task might include parameters or data, or be a part of a bigger workflow, it gets trickier. (2) you need a good way to detect when to adapt / fall back to the LLM. When the cache is only a playwright script, it can be difficult to know when it falls out of the existing trajectory. You can check for selector timeouts and things, but you might be missing a lot of false negatives.
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Yeah, I think its a little tricky to do this well + automatically but is essentially our goal - not necessarily literally writing a script but storing the actions taken by the LLM and being able to repeat them, and adapt only when needed
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
For context, we have no affiliation with KeysToHeaven (though we appreciate his comment). We do think our vision-first approach gives us a significant edge over other browser agents, though we probably could’ve made that aspect clearer in the title
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Both of them are "visually grounded" - meaning if you ask for the location of something in an image - they can output the exact x/y pixel coordinates! Not many models can do this, especially not many that are large enough to actually reason through sequences of actions well
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Yeah we've though about this approach a lot - but the problem is if your final program is a brittle script, you're gonna need a way to fix it again often - and then you're still depending on recurrently using LLMs/agents. So we think its better to have the program itself be resilient to change instead of you/your LLM assistant having to constantly ensure the program is working.
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Glad you were able to get it set up quickly!

We currently are optimizing for reliability and quality, which is why we suggest Claude - but it can get expensive in some cases. Using Qwen 2.5-VL-72B will be significantly cheaper, though may not be always reliable.

Most of our usage right now is for running test cases, and people seem to often prefer qwen for that use case - since typically test cases are clearer how to execute.

Something that is top of mind for is is figuring out a good way to "cache" workflows that get taken. This way you can repeat automations either with no LLM or with a smaller/cheap LLM. This will would enable deterministic, repeatable flows, that are also very affordable and fast. So even if each step on the first run is only 95% reliable - if it gets through it, it could repeat it with 100% reliability.

anerli··on Show HN: Magnitude – Open-source AI browser automation framework
I think depends a lot on how much you value your own time, since its quite time consuming to write and update playwright scripts. It's gonna save you developer hours to write automations using natural language rather than messing around with and fixing selectors. It's also able to handle tasks that playwright wouldn't be able to do at all - like extracting structured data from a messy/ambiguous DOM and adapting automatically to changing situations.

You can also use cheaper models depending on your needs, for example Qwen 2.5 VL 72B is pretty affordable and works pretty well for most situations.

anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Try it out and report back!
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Exactly :)
anerli··on Show HN: Magnitude – Open-source AI browser automation framework
Hey! To have a framework that can effectively control browser agents, you need systems to interact with the browser, but also pass relevant content from the page to the LLM. Our framework manages this agent loop in a way that enables flexible agentic execution that can mix with your own code - giving you control but in a convenient way. Claude and OpenAI computer use APIs/loops are slower, more expensive, and tailored for a limited set of desktop automation use cases rather than robust browser automations.
anerli··on Parallel Scaling Law for Language Models
Qwen team shows how parallel streams of inference-time thinking tokens could be far more efficient than a serial stream.

Compared to scaling parameters alone, the same performance increase using their technique may be achieved with 22x less increase in memory and 6x less latency increase.

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
The small VLM (Moondream) decides when interface changes / its actions no longer line up.

We say 100% open source because all of our code (test runner and AI agents) is completely open source. It’s also completely possible to run an entire OSS stack because you can configure with an open source planner LLM, and Moondream is open source. You could run it all locally even if you have solid hardware.

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
This is definitely top of mind for us! A lot of ways to potentially approach it. We want to make sure the test case execution works really well so our focus is there but also want to think about test case generation going forward. Recording a video especially with small VLMs that can tokenize videos would be super neat.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Huh that’s an interesting use case. Yeah using an AI driven system definitely opens up some cool possibilities that aren’t possible with playwright alone. Would be curious to hear more about what you’re trying to test this way with audio.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Looks cool! Thanks for sharing! The idea of having a hybrid framework for component unit testing + end to end testing is neat. Will definitely consider how this might be applicable to magnitude.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Definitely a good question. Using an actual LLM as the execution layer allows us to more easily swap to the planner agent in the case that the test needs to be adapted. We don’t want to store just a selector based test because it’s difficult to determine when it requires adaptation, and is inherently more brittle to subtle UI changes. We think using a tiny model like Moondream makes this cheap enough that these benefits outweigh an approach where we cache actual playwright code.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
So the problem is if we cache the coordinates and click blindly at the saved positions, there's no way to tell if the interface changes or if we are actually clicking the wring things (unless we try and do something hacky like listen for events on the DOM). Detecting whether elements have changed position though would definitely be feasible if re-running a test with Moondream, could compared against the coordinates of the last run.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Hey! We can add this pretty easily! We find that Gemini Pro 2.5 works the best as the planner model by a good margin, but we definitely want to support a variety of providers. I'll keep this in mind and implement soon!

edit: tracking here https://github.com/magnitudedev/magnitude/issues/6

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Hey, awesome to hear! We are definitely open to contributions :)

We plan to (very soon) enable mixing standard Playwright or other code in between Magnitude steps, which should enable doing exact assertions or anything else you want to do.

Definitely understand the need to reduce costs / increase speed, which mainly we think will be best enabled by our plan-caching system that will get executed by Moondream (a 2B model). Moondream is very fast and also has self-hosted options. However there's no reason we couldn't potentially have an option to generate pure Playwright for people who would prefer to do that instead.

We have a discord as well if you'd like to easily stay in touch about contributing: https://discord.gg/VcdpMh9tTy

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Well yeah it's kind of ambiguous, it's just our way of saying that we're trying to use AI to make testing easier!
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
You can run it against any URL, not just node projects! You'll still need a skeleton node project for the actual Magnitude tests, but you could configure some other public or staging URL as the target site.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
The planner can plan out multiple web actions at once, which Moondream can then execute in sequence on its own. So Moondream is never deciding how to execute more than one web action in a single prompt.

What this really means for developers writing the tests is you don't really have to worry about it. A "step" in Magnitude can map to any number of web actions dynamically based on the description, and the agents will figure out how to do it repeatably.

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Originally we were actually thinking about doing exactly this and building agents for usability testing. However, we think that LLMs are much better suited for tackling well defined tasks rather than trying to emulate human nuance, so we pivoted to end-to-end testing and figuring out how to make LLM browser agents act deterministically.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Yeah good criticism for sure. We definitely want to keep this in mind as we continue to build. Some kind of accessibility tests which run in parallel with each visual test that are only allowed to use the accessibility tree could make it much easier for developers to identify how to address different accessibility concerns.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
So this is a path that we definitely considered. However we think its a half-measure to generate actual Playwright code and just run that. Because if you do that, you still have a brittle test at the end of the day, and once it breaks you would need to pull in some LLM to try and adapt it anyway.

Instead of caching actual code, we cache a "plan" of specific web actions that are still described in natural language.

For example, a cached "typing" action might look like: { variant: 'type'; target: string; content: string; }

The target is a natural language description. The content is what to type. Moondream's job is simply to find the target, and then we will click into that target and type whatever content. This means it can be full vision and not rely on DOM at all, while still being very consistent. Moondream is also trivially cheap to run since it's only a 2B model. If it can't find the target or it's confidence changed significantly (using token probabilities), it's an indication that the action/plan requires adjustment, and we can dynamically swap in the planner LLM to decide how to adjust the test from there.

anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
So the architecture is built with determinism in mind. The plan-caching system is still a work in progress, but especially once fully implemented it should be very consistent. As long as your interface doesn't change (or changes in trivial ways), Moondream alone can execute the same exact web actions as previous test runs without relying on any DOM selectors. When the interface does eventually change, that's where it becomes non-deterministic again by necessity, since the planner will need to generatively update the test and continue building the new cache from there. However once it's been adapted, it can once again be executed that way every time until the interface changes again.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
So the prompts that are sent to the planner vs executor are completely distinct. We allow complete customization of the planner LLM with all major providers (Anthropic, OpenAI, Google AI Studio, Google Vertex AI, AWS Bedrock, OpenAI compatible). The executor LLM on the other hand has to fit very specific criteria, so we only support the Moondream model right now. For a model to act as the executor it needs to be able to specific specific pixel coordinates (only a few models support this, for example OpenAI/Anthropic computer use, Molmo, Moondream, and some others). We like Moondream because its super tiny and fast (2B). This means as long as we still have a "smart" planner LLM we can have very fast/cheap execution and precise UI interaction.
anerli··on Show HN: Magnitude – open-source, AI-native test framework for web apps
Oh this is interesting. In our case we are being very specific about which types of prompts go where, so the planner essentially creates prompts that will be executed by Moondream, instead of trying to route prompts generally to the appropriate model. The types of requests that our planner agent vs Moondream can handle are fundamentally different for our use case.
← PreviousPage 2 of 3Next →