Robot Framework: generic open source automation framework
robotframework.org
robotframework.org
When I worked together with a QA team for automated testing, they used Cucumber, which uses Gherkin as an english-but-not-really DSL. Now the idea is obviously, that someone who doesn't know how to code can still describe and read test scenarios. But what ended up happening eventually is that I spent a lot of time writing code to support different Gherkin expressions, but then had to take over writing the Gherkin expressions as well, because even though Gherkin -looks- like english, it still behaves like a programming language. In the end, it would've been much easier for me to just write the tests in a programming language from the start, instead of using an abstraction layer that turned out to serve no purpose.
What I took away from that frustrating experience is that a lot of people think that learning to code is learning a syntax of a programming language, and if you could just do away with that pesky syntax, everybody could code. But, obviously, the syntax ist just the initial barrier, and actually coding is crystallizing and formulating intent - something which programming languages are much better at than natural languages.
But for some reason, this low/no-code idea never seems to die down, especially in some spaces, and I have the feeling I'm just missing something here.
I see a single advantage to these frameworks though. They force you to write a test in a specific way that looks like a real-world scenario with clear delineations between setup stages, assertions, and so on. Not that it's impossible to achieve with regular programming language, obviously.
I've seen well maintained and designed cucumber / rspec / robotframework "code" and i've seen some really horrible shit. Keyword re-use is really a joke if you are not familiar with any good practices to actual software development. Granted, in "low code" scenarios, the worst selling point is "you don't need to know how to write code" - which imho is bullshit and from my pov, statements like i quoted echoes that.
Disclaimer; i partake and contribute to RF ecosystem. And i do prefer to write my tests or process automation with it even thought i can write code in multiple languages..
GIVEN("Some setup") {
// code to create the setup
WHEN("some use of the setup") {
// do something
}
WHEN("something else"){
// do something unrelated to the previous when
}
}If you can't code, you shouldn't be writing automated tests. Most of the "benefits" of Cucumber can be gained by using a good reporting framework. I like to use Serenity as I love its reports.
I'd also say "non-coding" rather than "non-technical". Coding is just one aspect of being "technical".
There is benefit in making the tests easy to read but this is better done in actual code, e.g. the screenplay pattern.
I have never seen it work in practice, but that what it is supposed to be about. If your managers are not looking at the gherkin and making suggestions then there is no point in doing it.
Automated tests are then written based on these but not "using" these. By not "using" the Gherkin scenarios for our tests we can be looser with the language and not have to worry about RegEx.
If a manager, or other non-coder, wants to see the tests, they can look at the very pretty report that was generated.
In a previous company I had one use-case that I thought cucumber helped out a lot at.
We were writing software that would be vigourously tested and approved by a customer who didn't have any technical knowledge. Their initial plan was the "Excel sheet of death" test case management system with really fuzzily named tests.
I convinced them to write Cucumber tests, which turned out to be around 100, using the very explit syntax that then got called stupid by the cucumber devs and moved to the "training-wheels" gem. I never expressed what we were doing as cucumber tests. I just told them "Write them like this and it is easy to understand" and they agreed and did so, then they were signed off by our project team and them.
After that I then automated them our side so that we could run them quickly and easily in our pipeline while the customer could still use the tests for their tedious manual user acceptance testing at the end.
It worked out quite well. The customer never knew we automated the tests our side, they were just thrilled that something like 99% tests passed first-time and the 1% that didn't were down to interpretation differences.
The code consists of a series of invocations of functions ("keywords" in RF terms) with parameters. The syntax is tabular.
When I read the intro, it reminds me of Forth.
Although I only gave it brief reading a couple of times. (RF was used by QA at my past job.)
Worst of both worlds, IMO.
I've never seen a manual tester follow the steps that are written down, but I've seen a lot of effort put into writing down the exact steps anyway.
And that's how a lot of testing is done in various companies, especially where manual testing is still common, at least in Poland.
Granted, I've had middling success convincing non technical people to write or even look at tests, but I think it's irresponsible not to try to create this bridge, depending on the audience of the code. Because very few projects are well documented, and this syntax can act as documentation. In my current work, I am using BDD style to create high level flows, where screens are automatically captured, which is useful both for verification and generated documentation, especially when versioned with code.
It can avoid duplication (and its pitfalls), because in unit style tests, there is a text description along with the code, which can drift away from being accurate. If the test is the description, that won't happen.
It creates a more abstract model, because the tests are forced to a natural language, which will tend ot make it less implementation dependant, so the implementation can be swapped out. I've had good success with this approach, seamlessly migrating a complex workflow from one backend to another, and the tests were kinda pricesless.
I also find keyword and flow re-use to be quite high, and it's great how this approach means the more tests you write, the less low level, implementation specific code and scenarios you need to write.
At every client I emphatically insist that if they are not using Gherkin as a “Rosetta stone” for the business folks then they’re wasting their time with cucumber. Don’t go to the trouble of writing gherkin if you’re not using it as living documentation. Save yourself loads of trouble and just use an ordinary test framework.
I've read something along these lines about the choice of OCaml at Jane Street, where they argued that their investor-types are able to proof-read new code, despite not being programmers, because the language "reads like english", similar to SQL or COBOL.
It works really well for us to get a shared understanding between a non-technical business, devs and QA.
That said, we then hand it the Gherkin over to the SDETs who fettle it according to our house style (which is strict and must pass enforced linting) and it does take gardening as the product matures. It's our SDETs who write the test code: we wouldn't dream of letting a non-technical person actually drive the test logic itself.
I'm not impressed at all by Gherkin, BDD, or Cucumber. Never once has any business person gone "oooh, let me read the tests!", And in the rare circumstance they do, what you'll find is a secondary grammar develops where particular step implementations gain loaded significance, far and above the natural language implementation of a step, making things complicated.
At this point, I just tell people to keep notes, make checklists, get familiar with the data model for test data authoring, then go eith God, and always put the User first.
Mileage varies greatly, but there's a culture problrm in a lot of places where "the fiture must go out" always takes precedence over "the spec must be right".
And no one ever wants to rip apart testing tools to figure out how they work.
I presume low code/no code comes from the simplest notion of programming language, which is a medium between human and machine. And the easier it is to read and write means the "better" it is. Therefore pitching it to a non-technical or even semi-technical people is as simple as teaching people to write code at surface level.
However what people don't understand deeply is that any language is a set of symbols and rules requiring a certain amount of cognitive space and a certain time to absorb it. The more ambiguous the symbols are, the closer the growth of the number of the rules to exponential. This is a fundamental point that happens in real-life but is rarely understood, moreover to programmers, surprisingly.
I believe a proper low-code/no-code should have natural language processing capability and understanding context, not just providing syntactical aliases.
I was also a Quality Engineer for a time, and one of our bigger projects was to implement precisely a Natural Language system or something so that PMs or whoever wanted to “live in the future” could write their tests directly in the User Story.
What we did was create a large set of commonly used “phrases” to interact with systems, be it UI or APIs, which we coded the actual test framework for, document them for anyone wanting to write User Stories (in Rally platform, like JIRA), the instructions on what to do if a phrase didn’t exist for their case, etc. It took us 6 focused months, but we delivered, and people were happy. Of course, we are talking about technical PMs, PMs that met with designers and set standard guidelines, PMs that required API documentation and read it, and the like; just not “coding much” PMs.
Albeit this being a company where we already wrote automation tests right next to the developer’s timeframe for their own coding, where a feature got both code/tests right around the same time, 99% of the time, so it was easier for people to take advantage of features instead of shallowly complaining.
So, yes, I concur with you that probably your case would be better if everyone were a developer, I don’t really see the need for those kind of DSL, but if you reeeeeeally need to have other people intervene, and want them to focus more on their end of the bargain (not coding), I think these DSLs have their place: Assuming the whole company has the stomach for it.
Honestly if you know your Python you are better off with Selenium, which is probably what Robot uses under the hood anyway.
What RF provides is by far the best reporting, test tagging, test filtering, and test data injection compared to any other test automation framework. And you can automate pretty much anything using it.
use PlayWright!!
-- Robot Framework builds on top of Python and usually wraps python libraries -- such as Selenium or Playwright -- in a different API (which IMHO is usually easier to use) and provides some better reporting out of the box, so, I can't really see how the tool made people unproductive here...
The one thing I did like was the separation of concerns between a "readable" specification written in a non turing complete markup language and test code. I find pytest tests difficult to parse.
The problem is that I found the markup language they built difficult not as expressive as I needed and the tooling around the framework made things difficult to debug.
In reaction to robot/gherkin I wrote my own equivalent framework based using strongly typed YAML with an emphasis on fail fast/crystal clear error messages:
https://hitchdev.com/hitchstory/
(in the process of creating it I discovered that YAML really needs type safety and built a library called strictyaml to underpin it).
Burn. It. With. Fire.
I was at Nokia and we used robot, what an absolute pain it was
In the end you end up programming inside english like DSL with weird quirks, that's somewhere connected to python, with different variable syntax, and not really good debugging.
I guess everyone has to pass through that stage at least once, but I will never go back to it.
I don't miss Robot at all.
I'd say that it does have some quirks, but if well used it can be a good tool to have in your toolbox.
Disclaimer: I'm currently working on the Robot Framework Language Server extension for VSCode (by Robocorp, which uses Robot Framework to work with RPA -- albeit Robocorp supports well both Python as well as Robot Framework in the platform, there's a high traction on the Robot Framework side and many users enjoy it vs using just Python).
The Python code would parse those HTML tables to execute the code, lovely.
Things have changed a lot from those times though and I am somewhat happy camper with robot in its current state. Its not without its warts but definitely has come a long way since "nokia days".
You need devs to write good tests in general.
In JavaScript projects I don't have such a favourite, Jest and Mocha have been okayish when I have tried them but didn't really spark joy. In multilanguage projects (like the Playwright based robotframework-browser which I have been developing) I have enjoyed writing integration tests with Robot Framework.
I don't have excessive experience with other comparable QA tools (like cucumber), but I would say Robot Framework's main advantage and disadvantage are both the fact that it is constantly so close to python. (E.g. you can easily do in-line python expressions, but if you were to have a team working with Robot Framework where some are lacking Python competence those might become very confusing pretty fast.) Also for writing advanced libraries Python usage is pretty much mandatory. But I guess it's pretty rare for a DSL to support writing advanced extensions with the same DSL.
First of all, it's not web testing tool, it's not mobile testing tool, it's not REST API testing tool. Or any other specific testing tool. You can automate pretty much any testing activity using RF, including web, mobile and REST, but not in any way limited to those.
The killer feature from quality assurance point of view is test tagging combined with really powerful way of selecting what to include in test run and getting good reports which are supported by many many tools. Another killer feature is dead simple test instrumentation.
Regarding tagging. Let's take this example:
*** Settings ***
Force Tags feature-xyz
Suite Setup Initialize Tests
*** Variables ***
${test_environment} dev
*** Test Cases ***
Foobar
[Tags] jira-id-001 jira-test-id-001 smoke
No Operation
Lorem
[Tags] jira-id-002
No Operation
Ipsum
[Tags] jira-id-003 bug-in-jira-001
No Operation
*** Keywords ***
Initialize Tests
Connect To Environment ${test_environment}
I can select any combination of those test to be included in given test run, specify which environment to connect and send results automatically to Jira. This allows me to run only "smoke" tests against "Pull Request" environment when ever PR is opened. This also allows me to automatically run all tests every hour against "Dev" environment and submit results to Jira.Like this:
robot --include smoke --variable test_environment:pr
That would only run the one test tagged smoke and Connect To Environment would get the value "pr". robot -i feature-xyz .
Would run all tests with tag feature-xyz (and in the example file that would be all tests) against dev environment. And then I could just `curl` the XML result file from the run to Jira (given it has XRAY installed) and Jira would automatically update all the Jira tickets in mentioned in the tags with the test results. If there is no Jira Test tagged in RF test, Jira would automatically create new Jira Test for me.And in order to display test statistics in Jenkins, just install RF plugin in Jenkins and instruct your job to read the output XML and you get nice statistics, reporting etc.
That way, when you need to know what is you test coverage, just open Jira and see it yourself.
Sadly, that same thing can be also a negative side.. There is really no proper way to expose the results if you are not using Jenkins as your main CI. Ofcourse one can use junit reports or whip up a own listener that will generate some sort of test reports for the given ci platform...
Not sure what else one might need. Surely any relevant CI can also store the HTML files RF generates.
EDIT: Just realized who you are. We've met few times at RF events. I'm not sure what other information one might want besides jUnit, console output and exit code, Allure reports, HTML files and RF output.xml.
Yeah, point was that jenkins integration and how/what reports are shown is really good but utilizing junit does not provide the same experience. Way back in the past, Peke (the lead dev for those who dont know) was even griefing about RF's build in junit support "there is no standard!" - things have changed now but when i came to RF ecosystem, the buildin junit support was *shit* and it was using a format that wasn't really even close to what some other CI systems where expecting (except jenkins) ..
It was pretty slow, just rewriting the same test using Pytest sped it up by about 50% on average and that was despite our test environment being the slow part.
The Robot language was quite error prone as it uses spaces inside single keyword as well as to separate parameters. And the IDE support (VSCode in our case) wasn't particularly great. We ended up having to write custom static analysis tools, to catch some of the more common Robot problems.
And pretty much every time something more complicated was required, it was easier to switch to Python to do it. So in the end two languages were required instead of one, making maintenance more painful.
That is not to say that Robot doesn't have its good points. For example retrying on error was much easier than in pure Python. But in the end it was much more work than just using pure Python with Pytest as testing framework.
*** Test Cases ***
Write my test using a DSL
Read examples of the DSL
Write a basic test and see that it works
Write a complex test and find something missing
Learn quirks about DSL
Implement missing things on the language that the DSL was implemented on
Write part of the test in said language
Write part of the test in DSL matching regexes
Run the test
And it works 9 out of 10 times
Install flaky extension
Add flaky tag to test
And it is green
[Teardown] Reflect about my life as a software developerRobot lacks features you'd need to support larger-scale embedded test, such as with a device farm, where you need the concept of tests leasing resources. If you build a test stand with 10 testable devices and need to run a suite of 100 tests, it's ideal to run the tests in parallel as devices become available. Some test setups might require more than one instance of a testable device, or instances of more than one type of device (for example, to ensure that a V2 product can still interoperate with a V3 product.) Robot doesn't really support this.
While a feature to support this could be made extremely general (resource classes, instances, and leases) the RF developers have been uninterested in incorporating this aspect of test into their framework. The result is everyone who does even mid scale embedded device testing has to writes their own.
Another complaint about robot framework is that when you have an expensive setup like a 4 minute flashing operation, you don't want to repeat it more than necessary. So in a file, you might make the expensive setup a suite-level setup, followed by the tests cases that depend upon it. When this file grows, you might want to refactor it into a multi-file test suite in it's own sub-directory. However, these tests no longer share a suite scope (because robot's "suite scope" is actually a file scope" for legacy reasons) so in practice you may need to tolerate 3000 line files to avoid long setups.
Have they been asked? I had the same problem and had different solutions built in house with different levels of success
The problem with 3-rd party keyword-based solutions is that the resource acquisition becomes a part of the test process and could a test report to look like "ran for 6 minutes, then failed" when what really happened is the actual test couldn't be run at all because resources weren't found. It would be nice to be able to queue tests based on the resources they'll need.
Also, another feature that's more difficult with custom implementations is static analysis of the test base. If resource requests were a declarative part of the RF language (much like tags) you could do things like "run all tests that exercise resource X" or "run all tests except those which require more than one Y". This is technically possible with a keyword library but requires an additional global lexing pass of the entire test base, and this is the kind of logic you don't want to decouple from RF and it's test discovery behavior.
Perhaps Pekka can be won over in the future.
Robotframework is a great tool for expressing tests. For simple setups, Jenkins has an RF plugin to kick off tests.
The tricky part is managing many test resources. One jenkins server triggering robot tests on one bench PC with one embedded target is a great way to develop an initial test suite. But once your org gives assigns you more hardware, the tricky part is figuring out how to use it maximally (either optimizing for fastest time to completion or for least time idle of your dedicated test hardware.) This means the tests themselves can no longer "drive" the test resources directly, they need to lease them from a resource manager service.
Larger companies with nearly homogeneous hardware may get by with a basic priority queue. But some test setups with an L1 or L2 switch between them may support a DSL to allow tests to request resources that are physically or logically connected, and may need abstraction layers for common tasks such as "physically disconnect power" which the Resource Manager could implement as a SNMP command to a PDU or as a SCPI command to a DC supply -- the test itself should be declaring what has to happen, not how.
I've scoured the internet and plied counterparts in other industries for any example of commercially available prior art. Everyone I've talked to wrote their own. It took a week to produce something acceptable for my use case. But if the RF devs were to maintain a similar solution, I'd switch to that in a heartbeat.
Writing code to call a bunch of test functions then generate a report is really not hard and having control over the whole thing is nice.
You can use it to automate - Web Applications (with Selenium or Playwright) - Rest APIs - Desktop Applications (Java Swing, WPF, SAP, ..) - All kinds of hardware
It offers control-structures like most programming languages (IF/ELSE/FOR/WHILE/TRY/EXCEPT)
The Extensions for VS Code or PyCharm/IntelliJ offer Code-Completion, step-wise Debugging, Linting.
It is very hackable and can be extended using Python. It has a great API, allowing you to connect it to all kind of other tools.
I guess a lot has happened since some of the people here used it.
But I guess, if somebody is focused on using a single programming Language (like JavaScript) and is not open to learn Python or the RF Syntax I recommend to look into another native test framework
Does this framework stream sensor data? No. Does this framework control actuators? No. Does this framework represent an actuated device’s configuration space? No. Does this framework perform collision detection? No. Does this framework keep an obstacle map? No. Does this framework have any planners? No.
Just stop already. Stop calling things a robot when they are not.
I agree the name is a misnomer, but it's now part of the Enterprise Lexicon, which is filled with jargon.
I run an RPA startup, but as much as I hate the term (e.g:https://news.ycombinator.com/item?id=30755118) , once something becomes part of common language patterns, you can't unshift it.
I hate the modern usage of 'AI' as ML, but I can't change the new definition.
But soon a realize I was wrong and start to notice the power of writing tests in robot. And how it allowed to easy the understand of what operation I did on code in the past... like 6 months before, 1 year before... and also the ability to understand more clearly how other teams' features worked and to fix their test and features many times, because we more clearly wrote test using Gherkins in RF... and see how new team members could ramp MUCH more easily inside our environments, by reading tests, most of time.
I was even able to implant RF inside some other business after that, bringing together that culture of writing clear tests with RF. Some of those places had the heavy culture of writing top-down requirements precisely. On those places, we integrate RF documentation and test case procedure on our process, generating final documentation using RF test case information.
Since them I'm helping to disseminate and evangelizing for the RF use. I saw system analysts embracing the RF, writing the first version of test cases for requirement in matter of minutes. RF provided us with much faster peace in maintaining some very complex environments.
We could grab requirements, and write down cases, that latter could be implemented, edited, removed or adjusted. Whatever is the test environment, testing it is a 'living' environment, and need catering from developers. RF had helped greatly maintaining such places, to produce great reports, to create this live ecosystem easy for uses.
Is perfect? Surely No.
But is much better that all other tools available around. Is open source, easy to expand, allow use of natural language. And when well used, allow us to greatly improve our QA and DEV lives :D
After a time, when I was starting a startup with a few junior developers which lacked the expertise, we used this tool in combination with Selenium to crawl data. It was fast & easy.
Always glad to see this tool is still alive. Good job Pekka Klarck!
Robot Framework – test automation in Python - https://news.ycombinator.com/item?id=10631074 - Nov 2015 (5 comments)
We also used Robot Framework at work for automated testing, worked like charm.
> An extension which brings support for RobotFramework to Visual Studio Code, including features like code completion, debugging, test explorer, refactoring and more!
Have you tried it?
RIDE is not at least officially "not maintained" but my POV is that the current maintainer might not be up to the task -- or at least doesn't have time to dedicate to its maintenance.
But why is Robot Framework being mentioned and getting attention now? It's not like it's something new and doing anything different from the existing frameworks.
We selected it specifically for its versatility and the DSL is easy enough so even less technically adept testers can pick it up quickly.
[1] = https://pabot.org
Robot Framework is (was?) heavily used by Nokia and then spread from there to its subcontractor network.
Write some code, or get playwright to generate it for you. Don't use DSLs if you value your sanity.