Flawless – Durable execution engine for Rust
flawless.dev
flawless.dev
For short lived workflows, you may not care about updating; just let it finish.
For longer jobs, you want some way to replace the current logic and either resume from where the job left off, or restart it idempotently. Especially if your workflow spans months or years (which at least some of these systems are designed for).
The challenge is these systens shine when you manage the job state in-memory, but they don't "store" the data in a traditional sense. They just replay your logic and replay the original I/O results. So if your logic changes, the replay breaks and your state goes bye-bye.
(I think of it similarly to React's "rules of hooks": you can't do anything that makes the function call key APIs in a different order than previous executions)
So you either accept that you can never update an in-flight job (in a meaningful way, at least), or you track job state in some other system and throw away the distinguishing feature of these systems.
I'm curious how people normally handle this. When I worked with Azure Durable Functions I couldn't find a way around this.
I wonder if there could be an approach where you have both versions live simultaneously, and introduce some sort of "checkpoint" into the old version that would act similar to a DB migration. When re-computing a workflow you could then start from the latest checkpoint, but any workflows that were created with the old version that haven't reached a checkpoint would continue to run the old code until it does.
I have spent a lot of time thinking about this and believe that the most straight forward solution for long running (or even forever running) workflows is to allow hot upgrades.
A hot upgrade would only succeed if you can exactly replay the existing side effect log history with the new code. Basically you do a catchup with the new code and just keep running once you catch up. If the new code diverges, the hot upgrade would fail and revert to the old one. In this case, a human would need to intervene and check what went wrong.
There are other approaches, but I feel like this is the simples one to understand and use in practice. During development you can already test if your code diverges, using existing logs.
Quick question: how do you prevent persisting the effect of a DOS attack on those systems?
This comment reads like ml-generated nonsense. Architecture is summed up as subjective opinions bounded by regulatory constraints, and medicine is still empirical knowledge validated through experiments. None of these fields have even a passing resemblance with engineering.
We have a long way to go.
However, I think you are very reductive in your descriptions of those fields.
Sorry I meant, as a large language model...
You conflated architecture and medicine to a builduo to "serious engineering culture". There is no way around it: it's nonsense. Even though they are highly specialized fields, neither architecture not medicine has anything to do with the basic principles of engineering. It's just nonsensical word soup.
See: ISO13485, IEC60601, QMS, etc
We know how, it’s just slow, expensive and the tooling isn’t some fantasy perfect environment where safety and effectiveness is already built in and developers just do the business logic.
You're confusing with mechanical engineering applied to medical devices with medicine.
It's like claiming bakeries have incredible rigorous engineering processes just because mechanical and industrial engineers design devices and automate processes.
Not the same thing, is it?
> See: ISO13485, IEC60601, QMS, etc
You're just pointing to standards that engineers working on specialized devices need to comply. One covers quality assurance and the other is focused on specialized electrical equipment.
As an analogy, it takes the bill-of-materials approach done by popular package managers and makes it a whole lot more rigorous.
Those would be engineers, mainly mechanical engineers or electrical engineers.
Not a fan of the new meta
[1] Which is a form of art. Go look at a school of architecture at a university it will be in the Art faculty.
[2] Which is an applied science like engineering but is not an engineering discipline.
It is deterministic in a single machine, but it unfortunately is not deterministic across machines (and specially not across architectures)
... however, if you restrict yourself to a subset of floating point, it's straightforward to provide cross-platform determinism. That's what's the Rapier physics engine does
[1]: Back in the day, CPU results could also vary based on path of execution; see tweet in article.
That's what you'll be paying for when they release the product.
// Update the `keyFrame`.
keyFrame += 0.5;
// This is just a huge state machine progression.
switch (keyFrame) {
// In the first 4 seconds we just run through the existing log messages.
case 0.5:
case 1:
case 1.5:
...
Wow! I love the design. So simple.I assume there is going to be some kind of test harness that allows developers to check their workflows.
So far, this looks more like a scripting runtime which uses Rust.
By using WebAssembly, it's kinda by default fool proof. WebAssemby explicitly requires you to declare host calls inside your modules. If you try to use another host call that is not provided by flawless, your module can't be instantiated.
It's also important to notice that there are multiple standardisation efforts going on in the WebAssembly space. For example if you are using the rust `rand` crate and compiling to WebAssembly, it's using the WASI host functions for generating random numbers. While I'm waiting for wasi, wasi-http and others to be standardised, I expose my own interface for now.
Obviously, this also has a big downside. You can't compile all Rust code to WebAssembly. However, I prefer this reject by default approach, so that you are guaranteed to never have unintended side effects.
I'm a bit worried about the limitation of having to use the flawless namespace functions for everything.
Founder of windmill.dev here which is another durability engine in Rust except it's a lot less elegant since we split our workflows into well defined steps in python/typescript/go/bash and can resume from incomplete steps only by restarting from the last step and storing the result of each step forever in our postgres db (using jsonb). The use-cases are clearly different and I can see flawless being so lightweight that you could use it to model UI flow state and scale it to millions on a small server as pointed out by the site.
This is fantastic, hopefully one day rust will power all distributed systems.
While I wouldn't normally argue for rigor or anything tedious, creating computationally expensive abstraction layers to let you develop capabilities faster without having to worry about defects or edge cases seems the opposite of what rust is designed for. I think the language will fight you the whole way. The only aid I can think of is that you can wildly unwrap stuff everywhere not worrying about panics. Which does save some time.
But, as I understand it, it works with 'normal Rust', which, if you enjoy Rust, would be normal right? It seems to compile the rust to wasm, run it in wasm, log the results while keeping pace of how far it came, and then rerun it with previous results. That's not quite nice to have, but, like you say I think, it has foot guns, and, one of the worst, is getting too comfortable thinking it has your back while this is not enough; side effects behind immutable variables are there and while this helps you a bit, it doesn't fix any of that and maybe makes it worse.
I am just a little annoyed this is not the norm but rather something special (and closed source?), why is it not everywhere? Probably because high amounts of magic?
It is. I recall it fooling me when I first got into Erlang. But it's wrong. Erlang has some tools that help lead you in a more robust direction, but you still have to work to use them. They are not automatic and it's trivial to write an Erlang service that has a single point of failure, even accidentally.
Now, I don't want to criticize the tools it has and the fact that it does rather strongly lead you in a robust direction more than many systems. This is somewhere between "an easy mistake to make for a newbie" and "some sloppy advocacy sometimes by some people", 100% a people thing, not a criticism of the code base or technicals at all. But it is important for people considering Erlang/Elixir or early in the process to understand that it does not simply automatically make all your code run robustly in a cluster. It is only a strong push in the right direction and you still must understand it enough to make sure you don't break it.
“It’s commonly known under the name durable execution, and is so new that most developers have never heard of it”
Serious? I feel like this post is very naive and perhaps disingenuous in thinking that Erlang/Elixir developers are not accounting for this. I’ve been building apps this way for a long time, regardless of language.
Rewriting everything in rust doesn’t solve any of these issues. State is state.
This tool is literally dagster or prefect but it’s claiming to be a revolutionary new idea.
I've used workflow engines before, which provide similar capabilities: execute a distributed process to completion in the presence of failures. However, you have to be really careful when building workflow applications not to introduce accidental side-effects (that are not modeled in the workflow directly), or else the result will be nondeterministic (unspecified at best).
I haven't used Dagster or Prefect but it looks like this tool uses a different approach. An application developer doesn't need to "model" their workflow using this tool - just implement it on top of the API that Flawless provides. It reminds me a bit of the AWS Flow framework [1] for AWS SimpleWorkflow, but when developing Flow applications (which are Java apps) you still have to be extremely careful not to accidentally introduce local side effects or nondeterminism.
Because this approach provides a deterministic runtime, I see it being plausible that developers could be significantly more confident that their code is, in fact, deterministic.
I haven't touched erlang in ages, but Joe Armstrong popped up in my youtube recommendations recently. After watching the video, I just thought about microservices and IaC for a while and thought "Ha, we're really fucking this up, aren't we?"
Here is a conference talk that is a pitch, if that's what you're looking for: https://www.youtube.com/watch?v=cNICGEwmXLU
yeah they should have been more humble about their system and called it Flawless :)
Erlang won’t do this out of the box but it’s a menial addition to your system.
I guess I just don’t understand how this is such a novel idea.
- Temporal / Cadence
- Amazon SWF
- Azure Durable Functions
Not even the web assembly part is novel - Temporal does that for several languages.
This seems to be a new implementation of an existing idea, and may end up being cleaner - that part is to be seen since there’s no actual product or code visible yet.
I find it a bit disheartening that these features are not more pervasive on language, runtimes, libraries and OSes, so very often the actual solutions are either very kludgy or very high maintenance, whereas the primitives are really interesting for many other endeavours.
It's not entirely clear to me how well this would work in which domains, but the overlap with domains where Erlang works well seems less than total.
That being said - what’s the relation to Lunatic [0]? Are you still working on Lunatic? Is this a side project? Or is it something completely separate?
How is this guaranteed? Isn't exactly once delivery in a distributed system impossible?
i.e., the message will be delivered exactly once if the system makes progress.
If you want an existence proof, NFSv3 had this working back in the 1980's. I doubt it was the first.
From the folks at Ziverge, who've worked on ZIO in Scala.
They use a similar approach I believe. It's discussed in this podcast: https://podcasters.spotify.com/pod/show/happypathprogramming...
like "just" WASM as a target doesn't have many of the APIs build in without non trivial "magic" trics
and for WASI might be tricky to get the exact form of determinism they choose right
like they might need to provide a custom standard library which is viable but not very convenient to use and even less convenient to maintain (if it's just about custom build flags for the "upstream" standard library it's on the other hand quite easy but likely not good enough)
through that's purely speculative
but they might be lucky there is quite a bit of benefit for rust to have a deterministic wasm build target (for wasm based derives which if combined with cryptographic hashing of inputs form a potential building block for having a shared (potential public) derive binary artifact cache speeding up CI systems and similar and allowing more complicated/advanced derive usage without having to worry about the derive compliation time adding to the initial build time)
I made a very simple snapshottable compute executer for a talk last year and it took surprisingly little code if you want a look at how it works[2], but the complexity of something like Flawless is that doing anything useful involves communicating with the outside world, where non-determinism can easily sneak in. I've been following Bernard's work on Lunatic so I'm incredibly excited to see he's tackling these hard problems.
[1] https://github.com/WebAssembly/design/blob/main/Nondetermini...
While I have an idea how to implement it, now after having read the article and comments here, how is this concept called? Does an implementation for Python exist already?
Imagine a simpler implementation for me where any earlier called subfunction would simply return the previous result, up to the point where the function was previously interrupted. Therefore these previous return values need to be stored (in what temporal calls a workflow), i.e. in a database. Sounds simple enough and decorators would probably help designing a good abstraction and keeping everything self-documenting and simple enough to understand.
That said, Temporal is very similar and supports multiple languages, including Python.
Such systems need tooling for diagnosing and fixing problems: metrics, logging, dead-letter queue, inspecting and evicting cached items, retrying dead jobs, adjusting worker settings. Flawless and similar systems will have the same problems and need the same tools.
(SCNR)
That’s not a problem per se but it does affect the ability to debug and see relevant stack traces. Its like how sometimes you see transpiled JS when what you really want is the typescript using source maps.
> Notice how flawless takes away the burden of persisting the state.
Reminds me of the Wisdom of the Well: "You'll never find a programming language that frees you from the burden of clarifying your ideas." [1]
Although solving durability of execution for arbitrary code sounds super cool, I suspect that writing the code in a way that could be checkpointed/resumed more naturally would probably be easier to debug and implement, in the end.
--