Yml Coding
cloud.google.com
cloud.google.com
And a restricted model is not there just so that non-programmers can use it. Restricting what you can do in the program means that there are more things you can do with the program. I haven't checked if GCP's Workflows supports everything below, but here are some things you can do in principle:
* You can visualize the entire program as a flowchart, and visualize the state of the program simply by pointing at a node in the flowchart. This is not possible with a general purpose language since there could be arbitrary levels of call frames.
* You can implement retry policies for each step entirely transparently, and possibly other things like authentication. Aspect-oriented programming is more practical when the programming model is restricted.
* You can schedule the steps onto different hosts, possibly in parallel.
* You can suspend and resume the workflow, since its entire state is just which step is being executed, plus a handful of variables (which presumably are always serializable).
Re the problem of extension: the idea seems to be that you put all "smartness" inside HTTP services that are written in real languages and only use this as a dumb glue language.
public void execute(String customerId) {
activities.onboardToFreeTrial(customerId);
try {
Workflow.sleep(Duration.ofDays(60));
activities.upgradeFromTrialToPaid(customerId);
while (true) {
Workflow.sleep(Duration.ofDays(30));
activities.chargeMonthlyFee(customerId);
}
} catch (CancellationException e) {
activities.processSubscriptionCancellation(customerId);
}
}If that's the case, then what you could do is still restricted by the protocol of the workflow engine - building the instructions in code gives you some dynamicism, but not a whole world of difference, and complicates things that are easier done statically. It is definitely a valid approach, but it doesn't invalidate the approach of writing out the workflow definition statically, especially if the paradigm is "put all smartness inside HTTP service and only use the workflow as a glue".
There are lot of advantages of using a general purpose programming language for implementing workflows. An incomplete list in no particular order:
* Strongly typed for languages that support it
* No need to learn a new programming language
* Practically unlimited complexity of the code
* IDE support which includes code analysis and refactoring
* Debuggers
* Reuse of existing libraries and data structures. For example can YAML based definition support ordered maps or priority lists without any modification?
* Standard error handling. In Java, for example, exception are used.
* Easy to implement handling of asynchronous events
* Standard toolchains just work. For example Gradle for Java and modules for Go.
* Standard logging and context propagation can be supported
And so on. Any new language has to have a ton of tools, libraries and frameworks to be useful. And using an existing language allows to benefit from the existing ecosystem out of the box.But what's usually more interesting when it comes to workflows is inspecting and debugging the workflows themselves, and you'd still need custom tooling for that, regardless of how the workflow is built.
The actual appeal of using YAML here is that the "code" is amenable to static analysis. YAML is irrelevant; the DSL embedded in YAML is. For example, you can easily count how many steps there are, how many edges there are between steps, etc. If you build the workflow in code, in general you can only know these after the orchestration code has executed.
In my experience, no developer ever asked for this information, especially if the price is writing code in turing complete YAML/XML/JSON based language.
I think our disagreement really boils down to different approaches towards workflow configuration. Static configuration has its place in a system where all the "smartness " can easily fit somewhere else. This is often the case when you are building a workflow that glues many in-house components together, which tend to have uniform behavior and you can easily extend them. On the other hand, if you are working with many heterogeneous components that are clumsy to extend, having a smarter, more dynamic workflow configuration API definitely makes more sense.
Configuration based languages are awesome for domain-specific use cases. For example, AWS Cloud Formation or HashiCorp Terraform configuration language are good examples of domain-specific workflow definition languages. They solve just one specific problem that allows them to be mostly declarative and omit most procedural complexity. Even in this case, I'm pretty sure that Pulumi folks would not be 100% in agreement.
The general-purpose workflow definition languages are procedural. And I believe that procedural code in YAML/XML/JSON is a bad idea. It looks ugly, doesn't add much value, and never matches any of the general-purpose languages in expressiveness and tooling. Such configuration languages work in limited situations, but developers quickly hit their boundaries in most real use cases and have to look for real solutions.
BTW Temporal and its predecessor Cadence are perfect platforms for supporting custom DSL workflow definitions. Many production use cases run custom DSLs on top of them.
Here is the stackoverflow answer with more info about the new features: https://stackoverflow.com/questions/61157400/temporal-workfl...
Why do you think AWS moved away from this model?
I cannot comment on why AWS made certain decisions as I left Amazon soon after SWF launched.
I guess that SWF was hard to use for novices. When developing Cadence and later Temporal, we fixed the majority of rough edges that SWF had. And it turned out that the core "workflow as code" model was something developers love. And with improved developer experience, the adoption is going very strong.
Here is Hacker News thread that compares the two: https://news.ycombinator.com/item?id=19733880
We've had 70 years to figure out what good programming language layout looks like. I think there is almost a consensus that it doesn't look like the linked code-wearing-a-fake-moustache .
Are people supposed to code these state machines using some sort of real language then translate it to YAML?
I can imagine that the developers chose YAML based on its popularity elsewhere (especially in the CI/CD space), not its merits.
(Oh, and a belated disclaimer: I work for GCP but not on this product. All opinions are my own, obviously.)
...what if I could store my configuration (data! right?!) in a way that is nicely separated from the logic?! No more scripting for me!
Oh wait. Some of the configuration can't be generalized about universally. Configurations fundamentally contain logic, I guess.
Oh, well then why don't I just represent the logic in the data?! That will be much better than representing the data in the logic!
....but now, you are back where you started, only instead of using something nice, standard, and powerful like Python, you have to use this... language... thing. This YAML convention you cooked up.
Separate the data from logic. Read the data into the logic. Make a nice organized place to call custom logic from... like, you know, a file directory full of scripts, which are called according to some scheme. It could be another data file. Stop there.
Like this:
$ ls -Ra
./config/do_something.yaml
./config/do_something_else.yaml
./config/config.yaml
./logs/ping.log
./scripts/do_something.py
./scripts/do_something_else.py
--- $ cat ./config/config.yaml
do_something:
target: my.stupid.server
exec_frequency: daily
...
do_something_else:
target: my.stupid.server
exec_frequency: monthly
...
--- $ cat ./config/do_something.yaml
ping: true
output:
directory: ./logs/ping.log
mode: append
--- $ cat ./config/do_something_else.yaml
delete: ./logs/ping.log
--- $ cat ./scripts/do_something.py
from stupid.library.task import Task, subscribe
class DoSomething(Task):
@subscribe
def ping(self, target):
return super().ping(target)
Then you code the program that runs collects all of the task methods, put them in an ordered list, and run them if the conditions in the config.yaml is met.Every other thing is done in a task method in python, or something like it. I don't think we need more abstraction than that.
The whole idea being that you don't want the story to be turing complete (there are no loops or conditionals with a story), but the code that executes it needs to be turing complete.
Everyone agrees we hate this, until one day it's us who is dreaming it up, and we just can't resist! There should be no concept of "flow of execution", except for maybe expansion of previously declared variables just so you can do something like concatenating two lines together. Otherwise, the lack of logical operations will make even that unnecessary.
This is why we get YAML based languages that edge towards turing completeness with conditionals and loops and ugly templating hacks. Mostly people get it wrong. I got it wrong several times when creating two earlier versions of that framework.
But it seems like this falls apart as soon as one service you need to interact with creates a requirement not anticipated by this very constrained tool-set. You need to query service A, extract something with a regex, base64 encode something else before you post to service B? Well we didn't include regexes, a module/import system, or the ability to introduce UDFs in a different language.
And if you had the resources to make all your services play into the expectations of this workflow system, you might not need to use this workflow system.
- define:
assign:
- array: ["foo", "ba", "r"]
- result: ""
- i: 0
- check_condition:
switch:
- condition: ${len(array) > i}
next: iterate
next: exit_loop
- iterate:
assign:
- result: ${result + array[i]}
- i: ${i+1}
next: check_condition
- exit_loop:
return:
concat_result: ${result}
edit: it's not April 1st yet, is it?Config languages for Ansible, K8S, etc.are just basically a bad implementation of a Lisp-ish language.
But this is missing some of the key things of lisps. It’s not homoiconic (it’s a tree-based syntax, yes, but with no tree structure unless you count nested arrays). There are no closures or lambdas. You can’t pass steps/sub-workflows around as parameters to variables (at least not what is shown) so you can’t do something like:
- a_step:
- assign:
- some_task: a_step
- b_step:
- call: $some_task
I mean, that’d be a basic element for passing function-like things around and getting the functional capabilities of lisp and let you have higher-order functions.This is a weird, basic, imperative language that’s aiming for deliberately limited capabilities. They seem to have chosen YAML only because it’s already used as a configuration language, but not because it actually provides any real value or novelty. We are very much approaching the 00s infatuation with XMLifying everything here.
array = ["foo", "ba", "r"]
result = ""
i = 0
while len(array) > i:
result = result + array[i]
i += 1
exactly 3 times less codeBut yeah, concourse configuration files are probably the worst YAML verbosity offenders, even worse than k8s manifests.
1 - https://helm.sh/
This technique isn't fool proof against maliciously crafted programs, but it might be sufficient depending on what you are doing.
This is basically Lisp without the parentheses.
It may be fair to compare the syntax and structure to s-expressions, but not to lisp itself. In the same way this language could as easily be expressed as XML.
And then they realize that it's not very useful without the other half and it dies.
https://bitbucket.org/djarvis/yamlp/
Along with Pandoc, it allows common prose from Markdown to be de-duplicated, as described in my Typesetting Markdown series:
https://dave.autonoma.ca/blog/2019/07/06/typesetting-markdow...
Obviously it's laborious insert to YAML keys everywhere, so I'm developing an editor that integrates YAML and plain text document formats (such as Markdown):
There are a number of tasks that benefit from something like this, and there are a huge number of advantages from being able to encode steps in domain specific languages such as YAML and have them execute with truly "no code".
this thing here is a language that serializes into yaml. there are zero benefits of using it except perhaps 'it is yaml so i don't have to compile my configuration'.
note that xml with xslt would have handle it better.
There are many, many benefits from having steps defined in a restrictive, tightly-controlled language that doesn't require sandboxing.
There are also many benefits to not using XML and XSLT, not least of all user experience.
"I want to call an API, and if it's a Tuesday and the response contains 'foo' trigger another API" seems pretty simple with this. That's not to say it wasn't "simple" before, but there was definitely _more code_.
And is it really that much better? I think not.
Luckily, JSON/YAML is mostly interchangeable these days as it's just a nested hierarchy of a few basic types. Heck, I mostly treat XML the same way as well.
for a straightforward collection of strings: probably good enough, for anything else xml is more likely to cover the usecase better
For example
start: !date 2020-09-12
could actually map to a native date object.