Pulumi 3.0
pulumi.com
pulumi.com
I like the multi-language premise of Pulumi but I have the feeling that over time one of the languages will gain a significant advantage over the others (better supported, more features, native...) and thus will become the primary language for Pulumi package authors and users (similar to HCL for Terraform and TypeScript for CDK).
Meaning that eventually all the "serious" Pulumi users will end up converging on one single language (similar to this comment of a CDK user migrating from Python to TypeScript). Feedback loops will only help amplify this effect (e.g. more Pulumi users on language A -> more examples online -> more package authors working in this language -> better established best practices -> more submitted issue tickets for language A -> language better supported by Pulumi -> more users...)
Curious about how Pulumi multi-language components [0] work under the hood. Isn't writing a "Pulumi schema" describing resources the same as writing declarative HCL at the end of the day?
[0] https://www.pulumi.com/blog/pulumiup-pulumi-packages-multi-l...
Indeed what you say is true of many other "multi-language" platforms. I was an early engineer on .NET at Microsoft, and although it was multi-language from the outset (COBAL.NET was a thing!), the reality is most folks write in C# these days. And yet, you still see a lot of excitement for PowerShell, Visual Basic, and F#, each of which has a rich community, but uses that same common core. A similar phenomenon has happened in the JVM ecosystem with Java dominating most usage until the late 2000s, at which point my impression is that Groovy, Scala, and now Kotlin won significant mindshare.
I have reasons to be optimistic the infrastructure language domain will play out similarly. Especially as we are fundamentally serving multiple audiences -- we see that infrastructure teams often elect Python, while developers often go with TypeScript or Go, because they are more familiar with it. For those scenarios, the new multi-language package support is essential, since many companies have both sorts of engineers working together.
A "default language for IaC" may emerge with time, but I suspect that is more likely to be Python than, say, HCL. (And even then, I'm not so sure it will happen.) One of the things I'm ridiculously excited about, by the way, is bringing IaC to new audiences -- many folks learn Python at school, not so much for other infrastructure languages. Again, I'm biased. But, even if a default emerges, I guarantee there will be reasons for the others to exist. I for one am a big functional language fan and especially for simple serverless apps, I love seeing that written in F#. And we've had a ton of interest in PowerShell support since many folks working with infrastructure for the Microsoft stack know it. And Ruby due to the Chef and Puppet journeys.
I also won't discount the idea of us introducing a cloud infrastructure-specific language ;-). But it would be more general purpose than not. I worked a lot on parallel computing in the mid-2000s and that temptation was always there, but I'm glad we resisted it and instead just added tasks/promises and await to existing languages.
As to the Pulumi schema, you're right, that's a step we aim to remove soon. For TypeScript, we'll generate it off the d.ts files; for Go we'll use struct tags; and so on. Now that the basic runtime is in place, we are now going to focus there. This issue tracks it: https://github.com/pulumi/pulumi/issues/6804. Our goal is to make this as ridiculously easy as just writing a type/class in your language of choice.
As someone that has worked largely on backend systems and infra for a long time, are colleges actually training these infra skills? Some app engineers at a place I consult for was just having this conversation last week, interested in how to train infra skills and pretty much everyone had sort of fallen into infra roles over their careers, with no formal training in it. Most of us were just *nix hackers as kids, learned to program at some point, and now we’re here. Working on infra is quite a lot more than just knowing the language.
I’m not mad at having something other than HCL — years ago when I worked at Engine Yard we developed a cross platform cloud library that let us write things in Ruby which was nice. But when thinking of solving infra problems I’ve never once thought “you know, if I could just write this in Python these problems would go away”. Actually I personally hate Python as a language, I’d much prefer to write Go or Rust or TypeScript, and it does feel like a bonus that everyone touching infra sort of just has to use HCL which removes a lot of bike shedding.
Totally open to improvements! More isn’t always better though.
I am using Terraform 100% now, but sometimes wish I had more than the HCL (hashicorp configuration language) syntax available to use in my code.
Pulimi on the other hand is closer to terraform, except you can write in a language which actually doesn't stand in your way and allows you to programmatically access values available during execution only. With CDK you can only reference them in CF template (lets say instance ID which you just created), but these are never "lifted" to your code. Also CDK misses native 'datasources' and offers limited mechanism for lookups.
Pulumi also has a killer Automatioin featre, where you can code infrastructure migrations, not unlike you'd do it for SQL. Neither TF nor CDK allows you to do that, and you'd need to code infrastructure state transitions in error prone bash wrappers.
As a developer, I quite liked Pulumi resource model, documentation and team responsiveness on Github and slack channel. I am working with CDK now, because of costs and I'd prefer Pulumi if there was a choice.
I must admit that I'm skeptical on how well this would actually work though. Infrastructure API's tend to be flaky, and I'm not sure how reusable the migration code would actually be. If its too flaky, people will just not use it.
CDK indeed is just another layer of abstraction on top of CloudFormation, itself already a leaky complex abstraction, and it bites you in a lot of ways.
I'm still much happier using CDK vs CF or Pulumi vs HCL though.
Will have to look into it now.
Pulumi is more like Terraform but without HCL. So imagine instead of having to find-the-HCL-way-to-do-things and work within those constraints to make terraform work, you can use one of a few languages to build up resources (similar to you would in HCL) and have pulumi provision/act on them for you including managing their state.
Quite different to CDK/CloudFormation and much more like Terraform without the handcuffs of HCL. (sometimes handcuffs are useful tho..)
It is nice to have one system or even repo for managing all of that.
https://www.pulumi.com/docs/intro/vs/cloud_template_transpil...
As mentioned in other comments here, CDK for Terraform is another option and it's not AWS specific - it should work with any Terraform provider and there are heaps available.
> This experimental repository contains software which is still being developed and in the alpha testing stage. It is not ready for production use.
Pulumi started with actual programming languages instead which gives you full expressiveness, IDE integration, and more, without learning yet another new language. And since infrastructure is just represented as objects in the language, creating and combining them in complex ways is much more natural than wiring up configs.
The good:
- Tooling for any major language (Typescript, Python)… is lightyears ahead of anything you will find for Terraform / HCL.
- You can write your code as declaratively as possible (like you would do using TF), but you always have the escape hatch of using all libs available to your language of choice.
- AWS CDK uses CloudFormation under the hood, so you get cool stuff like automatic rollbacks in case of failure.
- You have access to much more mature testing frameworks, compared to what is available to TF. Because AWS CDK synthetize CF templates, you can also snapshot those for regression testing. Applying good software engineering practices is overall much easier. Most languages are much easier to extend than HCL.
- AWS is constantly releasing higher level libraries so you don't need to fiddle with low level API details. When using Terraform, you generally must understand your cloud provider API at a very fine level to implement IaC.
The bad:
- AWS CDK is written in Typescript and automatically translated to other languages. Even though you can use Python, C#, etc... Most examples and tutorials exist in Typescript. Also, you always must install npm + cdk + the libraries in the language you are using, so using any language other than Typescript means supporting two toolchains, which is a pain in the ass. I started with AWS CDK in Python and now I'm migrating to Typescript.
- Some modules only exist for Typescript.
- Since CF is used under the hood, it only supports resources supported by CF. Sometimes CF support takes a while. Since TF is just a wrapper for a Cloud Provider API, it generally implements new resources much quicker.
- It has its own vocabulary and and ways to do stuff (L1/L2 Construct? Stack? Retention Policy?)... even if you've used a lot of TF and know AWS very well, there's a learning curve.
Some cool videos:
- An AWS dev shows how to create your own construct and test it (that's the less "toy example" tutorial I have found): https://www.youtube.com/watch?v=cTsSXYOYQPw
- Same guy shows how to contribute to AWS CDK. I've found that looking at their source code is a great way to learn about good practices and patterns: https://www.youtube.com/watch?v=OXQSSibrt-A
I never have to worry about losing the state or ending with the incomplete state, and no matter what, I can always just delete the whole CF stack (in rare cases, you might have to retry deletion, but it never loses resources.)
That said, I don’t think that “generates CloudFormation JSON” is a bad attribute of CDK, but it is a major difference between Pulumi and CDK. In particular it means arbitrary values cannot be used for control flow at run time with CDK, since the template has a static generation step. However, it also does not require an additional tool at runtime, which has value in some cases.
This is noteworthy because many times CDK's constraints match CloudFormation's constraints. I've noticed a few issues in the CDK github repo that a feature is incomplete because they are waiting for a change or feature to land in CloudFormation first.
I'm surprised this is possible. If it was, why didn't Terraform follow this approach a long time ago?
The cloud providers' APIs provide the endpoints for creating, updating, reading and deleting resources. Terraform and Pulumi's value is to provide an idempotent abstraction on top of that. And that abstraction is not straightforward to write, because it has to handle numerous nuances and anachronisms in the underlying APIs. For example, in the event of updating an existing storage bucket, the abstraction has to determine whether the bucket can be simply updated, or if it needs to be re-created (say if you were changing the name or location). And the underlying API will not necessarily reveal this kind of information.
Hence one would think completely avoiding handwritten code is incredibly difficult if not an insurmountable problem.
You are right that it's not easy. Thankfully the cloud providers themselves have moved in the direction of auto-generation for their own SDKs, documentation, etc., which has forced this to get better over time. This is motivated by much the same reason we've done it -- keeping one's own sanity, ensuring high quality and timeliness of updates, and ensuring good coverage and consistency, when supporting so many cloud resource types across many language SDKs.
Microsoft, for instance, created OpenAPI specifications very early on, and standardized on them (AFAIK, they are required for any new service). Those API specifications contain annotations to describe what properties are immutable (as you say, the need to "re-create" the resource rather than update it in place). Google Cloud's API specifications similarly contain such information but it's based on the presence of PATCH support for resources and their properties. Etc.
The good news is that we've partnered with the cloud providers themselves to build these and we expect this to be increasingly the direction things go.
Much sweat and tears has gone into hand writing Terraform provider code. The vast majority of which has come from and continues to be maintained by volunteers.
To have had to replicate this manual effort all over again just to create a competitor would have been silly.
There will surely be a lot of wrinkles to iron out with this 100% automated approach. But indisputably this is a positive development.
And even then, magic modules only covers some of the resources in the GCP provider. The remainder are completely written by hand. Presumably because the Ruby objects are not sufficiently expressive to cover all the edge cases.
Some things in the AWS API aren't (yet) available in CFN.
But I'm skeptical.
https://github.com/hashicorp/terraform-provider-aws/issues/7...
This would make it easier for software for dynamic schema built on top of a relational database. Dynamic schema is one area where most libraries don't help you out much, but it's an increasingly important feature, especially for businesses.
If your app is built with Rails, you can use this library to help you on that (I'm not affiliated with it): https://github.com/jrochkind/attr_json
It's super useful and compelling at scale, but there are some potential downsides. For one, it's a real program and IMHO you should treat it as such with the same rigor of documentation, testing, etc. as the production code you're deploying. This might actually make your processes a bit slower or more risky. Because if you don't do that then you just risk building unmaintainable, untrustworthy deployment code that nobody wants to use.
The second downside is you're taking a big bet on Pulumi to stick around. At this point, it's probably a safe bet. But the more you use their SDK and system the more tightly coupled your production deployment is to the continued existence of Pulumi. If they go bust and can't maintain the SDK you might have a production system that can't be deployed anymore.
Sounds like a neat and useful project for complex setups.
There's really no definitive 'winner' or best option yet. Each has strengths and weaknesses relative to others. It might just come down to how a team prefers to manage and work with infrastructure.
Still, Terraform has my needs at this stage and looks a lot more mature. But if I was a small web host or something this could be a really interesting integration.
In that case, you can save your state in an S3 bucket, just like you save your Terraform state somewhere.
For me the biggest benefit is the fact that I can write my infra with Types in TypeScript.
See native provider.
The first is on constructing the DAG of resources, the second is using that DAG to orchestrate the changes from the previous state.
The problem I had is that values that are only available in the second phase (eg a subnet ID) can't be accessed in the first phase directly. The subnet's ID is accessed via "my_subnet.id" but it's actually effectively a future.
You can't write something like "if my_subnet.id == "1234" (this is an arbitrary example), but effectively a reference to a resource's attributes is only available on the "right hand side" of an expression, and pretty much as a simple assignment only.
To use the value of an attribute, you essentially need to access the value in the second phase of the pulumi process, and to do that (at the time) involved essentially pushing through source code to be "eval"ed on the "other side".
I'll be interested to see what's changed :)
After about a year of use, I simply cannot go back to editing YAML files or clicking things in a web UI. It's ruined me completely.
Since all infra resources are defined as code you end up defining them in a procedural manner (Python), and struggle to arrange/resolve dependencies.
The next big stumbling block is your license...
> The SDK is a CLI and collection of libraries for defining and deploying cloud apps and infrastructure in code.
https://www.pulumi.com/pricing/
With Terraform the open source CLI and libraries for defining and deploying infrastructure are fully usable, even for large systems.
What I would like to know is, how usable Pulumi is without a paid subscription?
I thought when I first looked at Pulumi, the only non-local state backend was the paid Pulumi Service. But looking at it now, they seem to support the normal object store backends that Terraform does (https://www.pulumi.com/docs/intro/concepts/state/). Though the talk at lack of concurrency control seems to imply they don't support locking like Terraform does.
Pulumi is fully functional in open source form. The analogy I like to draw is Git and GitHub. You can use Git fully independent of GitHub, or you can choose to use them together, for a seamless experience within a team. (Not a perfect analogy since we built both Pulumi open source and the Pulumi SaaS, which causes this very confusion!) We don't hold anything back, if it's in the SDK, it's open.
We recently added concurrency control to the alternative backends. I'm sorry the docs are confusing on this matter -- we will get that fixed up. We also have many large customers in production on the open source alone. It's easier with the SaaS just because we handle security, reliability, and sharing with your team along with access controls, auditing, etc. But if you prefer to roll your own there we are entirely happy to have you in the community and help out. Admittedly our marketing materials aren't super clear here and we are working to fix this.
Hope this helps to clear things up and again apologies for the confusion.
I think its a reasonable business model. The SDK is completely open and free. The service is how they make money .
State is a snapshot of what is deployed at a given time, it is used at the next run to compare if there were any drifts since, eg.: "Did anyone delete a server manually?" etc.
By default it is stored with Pulumi, but if you are on a budget you can just use S3 with a few tradeoffs (concurrent deployments would need a lock that you implement yourself):
These tools are just configuration management for the cloud. No other configuration management tool requires your changes conform to something that was pre-existing in a state file. They simply modify the system state to reach what your intended configuration is, without needing to track state.
Orchestration tools that rely entirely on state files to make all changes are poorly designed. State snapshots are a limited view of the past, and do not reflect the actual current state. So you have to grab the state, then see what has changed, and wonder if what has changed was intentional and should be preserved, or if it needs to be overwritten. This is basically distributed change consensus, like merging Git trees, or Paxos. But tools like Terraform basically throw away the current real-world state, like ignoring what's in the mainline branch, because merging is, like, hard, man.
State files are useful when your CM tool cannot understand which resources need to exist in what form. Sometimes you may need to maintain certain infrastructure, but your tool doesn't have an easy way to determine what resources it is maintaining versus the resources that are not managed by it. But in most cases, there are various ways to detect that in code, rather than recording it in a state file. Every CM tool in the world does this, because you don't really care about the previous state as much as reaching your desired state!
The other way state files are useful is as a log of past actions. But that's not the way tools like Terraform use it; they lean on it like a crutch. Rather than just telling you what has changed since the last Terraform run, they can often cause Terraform to just refuse to apply changes, or fail to detect and import existing resources if they weren't put into the state file manually (terraform import) or during terraform apply*.
If you want, you can run pulumi with the --refresh flag and you can see just how much slower it is.
Ansible does not need a state file to manage AWS R53 records, but it does support a cache. (This is not an endorsement of the tire-fire that is Ansible.)
Terraform's .plan file is like a cache before you apply changes. Since it's synced up to the state file too, it will fail to apply if the state has changed, so it's useful to "safely" apply only changes that seem like they might work. The problem is the state file acts like a giant boat anchor tied to the HCL code the rest of the time.
You don't need a state file for Ansible, or Puppet, or Chef, or Salt, or CFEngine, or literally any other such tool.
If you have never tried to destroy all your Terraform resources before, it will probably not work. You'll have to modify your Terraform to get it to jump through the correct hoops in the correct order in order to destroy without dying.
At this very moment I am fighting with Terraform to destroy some resources as part of a code change. AWS wants them destroyed in a very particular way, and Terraform won't create the resources I do want because apply is dying because the destroy is failing.
"If this is your first time running pulumi new or most other pulumi commands, you will be prompted to log in to the Pulumi service. The Pulumi CLI works in tandem with the Pulumi service in order to deliver a reliable experience."[0]
[0] https://www.pulumi.com/docs/get-started/kubernetes/create-pr...
[0]: https://www.pulumi.com/docs/intro/concepts/state/#backends
Nice!
Not allowing self-hosting for non-paid tiers is cutting control off from the user and moving it to the producer.
Terraform cloud is a similar deal.
Apart from that, it's an amazing tool to work with. The company I'm working at right now uses it extensively. All of our microservices have Pulumi in the CI/CD pipeline. It's an extremely powerful tool and I'd even use it for personal projects over provisioning resources through a vendor like AWS.
And when I used it they also had quite permissive pricing for personal use, where you could store your tfstate or equivalent on their servers which is nice as well for solo devs, who doesn't want to build or shell our for a storage platform or CI for terraform.
The only thing I had some problems with was errors, which wasn't the easiest to decode. But it wasn't too bad.
What's holding me back right now is that supported for Azure Functions is still in its infancy for code inlining. It's a shame, since I'm stuck in Azure for now.
[0] https://www.pulumi.com/docs/intro/concepts/function-serializ...
- Docker: https://www.pulumi.com/docs/reference/pkg/docker/
- Postgresql: https://www.pulumi.com/docs/intro/cloud-providers/postgresql...
- Keycloak: https://www.pulumi.com/docs/reference/pkg/keycloak/
Is there a tutorial out there that demonstrates using _just_ the Pulumi SDK, perhaps in a CI/CD setting, to deploy a simple service to Kubernetes?
Many thanks!
Suppose you wrote a program to go and create your cloud infrastructure. Using AWS’ APIs you write an app to spin up a VM, setup a load balancer, and provision a database. You run your program, it works like a charm. Now you have all the infrastructure setup just the way you want it!
The problem is what happens when you want to _update_ that infrastructure? The program you wrote, using the aws.CreateVM(..) and aws.CreateLoadBalancer(..) API calls is no longer applicable. Instead, you probably want to use a different set of APIs, like aws.UpdateVM(…), to update your existing cloud resources. For example, change some port settings on your load balancer. So your app needs to be smart enough to check if the resources already exist, and if so, update them. Otherwise, create them fresh.
And it gets even worse. What if you want to create some new resources, such as attach an SSL certificate to your load balancer… but still keep all of your existing infrastructure as-is. Or what if you want to update an IAM usage policy that is already in-use by several other resources… Somehow your app needs to know the impact of that change, and how it will ripple out across other cloud resources.
Does that start to make sense? You don’t really want a “wrapper” for cloud APIs. You really want something that allows you to effectively describe your cloud infrastructure, and “make it happen”. And leave the specifics of “how” as an implementation detail… accomplished by a cloud provider’s APIs.
That is what Pulumi does — and other Infrastructure as Code tools, like Terraform. It provides you a way to describe your cloud infrastructure in a programming language, so that every time you run your app it will make the cloud reflect that target state. It will:
- Create resources if they don’t exist. - Update existing resources if they do. - Delete any resources that you no longer need.
I work at Pulumi and am happy to go into details about the joys of not dealing with cloud APIs directly, and just using Pulumi :)
And in addition to making it easier to manage cloud resources by defining that state in a programming language, Pulumi can do other interesting things with your resource graph too. For example, analyze resources and check that they are compliant with security best practices and what not. https://www.pulumi.com/docs/get-started/crossguard/
With it, you can bring up some infrastructure on machine A and tear it down on another machine.
Without the service, you can only tear down on the same machine you brought it up. Ditto for adding to existing infrastructure.
You don't need their service but you have to solve problems like knowing what is running, replicating the needed state etc.
The service is very nice and makes it easy to know what is running, how long it has been running, who created it etc- across clouds. But, it is not strictly needed. As an individual user, you could get away without using. For a team, it fills a very vital gap.
They will surely lose points for reduced brevity.
I'm the only ops-y person in my group. I've done things so far with Terraform HCL, but a) it's not a great language, and b) I can't really ask other folks to learn it. But they all know Python, so my theory is that if I can wrap our ops stuff up in familiar-looking code, they'll be able to work with it effectively.
Is one of these toolkits better than the other? I'm inclined to go with Terraform CDK just because they company's further along, but I'd love other people's takes.
I last tried Pulumi around the 2.0 release and it flat-out didn't work for me. After working through a few bugs with their tech sales I gave up on it. I'm guessing their TypeScript support is better than their Python - for me it just wasn't ready for production.
On the other hand I've been using Terraform CDK in production since the early alpha releases. Had to work around a couple of bugs here and there, but no blockers at all. It's been a game changer for my team and we're really enjoying it!
This was very off-putting for me, and I had to give up on Pulumi. Is something changed on this regard?
The cloud offering is just there to streamline the storage + fancy dashboard.
Identical to how terraform works.
So this is a case of me not looking well enough while exploring many different solutions.
I have to say that reading this passage on one of the getting started guides (https://www.pulumi.com/docs/get-started/azure/create-project...):
> If this is your first time running pulumi new or most other pulumi commands, you will be prompted to log in to the Pulumi service. > The Pulumi CLI works in tandem with the Pulumi service in order to deliver a reliable experience. It is free for individual use, with features available for teams.
did not help. At least mentioning that the Pulumi service could be self hosted would have helped.
Thanks again.
Unless your infrastructure is so large and managed by so many departments you need the least powerful declarative language to stay sane, my suggestion is to go with Pulumi.
It's not perfect, the whole infrastructure-as-code ecosystem is frankly pretty bad, but if Pulumi is a 6/10, Terraform is a 4/10. YMMV
If you're just getting started with Terraform be sure to check out CDK for Terraform. It allows you to drive Terraform from a variety of high-level programming languages instead of it's DSL.
What? :(
From an end-user perspective, once you have your cdktf environment setup you don't really ever have to deal with node.js at all.
I've never had a reason to dig into this very deep, but my understanding is the cdktf Python libraries use JSII Python libraries which interact with the main JSII implementation which runs under node.js. That's where the conversion from Python to HCL compatible JSON happens. Or in the case of the main AWS CDK, from your favorite programming language to CloudFormation.
More info on JSII here:
Other than that, it's really annoying way to keep infrastructure. I prefer the Terraform DSL myself, but I'm certain there is an audience that disagrees with me.
> if you think typecasting the AWS api in your local code repository is useful, then Pulumi is the way to go.
What do you mean by typecasting here? Isn't providing the types what Pulumi (or TF CDK for that sake) does by the libs it provides? I dot see how I, the consumer of those libs, needs to "type cast" anything.
> someone will find an api that AWS didn't properly typecast before releasing and then the hundreds of hours of typecasting in your local repo will be worth it.
?