Cue, an open-source data validation language
cuelang.org
cuelang.org
Cue is a project originally started by Marcel van Lohuizen who previously was part of BCL (Borg Config Lang) at Google. The main use is to generate config files.
See the Kubernetes examples at: https://cuelang.org/docs/tutorials/
Here are two posts discussing the motivations for Cue over BCL/Jsonnet:
- https://github.com/cuelang/cue/issues/33#issuecomment-483615...
- https://github.com/cuelang/cue/discussions/669
A very interesting development is that Grafana appears to be adopting Cue as a first-class configuration option. See: "Bring new CUE-based config schema system to release-readiness" https://github.com/grafana/grafana/issues/33139
This could mean that a future where Grafana dashboards can be two-way synced with a git repo will eventually exist.
----
Other tools with some industry adoption in the "Infrastructure as Code" space include
- Dhall
- Jsonnet (from BCL)
- kustomize
- Helm
- kubecfg
- Tanka
- SkyCfg
- jkcfg
- Krane
- HCL (Terraform)
And two tools that fall into a separate class of enabling "Infrastructure as Software"
- Pulumi (TypeScript/Go/Python/.NET)
- CDK
Two-way sync with a git repo is one possible path, and we've talked a lot internally at GL about how to best support it. My sense is that we can do it with relatively little friction and likely will - but if you're just syncing with a git repo, there's still a lot of arbitrary, opaque repo layout decisions that still have to be made (how do you map a filesystem position for a dashboard to a position in Grafana? In a way that places the dashboards next to the systems they're intended to observe? With many teams? With many Grafana instances?) which induce new kinds of friction at scale.
Fortunately - and not mutually exclusively with the above - by building the system for schema in CUE, we've made a composable thing that we can make into larger systems. That's what we're starting to do with Polly: https://github.com/pollypkg/polly
Conveniently, my parts of a Grafanaconline talk tomorrow discusses both of these https://grafana.com/go/grafanaconline/2021/dashboards-as-cod... :D
I recall seeing another project HN which created dashboards out of a yaml description. This seems like a fantastic idea, given that a lot of business panels and dashboard apps can be implemented with a limited set of UI interactions.
It is very flexible. Will publish more examples soon!
Although, writing apps in yaml or json works really well currently. We mostly express logic through Lowdefy operators. We like supporting yaml / json since it is easy to write code to create or update such apps in any language.
It'll be possible to serialize/represent the dashboards in CUE. Here's a handwavy, pseudocode-y example i use in the talk: https://gist.github.com/sdboyer/d76196f94ca78d1e84c739e95e64...
That said, it's not like we're planning on replacing all the "export JSON" buttons in the Grafana UI with "export CUE." One of the interesting properties of defining schemas in CUE is how it allows us to remove schema-defined default values from a dashboard's JSON. The JSON representation can actually look a lot more like a concise CUE representation.
> versioning
Versioning of the Grafana schema is the essential design goal of the "scuemata" system that is under discussion in that epic issue
Versioning of artifacts that are instances of the schema is a key goal with Polly https://docs.google.com/document/d/1GU0DGy-X6z4FVwbJYPsBKRdq...
> non-visual editing
Like, editing something other than raw code in an editor? Yes, this is also something directly enabled by the schema (again, see the Polly doc, the "Produce" heading). For data-intensive tasks, such editing experiences are the only way to see your logic in the context of data, and therefore IMO prerequisite for confidence
> reproducibility
Yup. This already isn't "hard" to do today, but reproducibility gets more complicated at scale - those questions about how to map what's on disk to what's in your Grafana (or whatever app) in my parent comment become more complicated, leading to friction, leading to staleness.
In most projects data validation becomes problematic. In a most of cases the schema could be a lot more defined than what type def offers. This allows for test cases to make sure data fits the model.
We've also been creating a DSL to build web apps. Check out Lowdefy [0] - I'm trying to come up with an "Infrastructure as Code" word for Lowdefy. "UI as config" is the closest fit, but not sure...
* "Inheritance, is not commutative and idempotent in the general case"
* "A value is always final in CUE, it can only be made more specific."
From an engineering perspective, the latter is definitely more appealing. But I lack well articulated stories to understand how inheritance fails short, and how graph unification fares better. I wonder if there is somewhere a simple concrete example to contrast the not-idempotent inheritance approach vs. the graph unification approach.
Inheritance allows you to override properties/attributes. If you inherit from 2 classes that both specify the same attribute/property, but with different types for the same attribute, one of them takes precedence and overrides the other. A inherits from B inherits from C is not the same as A inherits from C inherits from B if C says attribute X is a string and B says attribute X is an int.
From my understanding, the equivalent graph unification is invalid. If type A is a unification of type B and C, then B and C cannot have any overlap. Each property is either a member of B or a member of C, but never both. It's commutative because A = B | C (A is the unification of B and C) is the same as A = C | B (A is the unification of C and B). If x is a member of B, and I access A.x, I will always end up accessing B. With inheritance, there can be a B.x and a C.x. Which one I end up accessing depends on which one is A's parent.
Inheritance is not idempotent because if A inherits from B inherits from C, then A is implicitly also B and C. However, A can override B's and C's behavior, so I can't trust that calling C.x will always return the same value. It might return the type C has for that attribute, it might return the type B has for that attribute or it might return the type C has for that attribute. You can prevent overriding the types in children, but at that point you've basically built graph unification.
To give a concrete example, Python allows inheritance. If we are provided with this:
class MyCar:
# Epoch time for when the car was made
created_at: int
class MyCarV2(MyCar):
# Time it was created in RFC3339 format
created_at: str
class MyCarV3(MyCarV2):
# Using an actual datetime object
created_at: datetime.datetime
And we have a function like this: def time_since_created(car: MyCar) -> datetime.timedelta:
That function has no idea what the type of car.created_at will be. Mypy will complain at you because it's bad practice, but it's valid inheritance. Even if they all start with same conceptual time, MyCar.created_at, MyCarV2.created_at and MyCarV3.created_at return different types, despite all supposedly being valid instances of MyCar.Graph unification forces you to pick a single type for each attribute of a single type. Rather than having 3 types that behave differently, graph unification forces you condense them into one:
class MyCar:
created_at: typing.Union[int, str, datetime.datetime]
That time_since_created function now knows exactly what type created_at is. Nothing else can change the type of created_at. If you need to add another possible type you have to either add it to the typing.Union, or create a new class. You can't create a subclass of MyCar with a different type for created_at.I’ve been using Cue for over a year now, using it as the foundation for a new projet; and will gladly answer questions about our experience.
For example, for any given application we have several fixed environments--dev, staging, prod--as well as "on demand" environments for things like pull requests or individual developer environments. The configs for these environments are almost the same, but they vary based on a handful of parameters. I want to be able to write a "generic environment" module for each application and then parameterize it accordingly for each environment.
Cue doesn't seem to care much about this problem, but rather it's just trying to make sure your data is type checked. It seems more like an advanced JSONSchema rather than a typed Starlark. I think the latter would be more powerful (albeit Cue's type system is more powerful than an ordinary generic type system with things like range types).
Cue almost has an answer to the DRY problem, but you can't quite emulate functions as far as I can tell (due, I think, to shadowing problems). I wonder what people who are convinced that Cue is the future would say to this? Am I just thinking about the problem wrong?
This specifically looks like you want inheritance, which Cue eschews.
You can set up a generic config with default values for everything, and then have more specific configs that override the defaults.
A concrete example would help figure out if Cue can do what you want.
I was going to say the same thing - this pattern of having an "override" file in add to defaults, is something I've seen in multiple systems and liked. For example, it's used for JSON configuration in .NET, and with Docker Compose's YAML service configuration files.
FWIW, I'm not looking for "inheritance".
> I want to be able to write a "generic environment" module for each application and then parameterize it accordingly for each environment.
This is pretty solidly in the target use case range, i'd say - managing variations of the same "object type" over some dimension is a lot of what's targeted by the way that CUE treats directory hierarchies when loading files: https://cuelang.org/docs/concepts/packages/#instances
The main thing you have to consider in designing a layout is that you have to take a compositional approach to how you define individual config instances. That is, you can't start from prod's config, then override a value or two for staging.
If i were to do it - i have not, this is not how i currently use CUE - my first approach would probably be by defining defaults at the "policy" level (per the above link), which effectively allows you to get exactly one "override"-ish behavior.
Lots of possible approaches to this, though.
> but you can't quite emulate functions as far as I can tell
Function-like capability is present, just in a form that's less familiar. I think of them as "function structs." This post has a bunch of examples https://github.com/cuelang/cue/issues/139#issuecomment-55677.... It seems there's a plan to add a more comfortable notation (https://github.com/cuelang/cue/issues/943), but it's fundamentally possible now.
How would you solve this with directory structure and "function structs" respectively? I'm having trouble wrapping my head around the former and ran into shadowing problems with the latter.
test.cue:
#Job: {
command: string
args: string
cli: "\(command) \(args)"
}
#GoJob: #Job & {
command: "go"
}
#GoJobV: #GoJob & {
args: "-v ./..."
}
job: #GoJobV.cli
Results in: $ cue export test.cue
{
"job": "go -v ./..."
}I'd recommend starting with this article as it nicely lays out the motivations for Cue.
I'm one of the contributors. We created a DSL in the language to describe the data and create tests. You can then use that data description to validate against json, csv, avro... One of the neat things we came up with was the concept of a data trace which is like a stack trace but is a path through the data to a particular error.
My experience with Pulumi and AWS CDK is absolutely brilliant in this regard, hopefully good DevOps/SRE/WhateverNewTerm practices and patterns will reassemble good software development practices in the future.
In what sense is CUE not a real language?
Personally, on one hand I know allowing computations into configuration immediatly destroys any hope of having a tidy, rational schema in real word projects.
On the other hand though, i do believe configuration and code should be build with related tools, possibly the same tool- or at least tools using the same syntax!
(a bit like the json syntax is the same as a Python dict syntax, except this is the terrible example that is so poorly thought out that does more harm than good)
This unlocks a much greater degree of freedom and power than all the gluing together technologies that we have to do...
My 2c - thus far, i've found the language features enabled by this constraint much more useful than the expressiveness lost - at least, for the purposes i've chosen.
In my view, this is the insight and value proposition that sets cue apart from everything else, including general programming languages.
Inheritance + property overriding is the source of most problems in configuration because you can never know if a value is the source of truth.
So you could do data validation on the frontend, backend and the database server based on the same definitions.
It would save us a lot of bugs caused by different opinions of valid data in different layers of software.
Is there a second implementation in another programming language?
As I understand it, it would make sense to have libraries for processing Cue data in every language.
I'm a bit concerned about Cue relying on Go too much. A data validation language should be independent of the implementation language.
Overrides and inheritance are a world of pain. Unification and commutative operations restore sanity to the actual work of coding with a domain representation language because WYSIWYG. And you get type safety for your domain model.
The project is still at the “Read the Source, Luke” stage so caveat emptor until we get a respectable release out.
Being able to work with CUE natively in TS, though, would be a huge gamechanger for what we can do with CUE in Grafana
Dagger was demo'd at the recent Dockercon - gives a bit more of a sense of the possible with CUE: https://docker.events.cube365.net/dockercon-live/2021/conten...
1. https://github.com/timberio/vector/blob/master/CONTRIBUTING....
Our use of CUE in Grafana is an example of framework-style usage. It is a hard requirement that users never have to install the CUE CLI to perform any of our planned CUE-related tasks; rather, the needed tooling is baked into Go packages we export, and things like grafana-cli. (Avoiding a dependency on the CUE CLI also gives us a defense mechanism against breaking changes)
I've been playing with some different ways of using CUE:
- https://github.com/cuebernetes/cuebectl (declarative, continuous syncing with a kube cluster)
- https://github.com/ecordell/cuezel (bazel/make but with CUE)
both have some overlap philosophically with the "command" subsystem in CUE, but it's been fun to play with different approaches to the same problems.
I’d spend more time trying to answer that question myself but this site wants me to read an awful lot of philosophy before showing me any code. I dislike homepages like this. I’d rather you assume I buy into your philosophy, otherwise I’d leave, so you’re free to just show me what the language is like. When you do that, I’m more interested in examples than EBNF.
Eg I want to import data file and get a report that tells me 20 values in column A are negative, column B contains floats etc.
And can it munge stuff? Column C are stored as floats, but should have been of type int. If all values are whole numbers, accept it.
Mind sharing a reference?
But it doesn't seem like that's currently possible. I wonder if they have a WASM port (or similar) on a roadmap.