299 karma · joined March 1, 2022
https://twitter.com/GalenMarchetti
warehouse: you connect your warehouse AND astrobee will store your semantic layer (ontology) in your warehouse. it will also use your warehouse to do computation
connect source: you'll use our warehouse under the hood, and when you connect your source systems, data will end up being stored in our warehouse. thats where computation and ontology storage will happen
but i agree it's confusing. we are considering changing the naming around to account for this. i appreciate the feedback!
this is still an early beta, so at the moment everything is only available with OpenAI's API. however, for people who want to use it in a higher security environment, we'll support switching OpenAI with any hosted model API including on-premise or models held in private VPCs. that way people can manage their data with no exfiltration to a third party
you can have an attitude towards spending the short hours you have on this earth attempting to produce quality work that others appreciate and make their lives easier in some way, as opposed to writing those hours off as sold to someone else
wouldn't have guessed that. that NYC pull is huge
imo any hope of really leveraging llms in this context needs this + human review on additions to a shared ontology/semantic layer so most of the nuanced stuff is expressed simply and reviewed by engineering before business goes wild with it
could be cool to see that unbundled into its own thing though
you can even use ingress and special URL for multiple services in a one-off deviation, but things get really messy when you want to test state-level things too like db migrations.
besides making it more convenient to handle the chain of dependencies and duplicating all the services, we also want to make it easier to test db/state migrations and larger features without having to do an expensive e2e application deploy
we use trace headers to keep track of where the request originally came from to route it to the right database. it's transparent to service B as long as service B is properly propagating the request context from incoming request to outgoing requests
environment is an incredibly overused word for devs. development environment might even mean IntelliJ, not a k8s cluster with your dev versions inside of it
been going back and forth a lot on how to make this more clear without introducing additional confusion
in reality what you're doing with namespace-based deploys or telepresence is equally "lightweight" to what you're doing with Kardinal. it's just that (at least in our phrasing), those don't quite constitute a separate "environment" as state is shared between any developers working at the same time
Kardinal matches those approaches in terms of light-weightedness, but offers state isolation guarantees too (like isolation for your dbs, queues, caches, managed services, etc). so in comparison to "ephemeral environment" approaches that give state isolation, we do believe we're doing this in the most lightweight way possible by implementing that isolation at the layer of the network request rather than by duplicating deployed resources
and thanks for the +1 on the name haha, we called it Kardinal because the goal is to deploy only "unique elements of your ephemeral environments" across the organization, i.e. number of services deployed equals the "cardinality" of service versions
several commenters have mentioned Cue/Jsonnet/friends as great alternatives, others find them limiting / prefer pulumi with a general purpose language
our solution at kurtosis is another, and tilt.dev took the same route we did...adopt starlark as a balanced middle-ground between general-purpose languages and static configs. you do get the lovely experience of writing in something pythonic, but without the "oops this k8s deployment is not runnable/reproducible in other clusters because I had non-deterministic evaluation / relied on external, non-portable devices"
But I see from the comments it is decently confusing!
Although I see how the overloading of the term is pretty confusing here
Our very first users were test engineers who used our tool to spin up test environments for E2E and integration tests, and we called it Kurtosis because we imagined them catching “tail” errors that were out of the expected norm for devs focused on only one component of the system. We imagined a tool that detects bugs happening in a “high Kurtosis” error environment.
We’re also big fans of the book “antifragile” by Taleb, about strategies in general for excelling in the face of volatility environments (can stretch that to kind of look like high “Kurtosis” environments)
The combination of those two things felt right when we started building - but I can see all the confusion it’s now causing here!
To give a bit of context on the name, our first set of use cases was all about test engineers using this tool to testbeds for end to end or large scale integration tests.
Since we are pretty mathy ourselves, we called it “Kurtosis” because we imagined the distribution of errors arising from service to service interactions to have a high Kurtosis - in the sense that there’s a lot of errors you wouldn’t “expect” to happen from a first principles understanding of each component. You’d have put them all together to see those, they’d be “far off the mean”. There’s also a lot of stuff about how we view our work, where we like exposing ourselves to outlier opportunities that we hadn’t previously imagined to produce results that would only happen in a “high Kurtosis results distribution”.
Now that being said, I definitely hear what you’re saying. It’s not obvious that’s where the name came from…and just because we were thinking that when we named it doesn’t mean it resonates with our users the same way!