562 karma · joined March 1, 2020
I don't this is immediately obvious. There is an assumption here that stars and galaxies are these inherently well defined things, therefore they are countable. I don't think that's necessarily true.
Where does a galaxy really begin and end? Or a star for that matter? If there is a solar flare occurring and solar matter is being ejected into the surrounding area is that matter still part of the star?
We have this faculty which seems to be quite good at taking complex systems and approximating them to singular entities (planet, star, galaxy, table, audience, country), we can reason about. It seems to be a core aspect of human cognition but I'm not sure if this is a prerequisite for cognition in general.
The crux of the addiction is that doing is hard and consuming is easy. I can watch a YT video on homotopy type theory (I'm not a mathematician) and feel good about myself and how much I'm learning and how I'm going to use this newly found knowledge. The reality of the situation is I'm probably not learning very much and just procrastinating.
The moment you sit down and try to create, you are almost immediately slapped in the face by your own limitations and this is an unsettling feeling. You then have a reflex to escape this discomfort by opening a new tab and navigating to YouTube and sedate yourself with that sweet, sweet content - where anything is possible.
If you're an engineer or product person, I would recommend perhaps turning this problem into an opportunity. Can you solve this problem for yourself as a product or service? If you do I'd love to buy it.
Most 'child geniuses' are in-fact child prodigies, children which display extraordinary talent in one or multiple fields (say mathematics, piano, etc.).
Genius requires an element of creativity and not just proficiency. Schopenhauer made the distinction quite nicely: 'Talent hits a target no one else can hit; Genius hits a target no one else can see.'
I think very few people are both prodigies and geniuses. One notable example is Mozart - played the piano at an incredibly high level at a very young age and also composed some creative masterpieces later in life.
I think you're right here and it's going to be something that's really hard to get right. I think prioritising the developer experience is the only way to get this right - even at the expense of other things you want to optimise for like costs.
> The true power and success of Terraform is in the blunt simplicity of its interface. Every time you run it, it will spell out in very large writing, what exactly is it going to do
I agree with this statement 100%. The explicit, simple and clear information on what exactly is being modified is one great piece of devex. This precise workflow is something that's inspired us and we want to bake into shuttle.
Glad you like Shuttle! What do you mean exactly by support for GCP? Shuttle right now is built on AWS but that is mostly abstracted away from you. Do you want to use GCP products (say BigQuery) or do you want to host Shuttle on your own GCP project?
Why do you think it's a terrible idea?
Actually we do[0]! You can deploy a shuttle app with an annotation and a single cargo command -> `cargo shuttle deploy` :)
Thanks for this! I found a different source[0] but I'll make the update shortly!
> this feels like a terrifying way to make your infrastructure entirely implicit and remove anybody's ability to coherently reason about the stuff underneath your application.
There's two parts to this. The first is understanding what is happening under the hood since we're automatically provisioning so much for you (subdomain, LBs, DBs, etc.). Right now this is a little opaque, but a dashboard is in the roadmap to give users visibility on what's happening 'under the hood' - showing your provisioned infrastructure as well as how it's all wired up.
The second point is about potentially destructive actions - here we're going to be following a 'terraformesque' philosophy where infrastructure diffs are presented to the user and need to be accepted explicitly when deploying (or via a `--auto-approve` flag).
So I hacked together a little 'distraction free' landing page in 30 mins and it's worked like a charm. I thought someone else might find this useful.
edit: Thanks for pointing this out. Removed from Show HN.
We spent quite a bit of time on performance over the past month. We called out to a Python runtime to generate semantic data types (names, addresses etc.). As of 0.5.0 ripped it out and replaced it with a pure Rust semantic data generator called fake-rs.
We've been quite frustrated with the status quo of test data generation - after speaking to tons of other devs we've realised that many people are struggling when it comes to generating realistic looking test data.
Also, where people don’t want to copy sensitive production data to testing environments, data obfuscation can be a huge time-sink.
Enter Synth: a declarative data generator (see our website: https://getsynth.com/, github: https://github.com/getsynth/synth)
Synth enables devs and dev teams to have their application data models as code (basically a hierarchy of files) in their repos. These files can then be used to generate data for a local dev environment, automated testing in CI or even for sharing across organisations. The parameters of generation can also be tweaked to push the data model to its limits for QA, and even scaled for load testing / performance testing.
We're now working on taking the next step, and building a DSL around Synth. The Synth DSL will enable users to concisely define what data should look like and get going.
We're open source and written 100% in Rust. We believe that by making test data be as easy as using production data, we can improve the security and privacy for all of us. We'd love to get more early users as the initial feedback is positive but limited.
Thank you and looking forward to any feedback / ideas about how we can build a better tool for you!
P.S. Synth launched on HN a while back (https://news.ycombinator.com/item?id=24198114) as an ML solution to create realistic (and safe) copies of your sensitive production data as a service. This approach quickly hit several limitations which couldn't address the use-cases we are trying to solve, happy to go into more details on this if anyone is interested.
I love Terraform and have used it for years (before 0.12 I think). The workflow, meaningful diffs and reproducible 'infrastructure-as-code' gave a user experience that really was a massive step up to what I was used to (basically cloud console and scripts in CI).
In fact the Terraform workflow / philosophy inspired some of the design of an OSS 'data-as-code' tool (https://www.getsynth.com/) that we're building a company around. We wanted to use HCL instead of JSON for our config to start off with, but the Rust HCL parsers when we started the project weren't really robust so we settled.
Anyway, congratulations Hashicorp!
If you need to create test data with complex business logic, referential integrity and constraints we've been working on declarative data generator that is build exactly for this: https://github.com/openquery-io/synth.
> Anyone able to weigh in on what "at scale" means here? Just "N number of records" where N is larger than is tenable by manual curation?
Scale here refers to the complexity of the schema. (I admit perhaps the wording is misleading). If your schema is sufficiently complex, generating coherent data can be a massive PITA - Synth alleviates that pain :).
> What about generating realistic amounts of data for performance testing (millions of records), and pushing that to shared database servers, or making it available for developers to populate their local databases? Are there best practices for this?
Good question. As Synth currently stands, theoretically it can do millions of rows but it would take a while as it's making calls to a Python runtime to piggy-back off the Faker lib.
I think I'll make a follow up post about this as it is quite an interesting question.
But honestly, kudos to the YC team. At a time when remote working was not as established as it is today - things were still awkward as people were not necessarily used to socialising over Zoom - the array of tools which have been developed to make remote working easier during the pandemic were not as refined and so on - the YC team delivered a really excellent experience which was very fulfilling.
We had loads of time with the partners - we met a bunch of other founders (even became friends with some). Every event that I attended was executed perfectly, with almost 0 scheduling or connection issues.
The only reality (not much could be done about this) which was a bit awkward is juggling time zones. We are based in Europe so all events were deep in the PM.
(Also not paying an exorbitant rent for 3 months definitely helped our runway)
Yes we've seen this quite a lot in the wild. The truth is this is not very well defined - how do you get your data to tell a story depends on the story you are trying to tell.
We are trying to come up with a more rigorous framework for abstract representations of 'scenarios' . It's on our roadmap so keep an eye out for this :)
What I can say is that these were alternative[0] datasets.
[0] https://en.wikipedia.org/wiki/Alternative_data_(finance)
From what we have seen, Delphix is very much focused exclusively on large enterprise, and by extension does not look like a tool which is focused on the developer experience (could be wrong here).
We are much more focused at addressing the engineers in businesses - at the end of the day it's developers who will be using this tooling.
The gist of it is that if the the original data has a spike around -71, you will indeed see a spike in the synthetic copy as well. What it boils down to under the hood is a decision on the value of a continuous degree of freedom between two pieces of information:
- the information that you have a significant number of users located in Boston, and
- the information that any given particular user is located in Boston.
At a high level, we are taking the view that for your synthetic data to be realistic, it would need to spike around Boston if and only if most of your users are in Boston. This also means that you are not leaking information about any given individual user and that the behavior of the crowd is OK. Put more simply, if you have a single user located in Boston and all the others in, say, San Francisco, then your synthetic data should not end up having users in Boston at all.
Currently we do not have any bespoke support for lat/lon data, beyond it being like any other float of course. It is planned for the next release though! So check back in a couple weeks and it'll be there
Feel free to get in touch if you would like to discuss more about your use case :)
Yes - this is a WIP. Thanks for pointing it out
> `synth model inspect` output does not match the docs - how do I see the JSON?
Ah yes this is a typo in the docs. We'll fix it up. What you're looking for is: `synth --format json model inspect <model-id> | jq`
Thanks for the feedback!
Why is this important for you?
It turns out that picking a name for a startup/product which is representative of what you do is hard!
Hope this helps :)