162 karma · joined May 24, 2011
You're welcome to make the app id opaque. Also feel free to file an issue for making a setting to disable the app id - definitely not opposed to adding it as an option given interest.
-Fred
Eng @ Streak
Once we got to a mode where we were only doing computation per request, the raw cost is pretty minimal since there's not much CPU time either way.
More details from when I was working through it: https://twitter.com/fredwulff/status/1204861220165017600
Streak is hiring our first dedicated engineering manager who will be directly responsible for some or all of our engineering team (currently ~15 engineers, distributed between product, infrastructure, and mobile).
* Problem: Make Gmail powerful for all businesses
* Product: We build a sales/hiring/fundraising/dealflow tool all inside Gmail. We believe these workflows belong entirely in your inbox because thats where people spend their entire day.
* Traction: Product market fit, hundreds of thousands of users, tens of thousands of paying users
* Funding: $2M seed, profitable and growing ever since
* Stack: Java, Kotlin, Golang, React, all the modern JS tooling - built on GCP, largest user of Google Cloud Spanner
Interested? Visit and apply at https://www.streak.com/teams/engineeringWill put together a follow-up with some more information, but in summary: * yeah, no built-ins for graph operations in Spanner other than relational SQL for standard joins * the key thing that make this work are the ability to easily construct global indexes that aren't sharded by the primary key and reasonably fast joins between them * it's also helpful that Spanner does a reasonable job of parallelizing queries (e.g. a lot of times we'll get a 15x increase in speed vs. a sequential plan) * we then do the fan-out across the graph in our Java Spanner client - each distributed SQL index read takes ~10 ms so we can do multiple round trips of graph traversal in the client
The intention wasn't to hold out Spanner as a purpose-built graph database, but rather to talk about using it for a graph-y database use case. Will put together a follow-up with some more information, but in summary: * yeah, no built-ins other than relational SQL for graphs * the key thing that make this work are the ability to easily construct global indexes that aren't sharded by the primary key and reasonably fast joins between them * it's also helpful that Spanner does a reasonable job of parallelizing queries (e.g. a lot of times we'll get a 15x increase in speed vs. a sequential plan) * we then do the fan-out across the graph in our Java client
Definitely wasn't meaning this post's title to be clickbait-y. The intention wasn't to hold out Spanner as a purpose-built graph database, but rather to talk about using it for a graph-y database use case. Will put together a follow-up with some more information, but in summary: * yeah, no built-ins other than relational SQL for graphs * the key thing that make this work are the ability to easily construct global indexes that aren't sharded by the primary key and reasonably fast joins between them * it's also helpful that Spanner does a reasonable job of parallelizing queries (e.g. a lot of times we'll get a 15x increase in speed vs. a sequential plan) * we then do the fan-out across the graph in our Java client
Redshift: Pros: Has the most adoption, so most integrations from SaaS services etc. are built with Redshift as their sink. Relatively fast and battle-tested.
Cons: In an awkward middle ground where you're responsible for a lot of operations (e.g. capacity planning, setting up indexes), but don't have a lot of visibility. Some weirdness as a result of taking PostgreSQL and making it distributed.
BigQuery: Pros: Rich feature set. Pay-per-TB pricing. Recently released standard-ish SQL dialect. Very fast.
Cons: JDBC driver is recent and doesn't have support for e.g. CREATE TABLE AS SELECT (as of a couple of months ago) so harder to integrate with existing systems. There are ways to run out of resources (e.g. large ORDER BY results) without a good path to throw more money at the problem.
Athena: Pros: Built off of the open-source Presto database so can use the documentation there. Pay-per-TB pricing.
Cons: Slower than the other options listed here. Very early product so lacking some in documentation and some cryptic errors. Not a lot of extensibility, but you could theoretically move to just using open-source Presto.
Haven't had a chance to evaluate Snowflake.
(I work for Dropbox but don't speak for it in this capacity)
One question: on the throughput graphs, I understand why you normalized them per graph, but were there any differences between graphs (particularly in terms of EBS vs. ephemeral) that would be sufficient to drown out the variability within the throughput graphs?
We looked into a lot of logging solutions and couldn't find a solution that fit our requirements. In particular, 1) at the moment we don't run any servers (even virtualized) ourselves, and we're loving not having to worry about that level of infrastructure. Splunk and most other logging solutions would have required running a VM just for the daemon (either proprietary or syslog). 2) the total cost of developing Mache was significantly less than even a single 500 MB Splunk license. 3) most other solutions that we looked at really want to treat logs as lines of text data. App Engine logs provide more metadata, and we didn't want to lose and have to reconstruct that data.
Unfortunately, there's not really a good, widely used license that does what you want, mostly because you run into problems pretty quickly based on derivative works with what you want. Let's say somebody merges your email client with a browser - can they call that by a different name? Can they sell that?
Anyhow, I'd suggest looking at looking at the Apache or BSD licenses (if you want the broadest use) or the GPL (if you want to ensure that modifications to the code must be distributed with any binaries made from the code). Licenses are a bit of a pain, but you can just pick one of the common ones and it'll really help adoption.
We've found, both in talking to language teachers and in our own experience, that video helps communication and comprehension. However, it's definitely the case that for both privacy and bandwidth reasons, sometimes it's a liability. We don't require a webcam so for now you can just disconnect it if you're uncomfortable with it, but we're planning to add an option to intentionally omit the webcam feed in the future.
Thanks for the idea on topics; we've got a lot of ideas in that space, but we wanted to launch as soon as we had a viable product.
Glad you liked the call quality, it's central to many of the features we're hoping to introduce in the future. There's lots more coming. This is just the first step. :)