997 karma · joined April 5, 2013
If I click the associated PDF document, I can see there is legible text, so I think it might be an issue with the ingestion of the document itself.
EDIT: Context - https://www.wsj.com/podcasts/the-journal/people-of-the-state...
I wish they give support to Firefox in the future, getting text (and even modifying it!) is something I need to do often.
The way of backing up to people who already left their jobs was very unprofessional.
The risk I see with that is that you have to be extra careful when processing updates to the metastore, as sometimes there are competing processes writing metadata about the same partition and if not done properly you'll end up with a corruption that might cause problems (ie. you skip a whole file because the metadata was incorrect, causing a correctness problem down the road).
yes! that's something we are trying to do, since we have a way to create signatures for SQL statements and subqueries are just SQL statements then we can get all the "signatures" a query use and compare if other queries are using the same signatures. Then just sort those queries by number of times used and put some other perf metrics like IO/CPU needed to compute it and you get a good starting point.
Microsoft did something similar with Azure, using bipartite graphs, their solution was more advanced as they also baked in constraints like "the materialized view can't be more than X GB in size" but the end result is the same. (https://www.microsoft.com/en-us/research/uploads/prod/2018/0...)
We also created some other datasets that tell you how the tables are normally joined and another team created a ML model that now we use in an internal tool that automatically recommends the best join keys when you join 2 or more tables (since we have lots of historical info on that already stored).
We also had to create a very efficient table profiler to evaluate the candidates, because in some cases there are columns that are widely used in equi-where conditions, but their cardinality is very high, making them bad partition columns.
One guy created a parser that actually gets the most common values used to filter each column, I guess we could use that in the future to materialize some views; my original vision of the project was to continue with partial aggregations for common computations. The thing is there are several pieces of code that are pretty much copy/pasted and reused in many pipelines, so why not materialize those and rewrite the SQL of the subsequent pipelines to leverage the materialized version? huge savings there.
We are planning to present it in VLDB or a similar forum, there are some aspects of it that we need to 'clean' if we want to open source it.
Other parts of the system include the candidate evaluation and the module that computes the expected savings for the best candidate selected during evaluation; this system in particular has a lot of specific Presto and Spark logic that might need to get more general if we want to open source it.
Since some of the tables are multi-petabyte with thousand of downstream consumers, the savings have been in the millions of USD, because of CPU savings mostly, but also there is the benefit of improved wall time and data arriving much earlier to dashboards and reports.
- How do you deal with standardization across the different events you log? How do you handle standardized naming, representations (formatting) and types?
- What about validations? do you validate the payloads in any way? if so, at what layer? Client before logging? In Kafka?
- Any privacy capabilities? How do you make sure the data is not accessed by processes/people that's not supposed to access a datum?
Thanks!