HNHacker News
TopNewBestAskShowJobs

kn7

47 karma · joined December 3, 2010

submissionscomments
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
Please see my reply to "someone else" you mentioned. If there is anything else you think I might have mistaken to employ during benchmarks, I am all ears.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
We store the real-time content stream in a separate bulk storage unit (e.g., BigQuery) with a certain retention window, but the ETL'ed documents are always on ES. Given a plain event (i.e., not ETL'ed document) is not much of a value for search, I would not call the stream storage as the primary storage. It just assists us to re-build the ETL state in case of an emergency.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
See my response above, it was indeed a typo from my side. I am sorry to hear that it spoiled the rest of the post for you.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
It was indeed an incorrect snippet -- removed it. We do have a group of PostgreSQL experts within the company and we let them tune the database for the benchmark. Let me remind, this was not a tune once, run once operation. We spent close to a month to make sure that we are not missing anything obvious for each storage engine. But as I explained, nothing much worked in case of PostgreSQL.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
We generally use index refresh in ITs (running ES in Docker) and it fails occasionally, which I believe the case described here: "The (near) real-time capabilities depend on the index engine used." https://www.elastic.co/guide/en/elasticsearch/reference/curr...
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
That still doesn't guarantee that a consecutive read is gonna get the last state. Welcome to the wonderful world of "eventual consistency".
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
That still doesn't guarantee that a consecutive read is gonna get the last state.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
Bol has plenty of other ETL pipelines for BI. What I meant is the data cooked for search is not (much) of interest to BI, yet. Though we do have other means to feed BI for search-relevant content.
kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
Hey Debarsh! First, thanks for taking time to read such a lengthy post.

If I am not mistaken the majority of the PL/SQL glue is owned by Gert, though you might recall better. Quite some VCS history was lost while migrating from SVN to Git. ;-)

The reason we are "replicating" the entire data is to 1) determine the affected products and 2) re-execute the relevant configurations (facets, synonyms, etc.) while making retroactive changes. (For instance, say someone has changed the PL/SQL of "leeftijd" facet.) Here, the storage is required to allow querying on every field, for (1), and on id, for (2). While id-based bulk querying is (almost) supported by every ETL source, querying on every field is not. Hence, we "replicate" the sources on our side to suffice these needs. Actually, the entire point of the post was to explain this problem, but apparently it was not clear enough.

For your remarks on event sourcing and BI, I am a little bit puzzled. I will need some elaboration on these remarks. We do have event sourcing on our side (that is how we can replay in case of need) and BI is not really interested in ETL data. Maybe I misunderstood you?

I am also confused by how you relate scheduling/running PL/SQL jobs via Hadoop, Spark, Flink, etc. Did you see the link to Redwood Explorer I shared in the post?

kn7··on Using Elasticsearch as the Primary Data Store in an ETL Pipeline
That is indeed the case and we are also bitten by that. The most effective work around we managed to find is to flush periodically (say every 300ms) for a certain timeout period before reading from ES again for checks. Though even then, ITs still fail time to time.
kn7··on Confirmed: Microsoft Will Announce Acquisition of Skype Tomorrow Morning
Skype never properly worked in Linux, now we also lost the hope -- if there is any -- that it will in the future.