HNHacker News
TopNewBestAskShowJobs

mateuszklimek

13 karma · joined August 10, 2021

Data Engineer, wrote https://github.com/re-data/re-data, juggler
submissionscomments
mateuszklimek··on Launch HN: Elementary (YC W22) – Open-source data observability
Okay, I'm happy that you admit inspiration in this comment (in opposition to the previously deleted one).

Also, I think it's more than just following up re_data in a couple of places. Elementary's whole data monitoring part started much later than your Lineage part, and it seems to try to follow what re_data did there on the idea & implementation level. I'm sure the other 59 projects you mention were not dbt packages for data reliability (there were no other one in the dbt hub) which is what re_data is and now elementary also tries to copy this. (seeing our traction)

As mentioned, it's open-source. You can use our code. But if you are doing that, state that clearly in the LICENSE.

mateuszklimek··on Launch HN: Elementary (YC W22) – Open-source data observability
Sure, please compare: https://re-data.github.io/dbt-re-data/#!/overview?g_v=1 and https://docs.elementary-data.com/ graph png.

Elementary models like data_monitors_thread1, data_monitors_thread2, data_monitors_thread3, data_monitors_thread4, data_monitoring_metrics, latest_metrics, metrics_stats_for_anomalies, z_score, anomaly_detection, schema_schenages, etc. Existed before in re_data, are doing the same things and specifically for *_thread4 are not similar to anything you normally do in dbt.

And these similarities are also visible in code, for example here: the same usage of the undocumented dbt context feature.

# elementary

{% macro get_monitor_macro(monitor) %}

    {%- set macro_name = monitor + '_monitor' -%}
    {%- if context['elementary'].get(macro_name) -%}
        {%- set monitor_macro = context['elementary'][macro_name] -%}
    {%- else -%}
        {%- set monitor_macro = context['elementary']['no_monitor'] -%}
    {%- endif -%}

    {{- return(monitor_macro) -}}
{% endmacro %}

# re_data

{%- macro get_metric_macro(metric_name) %}

    {% set macro_name = 're_data_metric' + '_' + metric_name %}

    {% if context['re_data'].get(macro_name) %}
        {% set metric_macro = context['re_data'][macro_name] %}
    {%- else %}
        {% set metric_macro = context[project_name][macro_name] %}
    {% endif %}

    {{ return (metric_macro) }}
{% endmacro %}
mateuszklimek··on Launch HN: Elementary (YC W22) – Open-source data observability
It is! But it doesn't have Monte Carlo code in it :)

And it's open-source so it's generally okay to do that, but it should be reflected in the LICENSE.

mateuszklimek··on Launch HN: Elementary (YC W22) – Open-source data observability
It's "inspired" the dbt transformation part by using the same models and logic/part of code of generating them. We, for example, had a funny thing of computing metrics in 4 threads via multiple dbt models, and this is also done in elementary in a very similar way :)

The lineage part is independent (re_data uses lineage from dbt), so I haven't looked into that much.

mateuszklimek··on Launch HN: Elementary (YC W22) – Open-source data observability
Hey Maayan and Or, Nice project, at re_data we just got over a lot of your new updates and it seems a quite large part of your project is "inspired" by code from our library https://github.com/re-data/re-data. Even with parts, we are not especially proud of ;)

If you decide to copy not only ideas but a big part of internal implementation, I think you should include that information in your LICENSE.

Cheers