If you decide to copy not only ideas but a big part of internal implementation, I think you should include that information in your LICENSE.
Cheers
If you decide to copy not only ideas but a big part of internal implementation, I think you should include that information in your LICENSE.
Cheers
Looks like much of the lineage code is also largely a wrapper around this library: https://github.com/reata/sqllineage
Would be curious to understand the project's purpose and unique contributions vs. the underlying dependencies powering it as there seems to be some ambiguity. Is this just a wrapper around dbt transformations and a lineage library in one package? Can I just use them directly?
The lineage part is independent (re_data uses lineage from dbt), so I haven't looked into that much.
In terms of the lineage, you can see in the code that we mostly rely on query and access history that exist in Snowflake and Bigquery to parse the queries and learn about the connection between nodes in the graph. We use other python libraries like sqlfluff and sqllineage as low level parsers for some specific use cases which we extend and solve many things on top of them. Actually we're heavy open source users, depending on around 20 libraries, all MIT or Apache.
Also, I think it's more than just following up re_data in a couple of places. Elementary's whole data monitoring part started much later than your Lineage part, and it seems to try to follow what re_data did there on the idea & implementation level. I'm sure the other 59 projects you mention were not dbt packages for data reliability (there were no other one in the dbt hub) which is what re_data is and now elementary also tries to copy this. (seeing our traction)
As mentioned, it's open-source. You can use our code. But if you are doing that, state that clearly in the LICENSE.
> Copyright [yyyy] [name of copyright owner]
https://github.com/elementary-data/elementary/blob/master/LI...
Elementary models like data_monitors_thread1, data_monitors_thread2, data_monitors_thread3, data_monitors_thread4, data_monitoring_metrics, latest_metrics, metrics_stats_for_anomalies, z_score, anomaly_detection, schema_schenages, etc. Existed before in re_data, are doing the same things and specifically for *_thread4 are not similar to anything you normally do in dbt.
And these similarities are also visible in code, for example here: the same usage of the undocumented dbt context feature.
# elementary
{% macro get_monitor_macro(monitor) %}
{%- set macro_name = monitor + '_monitor' -%}
{%- if context['elementary'].get(macro_name) -%}
{%- set monitor_macro = context['elementary'][macro_name] -%}
{%- else -%}
{%- set monitor_macro = context['elementary']['no_monitor'] -%}
{%- endif -%}
{{- return(monitor_macro) -}}
{% endmacro %}# re_data
{%- macro get_metric_macro(metric_name) %}
{% set macro_name = 're_data_metric' + '_' + metric_name %}
{% if context['re_data'].get(macro_name) %}
{% set metric_macro = context['re_data'][macro_name] %}
{%- else %}
{% set metric_macro = context[project_name][macro_name] %}
{% endif %}
{{ return (metric_macro) }}
{% endmacro %}And it's open-source so it's generally okay to do that, but it should be reflected in the LICENSE.