HNHacker News
TopNewBestAskShowJobs

mona_rakibe

7 karma · joined February 23, 2021

submissionscomments
mona_rakibe··on Free Data Health Analysis for YC Companies (Data Lakes & Warehouses)
Great opportunity to get a free-of-cost data health assessment.
mona_rakibe··on [dead]
There is a lot spoken and written about Data Observability, Data Lineage, or even Lineage as a pillar of observability, and yet there is still a lot of confusion. So, What is Data Lineage? Why is it so crucial for Data Quality? What is Data Observability? Why is Data Lineage important for Data Observability? How is data lineage used in Data Observability? and How should we use Lineage in Observability?
mona_rakibe··on [dead]
There is a lot spoken and written about Data Observability, Data Lineage, or even Lineage as a pillar of observability, and yet there is still a lot of confusion. So, What is Data Lineage? Why is it so crucial for Data Quality? What is Data Observability? Why is Data Lineage important for Data Observability? How is data lineage used in Data Observability? and How should we use Lineage in Observability?
mona_rakibe··on [dead]
As databases were evolving and migrating to the cloud, the data models were changing too. Most modern cloud data warehouses or data lakes now support semi-structured schemas. Today pretty much any big data platform supports working with NDJSON, i.e., New Line Delimited File, where each row is a proper JSON representing a record. And even native support of JSON schema type: [https://cloud.google.com/bigquery/docs/reference/standard-sq...]

This brought a great opportunity for data architects to design the most efficient data model for storage and querying. However, at the same time, it created a challenge for Data Observability. And the reason is multi-value fields, aka arrays. Such data requires special logic to properly calculate basic data quality KPIs like completeness, uniqueness etc. as there can be some records with no values vs multiple values per record. This blog highlights how we address monitoring for semi-structured data

mona_rakibe··on [dead]
Get alerted by Telmai, instead of being alarmed by the Business team!

We are re-launching our complete data observability product.

We are excited to re-launch Telmai as full data observability product to this incredible community!

What is Telmai? A data observability product that helps data teams proactively detect and investigate data issues closest to their source before having a downstream impact.

How does Telmai work ? Watch this short 3 minutes video.

What needs do we serve? Data teams are struggling today to detect data issues as soon as they occur — these issues are often first experienced by data consumers and cause long and tedious investigation cycles.

With today's fragmented pipelines, every step in the pipeline can become a point of failure with myriad issues such as incomplete data, data type changed by the source system, a mapping error, regex and pattern changes, microservice failure, etc.

Telmai can proactively detect and alert on such issues at its source, and once alerted, finding the root cause is quick using the interactive investigator UI.

What are some core features of Telmai? Automatic monitoring on 40+ predefined data metrics like schema change, row count, completeness, uniqueness, patterns, distribution change, accuracy, etc Simple UI for expectations/rules - no more hand coded rules using Great Expectation or DBT expectations. Monitoring for streaming and batch data sources: AzureBlob, GCS, S3, Snowflake, BigQuery, Firebolt, Redshift, Pubsub, Kafta & kinesis. Support for semi-structured data i.e., nested and multi-valued attributes. Support multiple formats like JSON, Parquet, CSV, and Avro Intuitive human-in-loop model for fine-tuning thresholds, policies, or writing expectations Spark processing to enable infinite scale without performance impact Private cloud and SaaS offering What are the benefits of using Telmai? Simple, no-code on-boarding Reduce escalation on data issues Faster detection and resolution of issues High-security design

How to get stated ? Create a free Telmai account https://www.telm.ai/get-started Within minutes you will get your account with on-boarding instructions (Thanks @Arengu for helping us streamline this process) Connect to your datasource

You are all-set, if these steps more than 1 hours take extra 10% off ;-)

mona_rakibe··on Launch HN: Secoda (YC S21) – Searchable Company Data
Very nice ! congrats - pretty awesome tool
mona_rakibe··on Who's responsible for the Data and Data quality?
Data Quality needs to be democratized across business and IT? yes or no?
mona_rakibe··on Launch HN: Exams, tasks, K8, eCommerce, cell sites, health, data quality, travel
Thank you so much for this feedback. Regarding the pricing page we will tighten it and also add calculator, please stay tuned. For now the way it works is after first 500K values we charge $150 per 1M attribute values/month.

We dont monitor server or any infrastructure, we are designed for data quality monitoring we can flag issues like missing data, volume drifts, schema mismatch and where we stand out is monitoring actual accuracy of data at record value level. Example : We can flag anomalous titles, emails , overrepresented phone numbers etc

mona_rakibe··on Launch HN: Exams, tasks, K8, eCommerce, cell sites, health, data quality, travel
Telmai (https://www.telm.ai/) is a real-time data quality monitoring platform that can automatically detect and investigate data quality issues as data is getting ingested. Our tool uses a statistical and ML engine that helps data product owners understand data anomalies and intuitively define correct versus incorrect data. These definitions are then used to proactively monitor and alert on data quality problems.

We have decades of experience with enterprise data and find that this approach towards data quality addresses a huge gap in data platforms. Detecting and investigating data quality issues is extremely tedious, time-consuming and expensive. Using Telmai, companies like Dun & Bradstreet and Myers-Holum are able to find and resolve such issues across millions of records in minutes. Ask us anything!