HNHacker News
TopNewBestAskShowJobs

Arimbr

318 karma · joined November 11, 2012

Data & Content Engineer. Building toolsfordata.com
submissionscomments
Arimbr··on Show HN: Build Live AI and RAG Pipelines in Minutes with YAML Templates
I like the YAML abstraction. This should make it easier to programmatically try and evaluate multiple configurations for the whole AI pipeline (not just the LLM) against a dataset or real users through an API deployment.

Some feedback: It would be great to see in one place all the supported fields and values for the YAML config.

Arimbr··on Show HN: Multimodal Search Using GPT4o for Metadata Extraction and Hybrid Index
What is an hybrid index?
Arimbr··on Show HN: Pathway – Build Mission Critical ETL and RAG in Python (NATO, F1 Used)
Nice, thanks! I was reading https://pathway.com/developers/user-guide/deployment/persist.... If I understand correctly you persist both source data and internal state, including the intermediary state of the computational graph. And you only rely on the backend to recover from failures and upgrades. So if I want to clone a Pathway instance, I don't need to reprocess all source data, I can recover the intermediary state from the snapshot.

Is it the same logic for the VectorStoreServer? https://pathway.com/developers/user-guide/llm-xpack/vectorst...

Arimbr··on Show HN: Pathway – Build Mission Critical ETL and RAG in Python (NATO, F1 Used)
If all the pipeline and the vector index is keep in memory... does Pathway still persist state somewhere?
Arimbr··on Airbyte's Spring Release with a Preview of the Connector Builder AI
The AI Connector Builder from API docs is insane! Which API doc specifications will it support? Or does it even matter?
Arimbr··on Python ETL with Airbyte and Pathway
Interesting implementation! For complex stream and text processing, I also prefer processing data in memory with Python (ETL) rather than SQL in the warehouse (ELT).
Arimbr··on Show HN: LLM App – build a realtime LLM app in 30 lines, with no vector database
I see the ingested documents in the data folder don't have an id field, only a doc field.

{"doc": "Using Large Language Models in Pathway is simple: just call the functions from `pathway.stdlib.ml.nlp`!"}

What if I pass two contradictory statements? Is there a way to remove (or better update) a document with a new version?

For example, if I am ingesting some public docs, and I update a doc page. How do I make so that it only takes the answer from the latest document version?

Arimbr··on Show HN: LLM App – build a realtime LLM app in 30 lines, with no vector database
Hi, interesting!

> Then it processes and organizes these documents by building a 'vector index' using the Pathway package.

What is the Pathway package?

Arimbr··on Data Orchestrators 101: Everything You Need to Know to Get Started
Nice list of resources!
Arimbr··on Show HN: Airbyte, data integration platform with 300+ open-source connectors
Thanks for the feedback! Recently a community member created a Terraform provider: https://github.com/eabrouwer3/terraform-provider-airbyte
Arimbr··on Show HN: Airbyte, data integration platform with 300+ open-source connectors
Know that there is also a CLI to manage configurations defined in YAML files. And a few options to deploy Airbyte in "one click". It's all in the README, sorry to hear you didn't find your way around our docs... There are a number of growing features and deployment options now.
Arimbr··on Data Engineering Trends for 2023
My bet: Data testing, data monitoring and data catalog solutions will consolidate to cover data quality all together.
Arimbr··on Whats the difference between ETL and ELT
The future is EtLT! t for data privacy transformations and T for the rest.
Arimbr··on The evolution of the data engineer role
Oh, declarative doesn't necessarily mean no-code. Airbyte data integration connectors are built with an SDK in Python, Java, and a low-code SDK that was just released...

You can then build custom connectors on top of these and many users actually need to modify an existing connector, but would rather start from a template than from scratch.

Airbyte also provides a CLI and YAML configuration language that you can use to declare sources, destinations and connections without the UI: https://github.com/airbytehq/airbyte/blob/master/octavia-cli...

I agree with you that code is here to stay and power users need to see the code and modify it. That's why Airbyte code is open-source.

Arimbr··on The Shift from Data Pipelines to Data Products
Interesting to see how modern data orchestrators seem to be adding some of the features of data catalogs and data observability tools.
Arimbr··on SQL vs Python for Data Analysis
Sorry, wrong link, and I couldn't delete the post. I reposted with the correct link to: https://airbyte.com/blog/sql-vs-python-data-analysis
Arimbr··on Python vs. SQL for data processing in 2022
Nice article! I also tend to favor SQL for simple querying and data processing with dbt, but when I need to unit test some complex logic, I prefer Python.
Arimbr··on ELT pipelines with Prefect, Airbyte and dbt
I like how Prefect is positioned as an orchestrator for the modern data stack.

Airflow also started as an orchestrator, but then they tried to cover all sorts of other use cases like ETL/ELT pipelines with transfer and transformation operators...

I feel like Prefect focuses on doing one thing, orchestration, and then integrates with other data tools.

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
Airbyte CDC is based on Debezium, but Airbyte abstracts it away and make it easier to CDC from Postgres, MySQL, MSSQL to any supported destination (included S3). Here is the doc for CDC: https://docs.airbyte.io/understanding-airbyte/cdc

I guess one benefit is that you can use Airbyte for all your data syncs, CDC and non-CDC. You can give it a try with your own data, and see if it's easier for your team. You can run Airbyte locally with Docker Compose: https://docs.airbyte.io/quickstart/deploy-airbyte

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
Nice work there! I also think that the next challenge for data teams is all this data documentation and discovery work.

I still think that Airflow is great for power data engineers. Airbyte and dbt are positioned to empower data analysts (or lazy data engineers like me) to own the ELTs.

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
Oh, you should check Materialize. I feel Materialize is like dbt but with an ingestion layer and real-time materialized views.

To deliver that you need to centralize data on their the Materialize database, which is may main caveat. With dbt you can use any data warehouse.

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
Really good points! I don't think that Airflow is necessarily a problem if your data engineering team knows how to best use Airflow Operators, Hooks and DAGs for incremental loads. But because Airflow is not an opinionated ETL/ELT tool, most often I see a lot of custom code that could be improved...

You know there is this "data mesh" hype now. I think the idea behind is to empower data consumers within the company (data analysts) who know best the data to create and maintain the models. That's easier said than done, and most often turns out into a worst situation than when is only data engineers who can model data... I've only heard of Zalando who has successfully distributed data ownership within the company.

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
My understanding of dbt is that it builds a DAG based on the interdepencies between models. The interdepencies are parsed from 'ref' functions on the SQL files. The thing with dbt is that you transform the data within a single data warehouse.

So, you would normally first load all data to the data warehouse. Then dependencies between SQL models are easier to map.

Arimbr··on ETL Pipelines with Airflow: The Good, the Bad and the Ugly
[author of the article] My main concern about using Airflow for the EL parts is that sources and destinations are highly coupled with Airflow transfer operators (e.g. PostgresToBigQueryOperator). The community needs to provide M * N operators to cover all possible transfers. Other open-source projects like Airbyte, decouple sources from destinations, so the community only needs to contribute 2 * (M + N) connectors.

Another concern about using Airflow for the T part is that you need to code the dependencies between models both in your SQL files and your Airflow DAG. Other open-source projects like dbt create a DAG from the model dependencies in the SQL files.

So I advocate for integrating Airflow scheduler with Airbyte and dbt.

Curious to know how other use Airflow for ETL/ELT pipelines?

Arimbr··on Show HN: A collection of 600 tech articles by developers at French companies
Hello HN! My name is Ari. I've worked as a data engineer for three French startups. The first, didn't have an engineering blog. The second had one, but i never took the courage to write. The third, didn't consider it a priority.

I believe that all tech teams have tremendous value to share. With Guriosity, I want to encourage and support more teams to write about their work.

So, I gathered 60 software engineering blogs by French companies and classified 600+ manually picked articles in 10 categories: Backend, Data, Frontend, DevOps, Product...

For candidates, it can be a great window to know how is it going to be working for a company before joining.

Arimbr··on Show HN: I built a cron job scheduler
Nice, I also think that backend engineering needs more innovation!
Arimbr··on Show HN: Subreddit Finder - Trained on 4M Reddit Posts from 4K Subreddits
Thanks for your valuable feedback. This indeed answers a question such us "What community is used to hear what you have to say?".

It is not really based on your interests, it just takes your text and suggests subreddits where people have posted similar texts.

Otherwise, I agree on what you say. I would love to also see those kind of systems. Kind of what you get as a reaction when you talk with a mentor that surprises you ;)

Unfortunately, my skills are not there yet but I am working hard to eventually be able to build those "surprise/discovery" systems.

Arimbr··on Show HN: Subreddit Finder - Trained on 4M Reddit Posts from 4K Subreddits
Yes, that is the human solution to the problem. I manually tested a bit on what people ask there :)

On the other hand, the machine is faster and lot of people don't get an answer there or can wait for it. The machine is not necessarily better, just a complement.

Arimbr··on Show HN: Subreddit Finder - Trained on 4M Reddit Posts from 4K Subreddits
Yes, unfortunately is not part of the 4k subreddits I trained on. I will retrain it with more subreddits :)

The list of subreddits and an estimation of the performance for each one is on this Google Spreadsheet

https://docs.google.com/spreadsheets/d/1NBY1o85ZiNpcm4tcYhKk...

Arimbr··on Show HN: Subreddit Finder - Trained on 4M Reddit Posts from 4K Subreddits
Nice work @exegete! The Chrome extension idea is great ;) Nice to also see some metrics comparison. I will review your work. Looks like you achieved 0.6 accuracy with 600 classes. I got 0.4 f1-score on 4000 classes, but I have a ton of posts and subreddits with images and no text ;) For this case, it is also nice to report Recall@k. The current model has Recall@5 of 0.6. Meaning that on the test dataset, 60% of the time the human choice is within the first 5 suggestions. Currently, it is not supposed to automatically post but help the user discover new subreddits :)

I will probably retrain it on more subreddits, and fine tune a few things.

Page 1 of 2Next →