Some feedback: It would be great to see in one place all the supported fields and values for the YAML config.
318 karma · joined November 11, 2012
Some feedback: It would be great to see in one place all the supported fields and values for the YAML config.
Is it the same logic for the VectorStoreServer? https://pathway.com/developers/user-guide/llm-xpack/vectorst...
{"doc": "Using Large Language Models in Pathway is simple: just call the functions from `pathway.stdlib.ml.nlp`!"}
What if I pass two contradictory statements? Is there a way to remove (or better update) a document with a new version?
For example, if I am ingesting some public docs, and I update a doc page. How do I make so that it only takes the answer from the latest document version?
> Then it processes and organizes these documents by building a 'vector index' using the Pathway package.
What is the Pathway package?
You can then build custom connectors on top of these and many users actually need to modify an existing connector, but would rather start from a template than from scratch.
Airbyte also provides a CLI and YAML configuration language that you can use to declare sources, destinations and connections without the UI: https://github.com/airbytehq/airbyte/blob/master/octavia-cli...
I agree with you that code is here to stay and power users need to see the code and modify it. That's why Airbyte code is open-source.
Airflow also started as an orchestrator, but then they tried to cover all sorts of other use cases like ETL/ELT pipelines with transfer and transformation operators...
I feel like Prefect focuses on doing one thing, orchestration, and then integrates with other data tools.
I guess one benefit is that you can use Airbyte for all your data syncs, CDC and non-CDC. You can give it a try with your own data, and see if it's easier for your team. You can run Airbyte locally with Docker Compose: https://docs.airbyte.io/quickstart/deploy-airbyte
I still think that Airflow is great for power data engineers. Airbyte and dbt are positioned to empower data analysts (or lazy data engineers like me) to own the ELTs.
To deliver that you need to centralize data on their the Materialize database, which is may main caveat. With dbt you can use any data warehouse.
You know there is this "data mesh" hype now. I think the idea behind is to empower data consumers within the company (data analysts) who know best the data to create and maintain the models. That's easier said than done, and most often turns out into a worst situation than when is only data engineers who can model data... I've only heard of Zalando who has successfully distributed data ownership within the company.
So, you would normally first load all data to the data warehouse. Then dependencies between SQL models are easier to map.
Another concern about using Airflow for the T part is that you need to code the dependencies between models both in your SQL files and your Airflow DAG. Other open-source projects like dbt create a DAG from the model dependencies in the SQL files.
So I advocate for integrating Airflow scheduler with Airbyte and dbt.
Curious to know how other use Airflow for ETL/ELT pipelines?
I believe that all tech teams have tremendous value to share. With Guriosity, I want to encourage and support more teams to write about their work.
So, I gathered 60 software engineering blogs by French companies and classified 600+ manually picked articles in 10 categories: Backend, Data, Frontend, DevOps, Product...
For candidates, it can be a great window to know how is it going to be working for a company before joining.
It is not really based on your interests, it just takes your text and suggests subreddits where people have posted similar texts.
Otherwise, I agree on what you say. I would love to also see those kind of systems. Kind of what you get as a reaction when you talk with a mentor that surprises you ;)
Unfortunately, my skills are not there yet but I am working hard to eventually be able to build those "surprise/discovery" systems.
On the other hand, the machine is faster and lot of people don't get an answer there or can wait for it. The machine is not necessarily better, just a complement.
The list of subreddits and an estimation of the performance for each one is on this Google Spreadsheet
https://docs.google.com/spreadsheets/d/1NBY1o85ZiNpcm4tcYhKk...
I will probably retrain it on more subreddits, and fine tune a few things.