163 karma · joined July 24, 2016
We are building the EvaDB database system for AI apps -- https://github.com/georgia-tech-db/eva.
Email: arulraj@gatech.edu Website: https://faculty.cc.gatech.edu/~jarulraj/
1. Abstraction Lifting (Al) + Policy/Mechanism Separation (Pm): SQL states high-level intent with precise semantics, and logical operators are decoupled from physical operators.
2. Equivalence-based Planning (Ep) + Invariant-Guided Transformation (Ig): We apply algebraic rewrites that preserve semantics (e.g., join reordering, predicate pushdown) under stated invariants.
3. Cost-based Planning (Cm): We choose concrete physical operators and join orders using a cost model and so on..
To clarify: this is indeed just a taxonomy of classic system-design principles. The periodic-table styling is a familiar metaphor; there is no claim that principles repeat periodically. The goal was to outline a mostly orthogonal set of design principles and highlight cross-domain connections across computer systems so it is easier to discuss designs precisely. Thanks for all the thoughtful feedback!
What are your thoughts on reducing LLM cost?
We are also exploring LLM-based data wrangling using EvaDB and cost is an important concern [1, 2, 3].
[1] https://github.com/georgia-tech-db/evadb
[2] https://medium.com/evadb-blog/stargazers-reloaded-llm-powere...
What were the interesting problems you faced in processing the survey data?
If possible, can you share the prompt?
Are you measuring accuracy with data wrangling prompts? Would love to learn more about that.
It would be great if you could share an example of the inconsistent output problem -- we also faced it. GPT-4 was much better than GPT-3.5 in output quality.
The results were pretty good: https://gist.github.com/gaurav274/506337fa51f4df192de78d1280...
Another interesting aspect was the money spent on LLMs. We could have directly used GPT-4 to generate the "golden" table; however, it's a bit expensive — costing $60 to process the information of 1000 users. To maintain accuracy while reducing costs significantly, we set up an LLM model cascade in the EvaDB query, running GPT-3.5 before GPT-4, leading to a 11x cost reduction ($5.5).
Query 1: https://github.com/pchunduri6/stargazers-reloaded/blob/228e8...
Query 2: https://github.com/pchunduri6/stargazers-reloaded/blob/228e8...
We found some interesting insights. In the Langchain community, ~40% of the stargazers are from India. In the GPT4All community, we found that web developers love open-source LLMs -- more so that machine learning folks :)
Curious if there is any reason why you would not use it?
We went with the local mode as several Python AI apps are using Qdrant in that mode based on the suggestion here: https://qdrant.tech/documentation/quick-start/.
We also believe in open-sourced benchmark code. Please find the code here: https://github.com/jiashenC/vectordb-benchmark-and-optimize/....
1. What feature extractor is used to derive code embeddings?
2. Would support for more complex queries be useful inside the app?
--- Retrieve a subset of code snippets
SELECT name
FROM snippets
WHERE file_name LIKE "%py" AND author_name LIKE "John%"
ORDER BY
Similarity(
CodeFeatureExtractor(Open(query)),
CodeFeatureExtractor(data)
)
LIMIT 5;The YOLO example on your Github page is super interesting. We are finding it easier to get LLMs to write functions with a more constrained function interface in EvaDB. Here is an example of an YOLO function in EvaDB: https://github.com/georgia-tech-db/evadb/blob/staging/evadb/....
Once the function is loaded, it can be used in queries in this way:
SELECT id, Yolo(data)
FROM ObjectDetectionVideos
WHERE id < 20
LIMIT 5;
SELECT id
FROM ObjectDetectionVideos
WHERE ['pedestrian', 'car'] <@ Yolo(data).label;
Would love to hear your thoughts on ChatCraft and a more constrained function interface. --- Create a reference table that maps neighborhoods to zipcodes using ChatGPT
CREATE TABLE reference_table AS
SELECT parkname, parktype,
ChatGPT(
"Return the San Francisco neighborhood name when provided with a zipcode. The
possible neighborhoods are: {neighbourhoods_str}. The response should an item from the
provided list. Do not add any more words.",
zipcode)
FROM postgres_db.recreational_park_dataset;
--- Map Airbnb listings to park
SELECT airbnb_listing.neighbourhood
FROM postgres_db.airbnb_listing
JOIN reference_table ON airbnb_listing.neighbourhood = reference_table.response;
More details on LLM-powered joins and EvaDB: https://medium.com/evadb-blog/augmenting-postgresql-with-ai-..., https://github.com/georgia-tech-db/evadbMore details here: https://medium.com/evadb-blog/augmenting-postgresql-with-ai-...
1. Will that query look like this:
SELECT LLM("{user_question}", order_info)
FROM postgres_data.order_table
WHERE user_id = “101”;
2. How will a feature store, like Hopsworks, help in this app?Shameless self-plug: We are building EvaDB [1], a query engine for shipping fast AI-powered apps with SQL. Would love to exchange notes on such apps if you're up for it!
Doesn't this token sampling optimization require using a locally-running model like Llama?
I am presuming that OpenAI doesn't provide direct access to token probabilities in its API.
I guess this is the SQL query you have in mind that uses the LIKE operator:
SELECT ChatGPT("Respond to the review with a solution to address the reviewer's concern", review)
FROM postgres_data.review_table
WHERE ChatGPT("Is the review positive or negative?", review) LIKE "%positive%"
AND location = “waffle house”;
From a query processing standpoint, both queries should have equivalent performance -- unless we build an index over the output of the ChatGPT query in EvaDB, in which case the former query would be faster than this one.Shameless self-plug: We are building EvaDB [1], a query engine for shipping fast AI-powered apps with SQL. Here is an illustrative query for analyzing food reviews stored in Postgres and generating responses for negative reviews:
SELECT ChatGPT("Respond to the review with a solution to address the reviewer's concern", review)
FROM postgres_data.review_table
WHERE ChatGPT("Is the review positive or negative? Only reply 'positive' or 'negative'.", review) = "negative"
AND location = “waffle house”;
It would be interesting to learn about the queries needed for supporting HackYourNews application. Would love to exchange notes on this if you're up for it!