Pandas AI
medium.com
medium.com
In general though, this is a good idea, definitely some obstacles to overcome though - and I suspect many of them involve solving problems that have to be solved on the OpenAI side, not on the user side. For example "send my entire data frame to OpenAI" seems a bit like writing a blank cheque, at the least they could mitigate that on the user side by having a pandas_ai.estimate which could take a run command and analyze the tokens to come up with a cost estimate. The other thing is obviously you want to cache the dataframe on the OpenAI side but that's definitely going to be tricky for these devs.
I wonder if the authors have checked if the data analysis is actually correct too, I guess that the AI is generating the code, not the actual analysis so if the code is wrong you'll get the wrong analysis on the right data rather than the right analysis on the wrong data.
I use pandas all the time and could give quite a bit of feedback but not when you are basically trying to conflate and leverage the hard work of another project for yourself.
Not a big deal though because no one is going to use this tool as is.
I'm trying to reach out the pandas founder, and I'll definitely change the name of the repo if he prefers to.
As of the quality of the output, we are constantly working on improving the quality of it. It provides good enough results in most of the cases, but it still suffers from some hallucinations, although we have already fixed many either programmatically or with prompt engineering.
As of the dataframe, we don't pass the whole df, but only the head and 5 columns with shuffled data (also redacting the sensitive ones).
Any questions, feel free to ask!
I'm a co-founder of a non-AI data code-gen tool for data analysis -- but we also have a basic version of an LLM integration. The problem we see with tooling like Pandas AI (in practice! with real users at enterprises!) is that users make an edit like "remove NaN values" and then get a new dataframe -- but they have no way of checking if the edited dataframe is actually what they want. Maybe the LLM removed NaN values. Maybe it just deleted some random rows!
The key here: how can users build an understanding of how their data changed, and confirm that the changes made by the LLM are the changes they wanted. In other words, recon!
We've been experimenting more with this recon step in the AI flow (you can see the final PR here: https://github.com/mito-ds/monorepo/pull/751). It takes a similar approach to the top comment (passing a subset of the data to the LLM), and then really focuses in the UI around "what changes were made." There's a lot of opportunity for growth here, I think!
Any/all feedback appreciated :)
EDIT: Also, shout out to JupyterAI (https://github.com/jupyterlab/jupyter-ai) -- it's an official Jupyter project with some really awesome LLM support directly in JupyterLab. I saw it debuted at JupyterCon last week :)
Any questions, feel free to ask :)
df = df.dropna()
Wow
Edit: more seriously though… this is a non deterministic pipeline that seems to mutate your data in place with no indication of what it did.
/dev/random has been around for a while, i wonder why they didn't just use that. ;)
It would be much better if the AI just helped you write the code, it's more ergonomic and auditable, and it will scale better when your data frame is measured in the 10s to 10,000s of megabytes.
I expected that this would be a Copilot like approach that was tuned for pandas specifically, and you pass it your data frame schema/column names or something, and it passing back some code to evaluate locally.
I’m definitely not excited about passing the data frame itself and letting the llm handle it. Maybe with time I’ll get there, but I think this is a step too far for anyone that needs reproducible/verifiable results.
I think that you entered in a "niche" discussion now. Because I see that people are building tools with LLM in the data analysis space trying to go more in the business direction, where product managers (and other stakeholders, mainly in startups) don't need to write a single line of code and get the data visualization/analysis that they want.
> This because the function is no longer idempotent, each call to the AI can yield a different result.
Also, it makes it harder to verify correctness, and may make processing larger datasets (or repeatedly processing similar datasets) more expensive.
"You may not use or register, in whole or in part, our Marks as part of your own trademark, service mark, domain name, company name, trade name, product name or service name.
You may not use names or trademarks that are similar to ours without first obtaining written approval. You therefore may not use an obvious variation of any of our Marks or any phonetic equivalent, foreign language equivalent, takeoff, or abbreviation for a similar or compatible product or service."
(Numfocus is the legal entity behind pandas and other open source projects): https://numfocus.org/trademark-guidelines
can you visit https://pandas.pydata.org/about/governance.html and tell me if I am allowed to use the term 'pandas' in the name of another unaffiliated project, for example 'pandas-ai'
--
Based on the BSD 3-Clause License under which pandas is released, neither the name of the copyright holder nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission[ 1] . This means that to use the term 'pandas' in the name of another unaffiliated project such as 'pandas-ai', you would likely need to get written permission from the pandas project's copyright holders.
However, please note that this is not legal advice, and it would be a good idea to consult with a lawyer who specializes in open-source software or intellectual property law to ensure that you're in compliance with all legal requirements.
That said, I don’t get this product. Using query as code is no help at all. Integrating GPT into jupyter might be useful - I’m sure it’s been done - but not this.
This really limits this package as I frequently want to perform multiple operations on a data frame. Does OpenAI not allow sessions so the data frame must be passed each call?
I think the approach here is kind of dumb, though. It works significantly better to send the AI meta-data and let it write queries (with a limit on result count) iteratively to answer your questions.
His request was for PandasAI to -
"Plot the histogram of countries showing for each the gpd,..."
The histogram was constructed correctly according to what he thought he typed. Scary mind-reading stuff here with that AI stepping in and correcting our defective requests without even asking - " Did you mean to request GDP?"
This AI could serve as the annoying know-it-all at all of your parties, correcting everyone's anecdotes and misconceptions in real time. It only needs a high-pitched nasal voice with a simulated west Texas drawl playing from your Alexa or on-screen as a Poindexter-in-a-Stetson avatar.
(1) run notebook in VS-code https://code.visualstudio.com/docs/datascience/jupyter-noteb...
(2) install Github Copilot extension https://marketplace.visualstudio.com/items?itemName=GitHub.c...
Then if you load in pandas and a DataFrame you want to analyse, it quite often suggests the next stage of your data-analysis work, either right away, as you write the first step in some chain of pandas methods.
1. Polars has a narrower API that seems a bit cleaner and more logical vs pandas is more maximalist. Matter of taste but https://franzdiebold.github.io/polars-cheat-sheet/Polars_che... Will give you a flavour
2. Polars perf seems quite a lot better. This may not matter for your use case but if it does it does.
3. Polars supports both eager and lazy evaluation modes and you can choose whichever makes sense for your use case. I like this but sone people find lazy eval confusing to debug.
4. polars can do things like parallelization under the hood
5. Polars is less inclined to guess what you mean about things like type conversions and will instead fail and throw until you fix your code to be more explicit. Pandas is a bit more lassez faire. I guess this reflects the dev philosophy of the respective teams.
Here’s someone else’s writeup with some benchmarks https://towardsdatascience.com/pandas-vs-polars-a-syntax-and...
> pandas_ai.run(df, "plot the histogram of countries showing for each the gd, using different colors for each bar")
Am I right that this could potentially give different answers when run multiple times?
You can jump to the code at https://github.com/gventuri/pandas-ai to see more of what it's trying to do.
the larger question I have is... if you're already advanced enough to use pandas, do you really need/want this?
I have been building my own open source tool for wrangling pandas dataframes - Buckaroo[2]. The aim of Buckaroo is to automate common data cleaning and exploration techniques, and provide a GUI to quickly access transforms and see their results along with the python code to perform the transform. The jupyter notebook is great for presentation, but clunky for iterative exploration.
Here are the most common operations I do when exploring a new dataset
1. Assess - What are the columns in this dataset, what are their ranges of values, how do these columns vary together. Buckaroo offers toggable summary stats to show this info.
2. Initial Clean - Are all of the date columns parsed as datetimes? Are columns that contain integers typed as integers (a single string in the colum will force the whole column to type object). Clean the types with safeint/dropna...
3. Filter to a subset of data based on criteria. Row wise filtering.
4. Perform some type of analysis. Group by with different aggregations, a plot... Something more complex
Other operations that I regularly perform
5. Concattenating similar dataframes. pd.concat is simple, knowing whether your columns match up, along with their types is difficult. I want a UI that shows where they match and where they dont. This normally will require an extra data cleaning step
6. Joining tables/dataframe. I want a UI that tells me Is there a natural key to join on? What percentage of rows are joined vs not matched.
Buckaroo currently does an initial version 1,2, and 4. Buckaroo allows these steps to be done in a single notebook cell, instead of many separate cells iteratively built up. There are planned features that address all of these use cases. In addition Buckaroo has been engineered to enable easy extension by users. Plug your own functions in, and make them quickly accessible.
In my view a point and click gui is the best way to quickly accomplish these tasks as opposed to a conversational AI approach. The conversational AI approach doesn't provide enough structure around common tasks.
PS: This morning I added a "Related Projects" [3] Section to the Buckaroo docs. If Buckaroo doesn't solve your problem, look at one of the other linked projects (like Mito).
[1] https://github.com/approximatelabs/sketch