Automatic Generation of Visualizations and Infographics with LLMs
microsoft.github.io
microsoft.github.io
> On average, once every three months, there's a startup that makes a thing that they say is going to be amazing, and it's just PivotTables. They're like, "It works with Excel, and it does this amazing consolidation, and slicing and dicing of all your data, and it's amazing, and we're going to make a startup. I'm going to sell this for four hundred ninety-five dollars." And that happens at least once every three months. The trouble is, the VCs usually know about PivotTables.
Of course this product goes a little further, making suggestions what columns to analyze and chart with an LLM. But it's quite funny to me that this Microsoft Research product is reinventing the PivotTable (+PivotChart) part with Python and Pandas.
[0]: https://youtu.be/0nbkaYsR94c?si=kkfFHZ_fyGmG3Lnj&t=2988
And the focus of this research was probably not to invent PivotTables but the Interface for these through LLMs
no code pivot tables?
https://karimjedda.com/beautiful-data-visualizations-powered...
What's your recommended approach? Even if open-source, curious to learn.
There's no moat.
> Low Resource Grammars: ... LIDA depends on the underlying LLM s having some knowledge of visualization grammars as represented in text and code in its training dataset (e.g., examples of Altair, Vega, Vega-Lite, GGPLot, Matplotlib, represented in Github, Stackoverflow, etc.). For visualization grammars not well represented in these datasets (e.g., tools like Tableau, PowerBI, etc., that have graphical user interfaces as opposed to code representations) ), the performance of LIDA may be limited without additional model fine-tuning or translation.
In other words, open source programmatic visualizations are required to feed the LLM, which then can, e.g., be licensed to corporates to accelerate various internal exploratory data analyses. A win-win for corporates and LLM providers.
Spot the loser.
Most companies use OpenSource in one way or the other.
Nonetheless, a company like MS has probably already build visualizers purely commercially (see excel) or/and is absolutly able to write it themselfs.
If I release a novel visualization library on github under some open source license I want it to be attributed to me. I don't want some specialized LLM to be lifting and offering the same visualization concepts to unnamed corporates for a hefty fee without me ever even knowing about it and these corporates pretending they don't know where that concept is coming from.
It is you choice whether you think that is a problem and how "novel" it is. Theft after all has a very long history.
Furthermore, it isn't theft to learn from others' work and reproduce similar qualities.
Good to know that the prevailing commercial tech culture now sees plagiarism and stealing ideas without attribution as the modern way of doing business and hopes that dressing things up under some algorithmic veil will hide the act.
I guess the pit of moral decline has no bottom. The consolation is that theft has never been the road to wealth. Once the plundering is over the only thing that is left is a wasteland.
It seems that Microsoft has finally found a way to kill the open source "cancer".
I.e. what is being stolen without attribution? I'm genuinely not getting what you mean in this specific case.
It seems to me that such solutions are soon to be within the set potentially constructed by an LLM.
Let me break it down for you. If I ask for a visualization that squares the circle and there is one repo that has an example of squaring the circle, the LLM will "arrive" at a way of squaring the circle.
If (1) an LLM is able to arrive at solutions in the same class of difficulty as the solution for the target problem and (2) it's not possible to establish the provenance of the solution actually offered by the LLM, then what's the argument for assuming that the solution is based on IP rather than constructive reasoning?
Retrain the LLM without access to the repo data. Ask for the same solution. Enjoy the hallucination. Provenance established.
Here the viz-related prompts (generation, editing, etc), for those interested: https://github.com/microsoft/lida/tree/main/lida/components/...
I built a tool that lets you use GPT to analyze data and build interactive graphs on the browser (https://deepsheet.dylancastillo.co/). I may try to adapt it to use LIDA or a similar approach.
I don't think you're thinking creatively enough here. A good system that makes use of these concept (because it's a research project, not a product!) will likely ensure that actions the LLM takes are non-destructive and inherently undoable. For example, if the underlying data was changed by the LLM, you can statically verify that and show a warning, emit an error, or ... something else entirely!
We take a middle ground with louie.ai of showing the database queries, data transforms, chart config, and any other decision or generation. It's nice being able to watch & check each step and then write in natural language what you want changed, so ends up feeling more like the easier side of pair programming than a blackbox.
They prompted that things MUST be correct in their prompts and it reports any transformations it does to your data, it might give you some insight into its logic to test yourself against the data.
[1] A guidance language for controlling large language models. https://github.com/guidance-ai/guidance
[2] Knowledge Infused Decoding https://arxiv.org/abs/2204.03084
Keen to see more research into this part specially making the questions more specific to the dataset in question and overlaying real-world situations.
But knowing _what_ to look for in the data given a problem statement - that's valuable, and hard to teach. LLMs have such a broad base of "knowledge", they can be reasonably good at this in just about any domain.
system_prompt = """ You are a an experienced data analyst that can annotate datasets. Your instructions are as follows: i) ALWAYS generate the name of the dataset and the dataset_description ii) ALWAYS generate a field description. iv.) ALWAYS generate a semantic_type (a single word) for each field given its values e.g. company, city, number, supplier, location, gender, longitude, latitude, url, ip address, zip code, email, etc You return an updated JSON dictionary without any preamble or explanation. """
Some spelling errors, and where is number 3?