SpreadsheetLLM: Encoding Spreadsheets for Large Language Models
arxiv.org
arxiv.org
The kneejerk reaction of “ugh, LLM and spreadsheet?!” is understandable, but I encourage you to watch that demo. It makes clear some obvious potentials of LLMs in spreadsheets. They can basically be an advanced autofill. If you’ve used CoPilot in VSCode, you understand the satisfaction of feeling like an LLM is thinking one step ahead of you. This should be achievable in spreadsheets as well.
[0] https://youtube.com/watch?v=0SVilfbn-HY&t=1251 (queued to demo at 20:51)
I've come close to blocking the site on my network many times but it's an absolute goldmine of interesting info too... I'm not really sure if there's a solution, other than to practice emotionally disengaging from internet discussions.
What was i supposed to take away from that demo?
Except sometimes you can't seem to stop the prose.
I suck at spreadsheets. I know they can do both useful and amazing things, but my daily life does not revolve around spreadsheets and I simply do not understand most of the syntax and operations required to make even fairly basic things work. It requires a lot of time and effort for me to get simple things done with a spreadsheet on the rare occasion that I need to manipulate one.
There are things in life that I am very good at; spreadsheets are simply not amongst them.
But do I know what I want, and I generally even have a ballpark idea of what the results should look like, and how to calculate it by hand [horror]. I just don't always know how to articulate it in a way that LibreOffice or Google Sheets or whatever can understand.
LLMs have helped to bridge that gap for me, but it's a pain in the ass: I have to be very careful with the context that I give the LLM (because garbage in is garbage out).
But in the demo, the LLM has the context already. This skips a ton of preamble setup steps to get the LLM ready to provide potentially-useful work, and moves closer to just making a request and getting the desired output.
Having one unified interface saves even more steps.
(And no, this isn't for everyone.)
Yes, "satisfaction"
Tried it once. Didn't get "satisfaction"; instead felt deeply irritated by the "backseat driver". Maybe it works better if you're just churning out mediocre boilerplate.
That "satisfaction" vanished pretty damn quickly, once I realised that I have often more work correcting the stuff so generated than I would have had writing it myself in the first place.
LLMs in programming absolutely have their uses, Lots of them actually, and I don't wanna miss them. But they are not "thinking ahead" of the code I write, not by a long shot.
There are a few things that really help the AI to understand what you want to do, otherwise it might struggle and come up with not so good code.
Not to say it gets it right everytime, but definitely often enough for me not to even consider turning it off. The time save has been tremendous.
“Oh you couldn’t take a train to work? Must have been something you did, the Palantir is usually great and helps our society. It always works great for me and my friends.”
A "wow, that's a great start" to one could be a "damn there's an issue I need to fix with this" to another. To some, that great start really makes them more productive. To others that 80% solution slows them down.
For some reason, programmers just love to be zealots and run flamewars to promote their tool of choice. Probably because they genuinely experience it's fantastic for them, and the other guy's tool wasn't, and they want them to see the light, too.
I prefer to judge people on the quality of their output, not the tools they use to produce it. There's evidently great code being written with uEmacs (Linux, Git), and I assume that, all the way on the other end of the spectrum, there's probably great code being written with VSCode and Copilot.
Web server work in Go, Python, and front end work in JavaScript - it's pretty good. Only when I try to do something truly application specific that it starts to get tripped up.
Multi threading python work - not bad, but occasionally makes some mistakes in understanding scope or appropriate safe memory access, which can be show stopping.
Deep learning, computer vision work - it gets some common pytorch patterns down pat, and basic computer vision work that you'd typically find tutorials for but struggles on any unique task.
Reinforcement learning for simulated robotics environments - it really struggles to keep up.
ROS2 - fantastic for most of the framework code for robotics projects, really great and recommended for someone getting used to ROS.
C++ work - REALLY struggles with anything beyond basic stuff. I was working with threading the other day and turned it off as all of its suggestions would never compile let alone do anything sensible.
Absolutely. LLMs let you make more programming mistakes faster than any other invention with the possible exceptions of handguns and Tequila.
To be fair, it is also really good at spewing out industrial levels of boilerplate. As we all know, 99% of the effort in coding is the writing of code and the more boilerplate in your code base the better. /s
https://i.redd.it/xr8uxqayv68d1.jpeg
For demos sure, but I'm not super hopeful on this frankly. LLMs are inherently about generating next char in sequential order. Nothing about real world spreadsheets is linear like that - they're all interlinked chaos.
I’m not sure it’s so much ‘satisfaction’; it felt more like I was having a stroke until I turned it off. Its suggestions were, like, plausibly code, but completely contextually nonsensical in general; frankly IntelliJ’s old autocomplete-with-guessing functionality was better, as it at least _knows_ a certain amount about the codebase. Now, this was on a very large old codebase; no doubt it’s better if writing trivial new things.
I did not get that feeling from CoPilot. I usually got the feeling that it was interrupting me to complete my thought but getting it wrong. It was incredibly annoying and distracting. Instead of helping me to think it was making it harder to think. Pair programming with an LLM has been great. Better than with most humans. But autocomplete sucks for me.
With not very much effort, one can explain to an LLM "here is a spreadsheet, formatted as..." which takes about 150 word tokens, and then not much more mental effort in your favorite language to translate an arbitrary spreadsheet into that format, and one gets a very capable LLM interface that can help explain complex arbitrary spreadsheets as well as generate them on request.
I've got finance professionals and attorneys using a tool I wrote doing this to help them understand and debug complex spreadsheets given to them by peers and clients.
This is interesting, it's about how you can represent spreadsheets to llms.
I open an Excel spreadsheet and also the AI Copilot. Then whenever I want to do something with Excel like "Show me which cells have formulas" CoPilot will interact with Excel and issue some command I cannot remember to do that for me?
Menus are good but often hard to navigate and find. So the CoPilot can give me a whole new (prompt-based) user-interface to any MS-application? Is that how it works?
It's also really accelerated my knowledge / skills of the specifics of the excel language.
Having an LLM being able to directly read/write at the sheet level, instead of just generating formulas for one cell, would be amazing.
By using a prompt function like LABS.GENERATIVEAI in Excel you can create solutions that combine calculations, data, and Generative AI. In my experience, transforming data to and from CSV works best for prompting in spreadsheets. Getting data to and from CSV format can be done with other spreadsheet functions.
I've created a book and course (https://mitjamartini.com/resources/ai-engineering/ebooks/han...) that teaches how to do this (both more beginner level). Just working through the examples or the examples provided by Anthropic for Claude for Sheets should be enough to get going.
Honestly, though, I kind of kid - I love spreadsheets, and if this actually works, it could be interesting. God help whoever needs to troubleshoot the hallucinated results - it's already hard enough to figure out what byzantine knotwork someone created using existing Excel functions, but now we'll have to also guess and second-guess layers of prompts that were used to either generate those same functions, or just generate output that got mulched through some AI black-box.
But then I wanted to analyze all our portfolio data over time so i had to figure out then how handle multi dimensionality in my spreadsheets. Then I figured out how to integrate and transform and reduce portfolio characteristics into sensible components for risk management and portfolio optimizations across different asset classes.
I figured out how to do some absolutely ridiculous stuff in excel, it’s tough for me to think of tools that scratch the surface me if l that is nearly as good at helping working through
The paper doesn’t speak at all to actual uses of this approach, but that doesn’t stop the article writer from assuming this is probably a big step towards automated tools that analyze spreadsheet data for non-numerically inclined users.
This is not that.
supporting correct context for comments everywhere
…without validating the results. Otherwise why should we ever use LLMs for anything?
Many do, although it's true that's often not the case.
I’m sceptical that anyone really _should_ be using generative AI for anything where correctness matters at all, but spreadsheets in particular seem close to a worst-case scenario.
=HAL(9000)
This.will not end. Well