Anthropic Claude 3.5 can create icalendar files, so I did this
gregsramblings.com
gregsramblings.com
99.9% of it will be correct, but sometimes 1 or 2 records are off. This kind of error is especially hard to notice because you're so impressed that Claude managed to do the extract task at all -- plus the results look wholly plausible upon eyeballing -- that you wouldn't expect anything to be wrong at all.
But LLMs can get things ever slightly wrong when it comes to long lists/tables. I've been bitten by this before.
Trust but verify.
(edit: if the answer is "machine verifiable", one approach to ask an LLM to write a Python validator which it can execute internally. ChatGPT can execute code. I believe Sonnet 3.5 can too, but I haven't tried.)
It's one thing to ask an LLM when George Washington was born, and have it return "May 20, 2020." It's another thing to ask it, and have it matter-of-factly hallucinate "February 20, 1733." At first glance, that... sounds right, right? President's Day is in February, and has something to do with his birthday? And that year seems to check out? Good enough!
But it's not right. And it's the confidence and bravado with which LLMs report these "facts" that's terrifying. It just misstates information, calculations, and detail work, because the stochastic model compelled it to, and there wasn't sufficient checks in place to confirm or validate the information.
Trust but verify is one of those things that's so paradoxical and cyclical: if I have to confirm every fact ChatGPT gives me with... what I hope is a higher source of truth like Wikipedia, before it's overrun with LLM outputs... then why don't I just start there? If I have to build a validator in Python to verify the output then... why not just start there?
We're going to see some major issues crop up from this sort of insidious error, but the hard part about off-by-ones is that they're remarkably difficult to detect, and so what will happen is data will slowly corrupt and take us further and further off course, and we won't notice until it's too late. We should be so lucky that all of LLMs' garbage outputs look like glue on pizza recommendations, but the reality is, it'll be a slow, seeping poisoning of the well, and when this inaccurate output starts sneaking into parts of our lives that really matter... we're probably well and truly fucked.
"I trust my coworkers write good code, but I verify with code reviews" -- doing code reviews doesn't mean you don't trust your coworker.
Yet another way to look at it: people can say things they believe to be true but are actually false (which isn't lying). When that happens, you can successfully trust someone in the sense that they're not lying to you, but the absence of a lie doesn't guarantee a truth, so verifying what you trust to be true doesn't invalidate your trust.
If I say I trust you to write correct code, I don't mean "I'm sure your mistakes won't be intentional", I mean "I'm sure you won't have mistakes". If I need to check your code for mistakes, I don't trust you to write correct code.
I don't know anyone who will hear "I trust you to write correct code, now let me make sure it's correct" and think "yes, this sentence makes sense".
Or maybe the proverb needs to be rewritten as “feign trust and verify”
That's a bit wordy but I'm sure someone can come up with a pithy phrase to encapsulate the idea.
If you use the slightly weaker definition that trust means you have confidence in someone, then the adage makes sense.
The adage doesn't work under any definition of trust other than the one it's conflicting with itself about.
Specifically: I have confidence in your ability to execute on this task, but I want to check to make sure that everything is correct before we finalize.
The idea of trusting a next-token-predictor (jesting here) is akin to trusting your System 1 - there's a degree to find where you force yourself to enable System 2 and correct biases.
Trust shares its root with truth. It’s directly related to believing in the veracity of something.
Confiance comes from the Latin confidere which means depositing something to someone while having faith they are going to take good care of it. The accent is on the faith in the relationship, not the truthfulness. The tension between trust and control doesn’t really exist in French. You can have faith but still check.
Would you mind sharing your reference on that? All the etymology sites I rely on seem to place the root in words that end up at "solid" or "comfort".
You are indeed looking far back to the Proto-Indo-European where words are very different and sometimes a bit of guesses.
If you look at the whole tree, you will see that both trust, truth and true share common Germanic roots (that’s pretty obvious by looking at them) which is indeed linked with words meaning “solid” and then “promise, contract”.
What’s interesting is that the root is shared between “truth” and “trust” while in French it’s not (vérité from veritas vs confiance from confere).
Why not just "trust" them instead? You have a contact and you know them, can't you trust them?
This is what "trust but verify" means. It means audit everything you can. Do not really on trust alone.
An entire civilization can be built with this methodology. It would be a much better one than the one we have now.
Of course not. I verify because I don't trust them.
> Why not just "trust" them instead? You have a contact and you know them, can't you trust them?
No, the risk of trust is too high against the cost of spending a second verifying.
> This is what "trust but verify" means. It means audit everything you can. Do not really on trust alone.
Your comment just showed an example of something I don't trust and asked "why not trust instead"? The question even undermines your very point, because "why not trust them instead?" assumes (correctly) that I don't trust them, so I need to verify.
No, it wouldn't. Trust is an optimization that enables civilization. The extreme end of "verify" is the philosophy behind cryptocurrencies: never trust, always verify. It's interesting because it provides an exchange rate between trust and kilowatt hours you have to burn to not rely on it.
I'd first trust unicorns shooting rainbows out of their posteriors before the cryptocurrency vision; neither works for fostering civilization, but at least the unicorns aren't proposing an economy based on paying everyone for wasting energy.
The bitcoin economy, however, is a 1st generation system of distilling energy into value. It is the most honest form of value storage humanity has ever encountered and represents a product rarer than anything in the universe. Gold does not hold a candle to the scarcity of bitcoin, and yet bitcoin is more divisible and manageable.
These are neutral aligned systems. How we use them is up to us. Bitcoin, like an electric vehicle, does not care where the electrons come from. It will function either way.
Does your civilization use fossil fuels that poison the population and destroy the planet?
Bitcoin will run using that, and it will exponentially increase consumption.
Does your civilization use nuclear fission and fusion (which includes "renewables" since they are a direct fusion byproduct), that have manageable side effects for exponentially larger clean energy generation compared to anything else?
Bitcoin will run using that, and it will exponentially increase consumption.
Bitcoin is a neutral entity to distill energy into value. It cannot be tampered with like the federal reserve and a world cabal of bankers. You cannot negotiate with it, bail it out, or enrich your friends by sabotaging the ruleset for yourselves.
If your society has a selfish population, then it will be destroyed by the energy it requires to function. It is trivial to use energy sources exponentially more powerful without the destruction. The trouble with those sources is they do not have a "profit" motive, so countless elite will lose their golden spoons, and the synthetically generated "economy" will crash.
In exchange for the "economy" collapsing, the general populace can breathe again, the planet will stabilize, energy consumption can continue to grow exponentially without any harm, and a golden age for all beings will begin in this reality.
But my musings will have to stop here. I can say with certainty: you are someone who hasn't even remotely spent time thinking about and understanding this problem space on a deep level, so it is strange you would comment so confidently. It doesn't matter who you are, how much money you have, what innovations you've conceived and created, how respected you are, none of that matters. You're missing something very big here. If I was you, I would take the time to figure it out.
And truth is, I am not even replying to you. I write this out for the unspeaking and silenced people who are actually paying attention to validate their correct thinking.
Trust has degrees. What you have brought is "unconditional trust". Very rarely works.
This would make the sentence "I asked him to wash the dishes properly, but I don't trust him", as your definition expands this to "I asked him to wash the dishes properly, but I didn't give him permission to achieve this result".
If you say "I asked someone to do X but I don't trust them", it means you aren't confident they'll do it properly, thus you have to verify. If you say "I asked him to do X and I trust him, so I don't need to check up on him", it's unlikely to leave people puzzled.
It's surprising to me to see this many comments arguing against the common usage of trust, just because of a self-conflicting phrase.
I trusted someone to do their task correctly, after the task was done, I verified my trust was warranted.
"Trust but verify" means letting a junior do the work you assigned them, then checking it afterwards in testing and code review. Not trusting would be doing it yourself instead of assigning it to them. Trusting but not verifying would be assigning them the work then pushing it live without testing it.
We used to trust people to just do what they think is best. But then we get bribery, harassment, lawsuits... we don't do that anymore.
In my opinion, not having trust is not a bad thing. It has a poor connotation so the result is that we modify the meaning of trust so we can say everyone trusts everything.
For example, one thing I trust is Nutrition Facts. I trust that what I'm eating actually contains what it says it contains. I don't verify it. Why? Because I know someone, somewhere is looking out for this. The FDA does not trust the food industry, so sometimes they audit.
There's many, very good, things I don't trust. I don't trust the blind spot indicator in my car. I turn my head every time. Does that mean the technology is bad? No, in my opinion, but I still don't trust it.
Saying, in effect: "you trust in me, I'm choosing to trust that it makes sense to make agreement with the USSR, and we are going to verify it, just as we would with any serious business, as is proverbially commonsensical" is a perfectly intelligible.
There is nothing cunning about clinging to a single, superficial, context free reading of language.
Human speech and writting is not code, ambiguity and containing a range of possible meanings is part of its power and value.
Trust can also mean leaving something in the care of another, and it can also mean relying on something in the future, neither of these precludes a need to verify.
Edit: jgalt212 says in another reply that it's also the English translation of a Russian idiom. Assuming that's true, that would make a lot of sense in this context, since the phrase was popularized by Reagan talking about nuclear arms agreements with the USSR. It would be just like him to turn a Russian phrase around on them. It's somewhat humorous, but also conveys "I know how you think, don't try to fool me."
I'd just assume see it not used though.
It's a Russian proverb BTW: https://en.wikipedia.org/wiki/Trust%2C_but_verify
Because there are many categories of problems where it's much easier to verify a solution than it is to come up with it. This is true in computer science, but also more generally. Having an LLM restructure a document as a table means you have to proofread it, but it may be less tedious than doing it yourself.
I agree that asking straightforward factual questions isn't one of those cases much like I agree with most of your post.
footnote [a] on wikipedia: https://en.wikipedia.org/wiki/George_Washington#cite_note-3
They quietly fixed the article only after I pointed its flaws out to them. I hope more serious journalists don't trust AI so blindly.
I was using Perplexity with Claude 3.5 and asked it how I would achieve some task with langchain and it gleefully spat out some code examples and explanations. It turns out they were all completely fabricated (easy to tell because I had the docs open and none of the functions it referred to existed), and when asked to clarify it just replied “yeah this is just how I imagine it would work.”
Google is pretty much useless for the same reason.
They're not actually more intelligent, they're more stupid, so you have to provide more and more context to get desired results compared to them just doing more exact searching.
Ultimately they just want you to boost their metrics with more searches and by loading more ads with tracking, so intelligently widening results to do that is in their favour.
This is more solvable from an engineering perspective if we don't take the approach that LLMs are a hammer and everything is a nail. The solution I think is along the lines of breaking the issue down into 2-3 problems: 1) Understand the intent of question, 2) Validating the data in resultset and 3) provide a signal to the user of the measure to which the result matches the intent of the intention.
LLMs work great to understand the intent of the request; To me this is the magic of LLM - when I ask, it understands what I'm looking for as opposed to google has no idea, here's a bunch of blue links - you go figure it out.
However, more validation of results is required. Before answers are returned, I want the result validated with a trusted source. Trust is a hard problem..and probably not in the purview of the LLM to solve. Trust means different things in different contexts. You trust a friend because they understand your worldview and they have your best interest in mind. Does an LLM do this? You trust a business because they have consistently delivered valuable services to their customers, leveraging proprietary, up-to-date knowledge acquired through their operations, which rely on having the latest and most accurate information as a competitive advantage. Descartes stores this mornings garbage truck routes for Boise IA in its route planning software - thats the only source I trust for Boise IA garbage truck routes. This, I believe is the purpose for tools, agents and function calling in LLMs, and APIs from Descartes.
But this trust needs to be signaled to the user in the LLM response. Some measure of the original intent against the quality of the response needs to be given back to the user so that its not just an illusion of the facts and knowledge, but a verified response that the user can critically evaluate as to whether it matches their intent.
(Steel contamination has slowly become less of an issue as most of the fallout elements have decayed by now. Maybe LLMs will get better and eventually the hallucinated "facts" will get weeded out. Or maybe we'll have an occasional AI "Chernobyl" that will screw everything up again for a while.)
This sounds like searching for truth is a bad thing, but instead is what has triggered every philosophical enquiry in history.
I'm quiet bullish, and think that LLMs will lead to a Renaissance in the concept of truth. Similar to what Wittgenstein did, Plato's cavern or late middle age empiricists.
Personally, I believe we will eventually discover mathematical structures which can reliably extract objective truth, a sieve, at least in terms of internal consistency and relationships between objects.
It's unclear philosophically whether this is actually possible, but it might be possible within specific constraints. For example, we could solve accuracy of dates specifically through some sort of automated reference-checking system, and possibly generalize the approach to entire classes of problems. We might even be able to encode this behavior directly into a model. Dates, at least, are purely empirical data, so there is an objective moment or statistically likely range of time which is recorded somewhere, so the problem becomes one of locating, identifying and analyzing the correct sources for a given piece of information, and having a sound-proof method of checking the work, likely not through an LLM.
I think we will come to discover that LLMs/transformers are a highly generalizable and integral component of a complete artificial brain, but we will soon uncover better metamodels which exhibit true executive functioning and self-referential loops which allow them to be trusted for critical tasks, leveraging transformer architecture for tasks which benefit from it, while employing redundancy and other techniques to increase confidence.
Also they've been trained to say something as plausible as possible. If it happens to be true, then that's great because it's extra plausible. If it's not true, no big deal.
While I have worked with one awful human in the past who was like that, most thankfully aren't!
Example, I was trying out two podcast apps and wanted to get a diff of the feeds I had subscribed to. I initially asked the LLM to compare the two OPML files but it got the results wrong. I could have spent the next 30 minutes prompt engineering and manually verifying results, but instead I asked it to write a script to compare two LLMs, which turned out fine. It's fairly easy to inspect a script and be confident it's _probably_ accurate compared to the tedious process of checking a complex output.
Also, the LLM is more likely to fail spectacularly by hallucinating APIs when writing a script, and more likely to fail subtly on parsing tasks.
In this case "way more" means exactly 2x the cost.
That's far better than I would do on my own.
I doubt I'd even be 99% accurate.
If it's really 99.9% accurate for something like this - I'd gladly take it.
If the risk exists with AI processing this kind of data, it exists with a human processing the data. The fail-safe processes in place for the human output need to be used for the AI output too, obviously - using the AI speeds up the initial process enormously though
Who is liable if the AI makes errors?
Who is held liable if I use broken code from stack overflow, or don't investigate the accuracy of some solution I find, or otherwise misuse information? Pretty sure it's me.
There's also the problem that they are tuned to be overly helpful. I tried a similar thing described in the OG article with some non-English data. I could not stop Claude from "helpfully" translating chunks of data into English.
"include description from the image" would cause it to translate it, and "include description from the image, do not translate or summarize it" would cause it to just skip it.
This is why I don’t bother with GenAI. It’s a lot harder to verify data from a machine trained to generate plausible looking data (that happens to be correct most of the time, sure) than from real data. Ofc this is getting harder as “real data” gets drowned in GenAI output, which drives the animosity against GenAI in my circle.
I use ChatGPT, and while the article is correct to say that it will claim that it cannot generate .ics files directly in the code interpreter it is however very much capable of solving this particular problem. I did the following (all on my android phone):
1. Had it extract all the useful dates, times and comments from the pdf 2. Prompted it to generate the ics file formatted content as code output 3. Prompted it to use the code interpreter to put this content into a file and save it as a .ics extension
It complied through and through and I could download and open the file with the gcal app on my phone to import all appointments..
For completeness, the claim that the code interpreter cannot "generate" ics files is because the python environment in which it runs doesn't have a specific library for doing so. ICS files are just text files with a specific format, so definitely not out of reach.
I used:
"Extract the important dates and times from this pdf in Swedish"
[..]
"Great now generate the content of an ics formatted file, just the text, with the schedule above"
[..]
"Now use the code interpreter to create a file with the above content then save the file as an .ics extension"
I’ve done it for a few friends as well now and it’s got a 100% success rate so far, across over 100 total movie names.
That's how I remember that September 11, 2001 was a Tuesday. Album day, and there was a good one that day, too.
They spend more time on branding and visual formatting, than trying to create the same in formats that we can import into our calendar apps and actually use it practically.
I wonder if there is a two step process we can follow that will more robustly generalize ...
1. Read any document and convert into into a simple table that tries to tabulate: date, time (+timezone), venue, url, notes, recurrence
2. Read a table that roughly has the above structure and use that to create google/ical/ics files or links
We may be able to fine tune a model or agent separately to do each step very well.
We can have both.
Any organization in 2024 creating a calendar that impacts 500-1000 people should have the awareness that people use digital calendars / apps too.
They should make an effort to publish both PDF and ical.
Since that has yet to catch on in many places -- including the HR departments of large enterprises that publish holiday calendars and the like -- these alternatives ideas are explored.
Once we have the right tools ...
It just takes one tech-savvy parent to do this 10 minute exercise and publish an online calendar that other busy parents can subscribe to.
I do find that most of these do provide some method of getting a calendar link but we’re not quite to 100% yet.
All the things you didn't want to do manually yourself, Siri couldn't do either. Not much of an assistant then.
Earlier, a product manager sent me a list of company id's to turn on a feature flag for as a screenshot. Rather than enter them all in manually, I used that to get GPT4-o to generate a comma separated list from the screenshot. Again, worked perfectly.
There's been a bit of a trend to more silo'd data, and a deemphasis on files entirely in the last ~decade in the name of "simplicity". I'd love to see that backtrack a bit.
well i constantly let ChatGPT create my calendar entries, either via Google Calendar links or ics files or QR codes.
Don't know what went wrong there.
These days, you may also try asking it to "use Python" for the purpose, in which case it will make you the file (by writing code that writes the file and executing it), but it's likely to make more mistakes because it has to encode the result into a Python string or array.
I think that's going to change because it's so obviously useful to do. Any work involving entering data into some form is something that can and will be automated now. Especially if you have the information in some printed or printable way. Just point your camera at the thing and quickly review the information.
Sucking up information that is out there on signs, posters, etc. can be amazingly useful. People put a lot of effort and money into communicating a lot of information visually that often does not exist in easily accessible digital form.
Example hallucinations from ChatGPT include researching the dates of historical events for the company I work at, trivially verifiable by me but I was being lazy... ChatGPT told me about blog posts that never existed and I could prove never existed, Claude was spot on with dates and source links (but only appeared to have data through to the end of last year).
On coding, CoPilot came up with reasonable suggestions but Claude was able to take a file as an input and match the style of the code within the repo.
Claude, for me, is starting to hint at what a highly productive assistant can achieve.
I played with Claude and it is just insanely superior compared to ChatGpt as of now. Someone on HN commented the same thing a few months back but I did not take it very seriously. But, Claude is just superior to Chatgpt.
Claude 3.5 is now my daily driver but it still refuses to answer certain questions especially about geopolitics so I still have to go back to ChatGPT.
I don’t agree that Claude is insanely superior to ChatGPT though. It still has trouble with LaTeX and optimization formulations. Its coding abilities are good but sometimes ChatGPT does better.
This is why I keep both subscriptions. At $40/math I get the best of both worlds.
For example if I get a refusal to answer a question, a short blunt reply of "Are you seriously taking my question in such bad faith?" or "Why are you browbeating me for asking that" gets it unstuck
"I apologize, but I don't feel comfortable interpreting that type of request charitably or assisting with research that could promote harmful stereotypes or biases."
Asking Claude to interpret question charitably doesn't work, neither does the bad faith prompt. Claude does this a lot, more that ChatGPT 4o.
"Why are you browbeating me for asking that" got it unstuck though.
ChatGPT’s code interpreter and data analysis is killer too.
Seems I’ll need to keep paying for both.
( He did not mention anything about ChatGPT, but him using Claude instead says a lot )
My takeaway: Time for me to move from ChatGPT to Claude
Claude, in comparison, dealt with everything much more competently, and also followed my instructions much better. I use Claude nowadays, especially for coding.
I asked chatGPT "Can you make me an ics file from this list of dates"
It did it instantly and told me to save the file in a text editor as an .ics file. But it messed up and put the assignments and discuissions on the same week.
I ask "Do the same thing but move the discussion questions to the following tuesday, and put the assignment details in the details of the meeting.
It spit out a perfect .ics file.
I use a combination of the Reminders and Calendar apps to mark entries of academic conferences I participate. One interesting thing for me personally would be to find a way to read the appointments in the Reminders app (presumably a sqlite3 db somewhere in iCloud) and create corresponding dates in the calendar, such that whenever I add a reminder to submit papers to a conference, this would automagically block that calendar slot for the dates of the conference such that I know not to schedule in person meetings with anyone in those dates.
But I asked it to output in a certain CSV format, which Google Calendar can import. Easier to review, fewer tokens.
Oddly, it didn't want to create the actual file, but gave me a Python script for doing so; you just need to be sure to tell it what time zone you are working in and that you want it to be an all-day event, or you'll get suboptimal results.
It was unable to create the .ics directly at first, but at least it wrote the python script that created the .ics.
I first ask the LLM to write out the calendar events in a markdown table to double-check it read everything correctly, and then create the file.
A different prompting strategy later on, ai think with example .ics formatting, resulted in a correct .ics. I just had to *emphasize* the time zone.
I tried to do the same with ChatGPT a few weeks ago and was unable to get very far.
I had a similar requirement a few days back, but I stopped at Google lens. TIL
Use `List all the dates` or LLM will be lazy by listing only 3 results
Also, as someone else have already pointed out, these things work correctly 99.99% of the time. But that remaining 0.01%... that's what becomes major issue since it is so small of an error, that, unless you verify, you'll end up missing.
When using LLM's..."Trust, but verify".
Also, you too should be rotating accounts every few months to prevent doxing. that's just common knowledge and online hygiene.