Also they did the thing that junior developers tend to do, where you have a race condition of some sort, and they just work around it by adding some if checks. The app is at around 400 lines right now, it works but feels pretty brittle. Adding a tiny feature here or there breaks something else, and GPT does the wrong thing half the time.
All in all, I'm not complaining, because I made an app in two days, but it won't replace a developer yet, no matter how much I want it to.
For example, I'm currently working on a Rust/Qt desktop app so I have a project with the whole Qt6 book attached to ask questions about Qt, a project with my SQL schema and ORM/Sqlite docs to ask questions about the app's data and generate models without dealing with hallucinations, a project with all my QML files and Rust QML element code, a project with a bunch of Rust crate docs, and so on and on.
GPTs allow attaching files too but Claude Projects dump the entire contents of the files into the context rather than trying to do some hacky RAG that never works like I want it to.
The UX of drag and dropping a few monolithic markdown files to include entire chunks of a large project outweighs the downsides of including irrelevant context in my experience.
The more context you give the llm, the better.
The key takeaway from that paper is to keep your instructions/questions/direction in the beginning or at the end of the context. Any information can go anywhere.
Not to be too dismissive, it's a good paper, but we're one year further and in practice this issue seems to have been tackled by training on better data.
This can differ a lot depending on what model you're using, but in the case of claude sonnet 3.5, more relevant context is generally better for anything except for speed.
It does remain true that you need to keep your most important instructions at the beginning or at the end however.
c.f.
https://pbs.twimg.com/media/GH2NJMxbYAAcRL3?format=jpg&name=...