Show HN: Transform your codebase into a single Markdown doc for feeding into AI
tesserato.web.app
tesserato.web.app
https://en.wikipedia.org/wiki/CodeWeavers
Trademark is active. It's an Ⓡ not just a ™, registered not just trademarked. To keep it, they have to demonstrate they defend it.
https://www.trademarkia.com/codeweavers-76546826
While this project drops the final "s", you don't get to launch an OS called "Window". The test is a fuzzy match based on likelihood of confusion.
This project is definitely going to get C&D'd.
find . -print -exec cat {} \; -exec echo \;
Which will return for each file (and subfolders) the filename and then the content of the file.Then `| pbcopy` to copy to clipboard and paste it into ChatGPT or similar.
Filename: demo.py
```python
...python code here...
```https://github.com/jzombie/globcat.sh
Nothing fancy, but gets the job done.
You should use Aider/Cursor for proper indexing/intelligent codebase referencing
any tips?
found the problem
Certainly, I do that several times a day.
The best results come from feeding precisely targeted context directly into the prompt, where you know exactly what the model sees and how it processes it. The prompt receives the most accurate use of attention—whereas god knows what the pipeline is for cursor or what extra layers and context restrictions they add on top of base Claude.
Giving the model a clean project hierarchy accomplishes a lot efficiently in terms of context tokens. The key is ensuring it only sees what's relevant, without diluting its attention.
Tools like reopmix and OP's version, feeding targeted context straight into models like Claude or Google's offerings, outperform Copilot and Cursor in my experience, even though they use the same base models. Use the highest-quality attention (the prompt context) directly, rather than layers of uncertainty and "proper indexing".
$ head -10000 *
==> package.json <==
{
"name": ...
...
==> tsconfig.json <==
{
"extends": ...
...
$ head -10000 * | llm -s "generate a patch to switch this project to esm"This will open a website that creates a copy of all the file contents of the repo (code, docs, ...) It's a great tool to use when using new/obscure code with LLMs in my opinion.
The UX is so just easy and great, change the URL from <https://github.com/user_name/repo_name> to <https://gitingest.com/user_name/repo_name>
//edit: fixed URLs
It's not mentioned on the page but is it using [0] in the background? Edit -> It's a Go program so I guess not.
I wrote this library [1] and hope to add the fine-grained "reference resolution" utility to it at some point, which could make implementing such a tool a lot simpler.
https://aider.chat/docs/usage/copypaste.html
and with /paste you can apply the changes.
To add some more detail, aider has a mode/UX that is optimized for "copy and paste" coding with LLM web chats. The "big brain" LLM in the web chat does the hard work, and a cheap/local LLM works with aider to apply edits to your local files.
There's a little demo video in the link above that should give you the gist.
https://github.com/kasperjunge/copcon
Point it at a code project directory to get a file tree and content, optionally with a git diff, copied to the clipboard - ready for copy pasting into ChatGPT.
It is very true that this only works for small projects, as you will bloat the LLM’s context with large codebases.
My solution to this is two files you can use to steer the tool’s behavior:
- .copconignore: For ignoring specific files and directories.
- .copcontarget: For targeting specific files and directories (applied before .copconignore).
These two files provide great control over what to include and exclude in the copied context.
find . -type f -name '*.py' -exec sh -c 'echo "# $1"; cat "$1"; echo ""' _ {} \; | pbcopyIt's actually pretty straightforward if you're in a language with lexical scoping, and it simplifies some things, like includes / cyclical, no modules, no hunting through files, etc.
I feel like this set up could integrate really well w/ AI models.
I've found that the only real limitation, at least in my experiment, was a lack of decent editor support. I use vim so this wasn't really much of an issue for me with many great ways to navigate a file, and a combination of vertical and horizontal splits on a large screen, but when I opened it up in other "modern" editors the ergonomics fell apart quite a bit.
I think the biggest downside was re-using variable names between large scopes occasionally made it hard to find the reference I wanted (E.g. i, x, key, val), but again, better editor support allowing you to limit your search to within the current scope would help. Also easily mitigated with more verbose throwaway variable naming.
I’m thinking the same approach would also work well in F#, Haskell, OCaml.
It’s easy to switch to files by name with a few keystrokes. Files are names to group things I’m looking for.
I would much rather do that than try to search through a 7,000 line file for what I need.
> I feel like this set up could integrate really well w/ AI models.
Massive files or too many files break AI models. Grouping functionality into smaller files and including only relevant files is key. The file and folder names can be hints about where to find the right files to include.
I mean I'm not arguing for it as a best practice. I did it as an experiment (as I stated), and discovered it's actually really easy, and snappy for me to navigate in Vim. Mileage may vary with other editors. Have you tried it?
> Massive files or too many files break AI models
It's growing faster than I code! With the latest Gemeni at least it's much larger at 1-2 mil tokens. I'm sure we'll hit a ceiling though, but I also think we may find some context caching / rag type optimizations eventually.
Pretty big flag that this isn't ready for primetime.
https://github.com/franzenzenhofer/thisismy
supports files, resursive directories, .gitignor and .thisismyignore and online ressources / URLs + tree commands
also available as a chrome extension https://thisismy.franzai.com/
Not a replacement for full 4M lines but it might work for some tasks/prompts
The context window for Gemini 2.0 Flash can handle roughly 50000 lines of code, and 2.0 Pro can handle twice that.
(unless the reason you're giving AI the code is that you don't have any docs for either humans or machines)
I don't know what kind of agent architecture Cursor uses internally but it seems much better designed at finding where changes need to be made.
But the approach to fit your entire codebase into one document so you can include it in your prompt context seems a dead end, instead the llm can use an agent to do targeted search through your code.
I prefer the idea of the other comment reply where you use AI as a tool to explore a codebase and assist you, not something you instruct to do the work. It can accelerate you building that experience and intuition at a level we've never been able to do before.
Regarding testing, I’ve had an interaction with windsurf where I told it there was a bug in the application it generated. It replied “I’ve added some log statements, can you run it and tell me what you see, then I’ll know what to fix”… The llm was instructing me…
In Emacs I’ve had good experience with gptel as well but I prefer aider for the coding workflow
The challenge is maintaining it... But you'd maybe ask the model to do that incrementally on every commit, or just throw it away and regenerate from scratch occasionally.
Then we can easily drop in and out of using LLMs in the code space.
Service Oriented Architecture lends itself well to the limited context of these models.
Maybe we can revive literate programming and simply build everything from a single markdown document..
It's one thing to have it be trained in billions of loc and be useful, its another for it to have enough quality dataset to have enough context and understanding of something like Kafka partition ordering and its possible interactions with something like a database and at-least once delivery. It will give you an explanation of those things in isolation, but not in combination.
https://paste.mozilla.org/9rD95yAy
I would like to be able to create sets of files that I can easily send to the clipboard in this kind of format. The files could correspond to the ones relevant to a particular feature, etc. They don't always fall under the same subtree of the source code, and the entire source code is too big for the context.
Use better tools people!
I’ve been getting complete working code with this strategy but I’m creating projects that are relatively simple.
I also notice that I have to give a little deeper context about “how” it should work, which I normally wouldn’t do.
I use it with ChatGPT's o1 pro (which can handle around 100,000 tokens).
1. Open all of the files I think are relevant
2. Use the extension to combine them
3. Copy and paste into ChatGPT
https://marketplace.visualstudio.com/items?itemName=DVYIO.co...
I think cherry-picking relevant sections would be necessary to make it function effectively. Has anyone tried using tree-sitter to recursively feed it the source for functions used in the section we want to analyze to optimize for this?
https://github.com/Dicklesworthstone/your-source-to-prompt.h...
treedump is particularly helpful.
I'm always baffled by the response they get since doing this is also the most impractical, poorly scaling, way to insert an LLM into your development process.
On one hand if you realize that, there may be times where you get lucky with the size of a codebase and the nature of your questions and it works acceptably.
But on the other, this feels like the kind of thing someone who's hearing others rave about the utility of AI will try with too large of a codebase, insert the result into ChatGPT, and then get an LLM underperforming because it's being flooded with irrelevant context for every basic operation it's being asked to do.
There are very few times when providing the entire codebase in the context window instead of the relevant code to a single operation makes sense.
Now no one will need something that can handle all of the edge cases, but whatever edge cases they need to be handled will already be handled. The overall time and frustration saved this way can be huge.
https://github.com/manfrin/bundle-codebases
I don't see much merit in things like markdown or syntax highlighting as that's just extra noise for the AI. My script tries to cut down on any extraneous data since the things I'm working on are near the context limit of consumer AIs.
My script also ignores anything in .gitignore and will take a .codebundlerwhitelist (i hate this name and have meant to change it) to only bundle files matching patterns you specify.
Got it.
Can I call this c++ code “machine code” now?