The context length is limited, for gpt-3.5 it's 4k tokens, there are other offerings which offer upto 100k(claude). 100k tokens is ~1 book., but it priced steeply for each call. It's often wiser, cheaper to Retrieve the context from your text & Augment your query to the LLM to Generate more contextual answers. That's the reason for the name Retrieval Augmented Generation (RAG).
For Retrieving - you'd need a vector database (for similarity comparison you can use semantic or vector embedding based similarity search)