Full threadSimFG·Maybe GPTCache can reduce the count of llm request for lower cost. detail: https://github.com/zilliztech/GPTCacheView on HN