https://aws.amazon.com/blogs/opensource/using-strands-agents...
27 karma · joined September 21, 2019
https://aws.amazon.com/blogs/opensource/using-strands-agents...
Genuine question. Not familiar with this company or the CLI product.
The key idea lies in dynamically determining how queries are handled:
- Strong matches (≥80% similarity): Responses are directly served from the cache.
- Partial matches (60–80% similarity): Verified answers are used as few-shot examples to guide the LLM.
- No matches (<60% similarity): The query is processed by the LLM as usual.
This not only minimizes hallucinations but also reduces costs and improves response times.
Here's a Jupyter notebook walkthrough if anyone's interested in diving deeper: https://github.com/aws-samples/Reducing-Hallucinations-in-LL...
Would love to hear your thoughts—anyone else working on similar techniques or approaches? Thanks.
Excited to see what y'all cook up.
Do you mean specifically the Bedrock Knowledgebase/RAG -- that uses serverless OpenSearch which costs at minimum $200ish/month bc it doesn't scale to zero?
Makes coding with AI much faster and less of a chore. You still need to know what you're doing as the models are fantastic at boilerplate but often get things wrong. But it's significantly sped up iteration for me, and hence translates to more fun.
Not agreeing/disagreeing here, just stating author's intent.
What are some other common ways to build a portfolio towards an O-1 application other than research citations or founding a startup?
Thanks for this AMA!