Retrieval-Augmented Generation
What RAG is
The story: An open-book exam. The student is smart, but they've never read your company's manuals. Before answering each question, they quickly flip to the most relevant pages, then write the answer using those pages. Their answer is accurate and up to date because the book is.
In AI/AWS terms: On their own, FMs don't know an enterprise's documents, manuals, reports, or transactions. RAG searches that enterprise data using the prompt and adds the most relevant pieces to the prompt as context, so the answer is accurate and up to date.
For the exam: RAG adds relevant company data to the prompt at answer time. It doesn't retrain the model.
How it works
The story: How a very well-organized librarian finds the right book fast:
- Every book gets a spot on a giant map of topics. Books about the sea sit right next to books about the ocean, and far from books about office supplies.
- The map and the books' details are kept in a special catalog designed to find neighbors on that map very fast, even among billions of books.
- When you ask a question, the librarian puts your question on the same map and grabs the books closest to it.
- They hand those books to the writer, who answers your question using them.
In AI/AWS terms:
- Embed: an ML model turns documents, images, or audio into vector embeddings, numbers in an n-dimensional space. Related items have nearby vectors: "sea" is close to "ocean" and far from "stapler".
- Store: the vectors and their metadata go into a vector database built for fast similarity search across billions of vectors.
- Retrieve: the prompt is embedded too, and the database returns the closest matches using k-nearest neighbors (k-NN) or cosine similarity.
- Generate: the FM answers using the prompt plus the retrieved context.
For the exam: RAG steps: embed, store in a vector database, retrieve with k-NN or cosine similarity, generate.
AWS vector database options
The story: There are several brands of that special catalog to choose from, some you manage and some that manage themselves.
In AI/AWS terms:
- Amazon OpenSearch Service (provisioned)
- Amazon OpenSearch Serverless
- pgvector in Amazon RDS for PostgreSQL
- pgvector in Amazon Aurora PostgreSQL-Compatible Edition
- Amazon Kendra
In the AnyCompany case, the chatbot now queries a database of enterprise data and gives contextual answers.
For the exam: Vector stores on AWS: OpenSearch (provisioned or Serverless), pgvector in RDS or Aurora PostgreSQL, and Kendra.