Generative AI Fundamentals
Foundation models
The story: In the old days, a town hired a different person for every job: one to write letters, one to summarize reports, one to answer questions. Then someone arrives who has read almost every book in the library. With a few instructions, this one person can write letters, summarize, answer questions, and chat.
In AI/AWS terms: The well-read person is a foundation model (FM): a very large model pre-trained on internet-scale data. Instead of training a separate model for each task, one FM is adapted to many: text generation, summarization, information extraction, image generation, chat, and question answering.
For the exam: One foundation model, pre-trained on huge data, can be adapted to many tasks.
The FM lifecycle
The story: Raising a new chef for a restaurant chain goes in a loop:
- They read every cookbook and food blog they can find, most with no teacher marking anything.
- They learn by covering the end of a recipe and guessing what comes next, then checking. Later, they keep reading new cookbooks to stay current.
- The chain sharpens them for its menu: better instructions, a recipe binder at the station, or a short focused course.
- Head chefs test them: taste tests, and "does this sell on our menu?"
- They start cooking in real restaurants.
- Customers send feedback, managers watch for problems, and the chef keeps improving.
In AI/AWS terms:
- Data selection: massive, diverse, mostly unlabeled data, because that's easy to get at scale.
- Pre-training: usually self-supervised learning, where the model makes its own labels from the structure of the data (like hiding the next word and guessing it). Continuous pre-training adds knowledge later.
- Optimization: prompt engineering, RAG, or fine-tuning.
- Evaluation: metrics and benchmarks, plus whether it meets business needs.
- Deployment: plugged into applications and APIs.
- Feedback and continuous improvement: monitoring, detecting bias and drift, and improving.
The loop repeats: it's iterative, not a straight line.
For the exam: FMs are pre-trained with self-supervised learning on mostly unlabeled data. The lifecycle is iterative.
Amazon Bedrock
The story: A food court has one payment counter, and behind it are stalls from many famous restaurants. You order from any of them at the same counter.
In AI/AWS terms: Amazon Bedrock is the food court. Through one API you reach FMs from AI21 Labs, Anthropic, Cohere, Meta, Mistral AI, Stability AI, and Amazon.
For the exam: Amazon Bedrock gives API access to FMs from many providers.
LLMs, tokens, and embeddings
The story: A reader doesn't take in a sentence all at once. They read it in small chunks: whole words, or pieces like "un-", "break", "-able". In their head, each chunk has a spot on a mental map, where related ideas sit near each other: "cat" is right next to "kitten", and both are far from "tractor".
In AI/AWS terms: Large language models (LLMs) are usually built on the transformer architecture. The chunks are tokens: words, parts of words, or characters. Each token's spot on the map is an embedding, a vector (list of numbers). Similar meanings get nearby vectors.
For the exam: Tokens are the units of text an LLM processes. Embeddings are vectors where similar meanings sit close together.
Diffusion models
The story: Take a clear photo and sprinkle a little static on it. Then more, and more, until it's pure TV snow. Now practice doing that in reverse: remove a little static at a time. Get good enough, and you can start from fresh snow and "clean" your way to a photo that never existed.
In AI/AWS terms: Adding noise until only noise is left is forward diffusion. Learning to remove it step by step to create a new image is reverse diffusion. Diffusion models are best known for text-to-image.
For the exam: Diffusion models add noise, then learn to remove it to generate new images.
Multimodal models, GANs, and VAEs
The story:
- You show a friend a holiday photo and ask "write a caption". They look at a picture and read your words, then answer in words.
- An art forger paints fakes, and an art detective tries to spot them. Every time the detective catches one, the forger gets better, and so does the detective.
- You describe a whole movie to a friend in three sentences. Later, from just a short summary like that, they can imagine and tell a full, new movie.
In AI/AWS terms:
- The friend with the photo is a multimodal model: it takes and produces several data types at once, like image plus text in, caption out.
- The forger and detective are a GAN (generative adversarial network): a generator makes fake data and a discriminator tries to tell real from fake.
- The movie summary is a VAE (variational autoencoder): an encoder compresses data into a small latent space, and a decoder generates data from samples of that space.
For the exam: Multimodal = several data types. GAN = generator versus discriminator. VAE = encoder into latent space, decoder back out.
Optimizing FM output
The story: Your new chef makes food that's not quite right for your restaurant. You have three options:
- Give clearer orders: "less salt, serve it on a banana leaf". Quick and free.
- Put your restaurant's recipe binder next to their station so they can look things up while cooking.
- Send them on a short course on your cuisine, so their habits actually change.
In AI/AWS terms:
| Story | Technique | What it does | Changes model weights? |
|---|---|---|---|
| Clearer orders | Prompt engineering | Designs the instructions and context given to the model | No |
| Recipe binder | Retrieval-augmented generation (RAG) | Retrieves relevant documents and adds them to the prompt | No |
| Short course | Fine-tuning | Supervised training on a smaller, task-specific labeled dataset | Yes |
They're listed from fastest and cheapest to most involved.
For the exam: Prompt engineering and RAG don't change the model's weights. Fine-tuning does.