AIF-C01 notes
Fundamentals of ML and AI

Machine Learning Fundamentals

How a model gets built

The story: You're teaching a friend to cook fried rice. First you go shopping and wash the ingredients. Then you pick a recipe, let them practice, taste the result, and adjust. If you bought rotten eggs, no recipe and no amount of practice will save the dish.

In AI/AWS terms: Shopping and washing is collecting and preparing data. Picking a recipe is choosing an algorithm. Practicing is training. Tasting and adjusting is evaluating and iterating. Rotten eggs are bad data: garbage in, garbage out.

For the exam: The ML process is data preparation, algorithm choice, training, then evaluation and iteration. A model is only as good as its data.

Labeled and unlabeled data

The story: One stack of flashcards has a photo on the front and the answer on the back: "cat", "dog". Another stack has only the photos, with no answers anywhere.

In AI/AWS terms: The cards with answers are labeled data: every example comes with a target value. The cards with only photos are unlabeled data: they have input features but no answer.

For the exam: Labeled data has a target value for each example. Unlabeled data has only the input features.

Structured and unstructured data

The story: A corner shop keeps a cash book with neat columns for date, item, and price. The shop also has a fridge thermometer that writes down the temperature every hour. And it has a pile of customer chat messages and photos of receipts that nobody has sorted.

In AI/AWS terms:

  • The cash book is structured, tabular data: rows and columns, like spreadsheets, databases, and CSV files.
  • The thermometer log is structured, time series data: values over time, like stock prices, sensor readings, or weather.
  • The chats and receipt photos are unstructured data: text and images with no fixed format. They need more advanced techniques to make sense of.

For the exam: Structured data is tabular or time series. Unstructured data is text and images, and needs more advanced techniques.

Three ways to learn

The story:

  • A student works through practice questions that have an answer key at the back, and checks each answer.
  • Someone hands you a bag of mixed buttons and says "sort these however makes sense". Nobody tells you the right groups.
  • You train a puppy with treats. Sit when asked, get a treat. Jump on the sofa, no treat. Over time the puppy works out what earns treats.

In AI/AWS terms:

StoryLearning typeDataGoal
Answer keySupervisedLabeledLearn how inputs map to known outputs, then predict outputs for new data
Button sortingUnsupervisedUnlabeledDiscover hidden patterns, structure, or groups
Puppy treatsReinforcementRewards and penalties from an environmentLearn the actions that earn the most reward over time, by trial and error

If only some of the practice questions have answers in the key, that's semi-supervised learning: training on data where only part is labeled.

For the exam: Supervised uses labeled data, unsupervised finds patterns in unlabeled data, reinforcement learns from rewards and penalties. Semi-supervised uses partly labeled data.

Inferencing

The story: Once your friend has learned to cook, they start cooking for real customers. A caterer cooks 500 boxes tonight for tomorrow's event, with time to check every box. A street food stall cooks each order the moment the customer asks.

In AI/AWS terms: Cooking for customers is inferencing: using a trained model to make predictions. The caterer is batch inferencing: a large set of data processed at once, when accuracy matters more than speed, as in data analysis. The street stall is real-time inferencing: an instant response to each new piece of data, as in chatbots and self-driving cars.

For the exam: Batch inferencing when accuracy matters more than speed. Real-time inferencing when you need an instant response.

On this page