AIF-C01 notes
Developing ML solutions

Fundamental Concepts of MLOps

What MLOps is, and why ML needs it

The story: A map app is perfect on the day it launches. Then roads close, new ones open, and traffic habits change. Without regular updates, the map slowly gets worse. You need a system that keeps checking and redrawing it, not a one-off effort.

In AI/AWS terms: MLOps applies DevOps practices to machine learning. It operationalizes the whole lifecycle so models are developed, deployed, monitored, and retrained systematically and repeatedly. It combines people, process, and technology.

ML needs its own operations because models degrade as data changes. A model that works at launch can get worse over weeks or months.

For the exam: Models degrade as data changes. MLOps makes retraining and redeployment systematic and repeatable.

Goals

The story: The map company wants updates to go out faster by automating them, fewer errors by testing and watching, surveyors and programmers working together, and a clear record of every change for regulators.

In AI/AWS terms:

  • Speed up the lifecycle through automation
  • Improve quality through testing and monitoring
  • Encourage collaboration between data scientists, data engineers, software engineers, and IT operations
  • Provide transparency, explainability, auditability, and security through model governance

For the exam: MLOps goals: automation, quality, collaboration, governance.

Key principles

The story: The map company:

  • Keeps every old version of the map, the survey data, and the drawing software, so it can go back if an update is wrong.
  • Has machines collect, clean, and draw the data automatically.
  • Checks every change, publishes approved maps automatically, redraws them on a schedule, and watches for complaints.
  • Has a manager approve each new map before release, checking it's fair to every neighborhood.
  • Runs one line that builds new maps and a separate line that publishes them. A new survey arriving starts the building line.

In AI/AWS terms:

  • Version control of data, code, and models, for reproducibility and rollback
  • Automation of ingestion, preprocessing, training, validation, and deployment
  • CI/CD for ML:
    • Continuous integration: tests code, data, and models
    • Continuous delivery: deploys new models automatically
    • Continuous training: retrains models automatically
    • Continuous monitoring: tracks data, models, and business metrics
  • Model governance: documentation, review and approval before deployment (checking fairness and bias), data protection, and compliance

A production ML setup usually has separate model build (training) and deployment pipelines. The build pipeline runs, for example, when new data arrives.

For the exam: Version data, code, and models. Continuous training is the ML-specific addition to CI/CD. Build and deployment pipelines are separate.

AWS services for MLOps

The story: Each step of the map company's process has its own department.

In AI/AWS terms:

  • Prepare data: SageMaker Data Wrangler (low-code) or the SageMaker Processing API
  • Store features: SageMaker Feature Store
  • Train and tune: SageMaker training and automatic model tuning
  • Track experiments: SageMaker Experiments
  • Register models: SageMaker Model Registry
  • Orchestrate: SageMaker Pipelines
  • Monitor: SageMaker Model Monitor

For the exam: Orchestration = SageMaker Pipelines. Versioned, approved models = Model Registry. Drift in production = Model Monitor.

On this page