AIF-C01 notes
Developing ML solutions

Model Deployment

What deployment is

The story: A chef has perfected a recipe in the test kitchen. Deployment is putting that chef in a real restaurant, with a stove, staff, and an order window, so customers can actually get food.

In AI/AWS terms: Deployment puts the model and its resources into production so it can make predictions.

For the exam: Deployment makes a trained model available to make predictions in production.

Self-hosted versus managed

The story: You can build and run your own restaurant: you choose everything, it might be cheaper, but you fix the plumbing yourself. Or you rent a spot in a food court: the building, cleaning, and extra seating on busy days are handled for you, and one counter can serve several of your dishes.

In AI/AWS terms:

  • Self-hosted API: you run the model on your own infrastructure (on premises, or VMs and containers in the cloud). More control and customization, possibly lower cost, but more operational overhead.
  • Managed API: a cloud service such as SageMaker AI handles the infrastructure. One-click or single-API deployment, automatic scaling, model hosting, and HTTPS endpoints that can host multiple models.

For the exam: Self-hosted = more control, more overhead. Managed = less overhead, automatic scaling.

SageMaker inference options

The story: Four ways to run a food business:

  • A restaurant open all day, where each order is cooked right away.
  • A caterer who cooks a huge order once and then closes the kitchen.
  • A cake shop: you drop off a large, complicated order, it goes in the queue, and they call you when it's ready, usually soon.
  • A pop-up stall that only opens when someone shows up. It takes a moment to fire up the grill, but you pay nothing when nobody comes.

In AI/AWS terms:

BusinessOptionBest for
All-day restaurantReal-timeInteractive, low-latency requests on a persistent endpoint
CatererBatch transformPredictions on large datasets with no persistent endpoint, or preprocessing datasets
Cake shop queueAsynchronousLarge payloads (up to 1 GB) and long processing (up to 1 hour), with near real-time latency. Requests are queued
Pop-up stallServerlessIntermittent traffic with idle periods, when cold starts are acceptable. No infrastructure to manage

For the exam: Low latency → real-time. Big dataset, no endpoint → batch. Payloads up to 1 GB or processing up to 1 hour → asynchronous. Intermittent traffic, cold starts OK → serverless.

On this page