Model Deployment
What deployment is
The story: A chef has perfected a recipe in the test kitchen. Deployment is putting that chef in a real restaurant, with a stove, staff, and an order window, so customers can actually get food.
In AI/AWS terms: Deployment puts the model and its resources into production so it can make predictions.
For the exam: Deployment makes a trained model available to make predictions in production.
Self-hosted versus managed
The story: You can build and run your own restaurant: you choose everything, it might be cheaper, but you fix the plumbing yourself. Or you rent a spot in a food court: the building, cleaning, and extra seating on busy days are handled for you, and one counter can serve several of your dishes.
In AI/AWS terms:
- Self-hosted API: you run the model on your own infrastructure (on premises, or VMs and containers in the cloud). More control and customization, possibly lower cost, but more operational overhead.
- Managed API: a cloud service such as SageMaker AI handles the infrastructure. One-click or single-API deployment, automatic scaling, model hosting, and HTTPS endpoints that can host multiple models.
For the exam: Self-hosted = more control, more overhead. Managed = less overhead, automatic scaling.
SageMaker inference options
The story: Four ways to run a food business:
- A restaurant open all day, where each order is cooked right away.
- A caterer who cooks a huge order once and then closes the kitchen.
- A cake shop: you drop off a large, complicated order, it goes in the queue, and they call you when it's ready, usually soon.
- A pop-up stall that only opens when someone shows up. It takes a moment to fire up the grill, but you pay nothing when nobody comes.
In AI/AWS terms:
| Business | Option | Best for |
|---|---|---|
| All-day restaurant | Real-time | Interactive, low-latency requests on a persistent endpoint |
| Caterer | Batch transform | Predictions on large datasets with no persistent endpoint, or preprocessing datasets |
| Cake shop queue | Asynchronous | Large payloads (up to 1 GB) and long processing (up to 1 hour), with near real-time latency. Requests are queued |
| Pop-up stall | Serverless | Intermittent traffic with idle periods, when cold starts are acceptable. No infrastructure to manage |
For the exam: Low latency → real-time. Big dataset, no endpoint → batch. Payloads up to 1 GB or processing up to 1 hour → asynchronous. Intermittent traffic, cold starts OK → serverless.