AWS AI Practitioner — Day 4: Deploy Inference
Learn: differences between real-time endpoints and serverless/Lambda patterns; health checks and basic monitoring.
Hands-on: deploy a minimal endpoint and send test requests.
Practice question:
Q1: What is a key tradeoff between SageMaker real-time endpoints and batch transform?
A) Real-time endpoints are always cheaper B) Batch transform suits low-latency needs C) Real-time endpoints provide low-latency predictions; batch is cost-efficient for bulk jobs D) There is no difference
Answer: C — real-time endpoints are for low-latency, batch for bulk cost-efficiency.
Daily Practice Questions (new)
Q1: How do you test a SageMaker endpoint from Python?
A) boto3.invoke_endpoint B) aws s3 ls C) ping endpoint D) aws iam create-role
Answer: A — use boto3.invoke_endpoint to call real-time endpoints.
Q2: What is a cold start for model endpoints?
A) Initial latency when container spins up B) A security breach C) A type of S3 storage class D) A logging level
Answer: A — cold start is the initial startup latency.
Q3: How can you reduce cold starts?
A) Use provisioned capacity or warm containers B) Disable monitoring C) Use smaller models only D) Delete model artifacts
Answer: A — provisioned concurrency or warm pools reduce cold starts.
Q4: What does a feature flag approach enable?
A) Toggle rollout by config or flag service B) Encrypt data at rest C) Increase S3 costs D) Disable CI
Answer: A — feature flags control rollout scope safely.
Q5: Why add health checks to endpoints?
A) Auto-detect failures and improve reliability B) Increase costs C) Expose secrets D) Reduce logs
Answer: A — health checks help detect issues and trigger remediation.
Q6: Which HTTP code normally indicates success?
A) 200 OK B) 404 Not Found C) 500 Server Error D) 302 Redirect
Answer: A — 200 indicates a successful response.
Q7: How to secure an endpoint?
A) VPC endpoints, IAM policies, or API Gateway auth B) Open it publicly with no auth C) Store keys in code D) Disable TLS
Answer: A — use network/IAM/API controls to secure endpoints.
Q8: What is a model container?
A) The runtime image serving the model B) A physical shipping container C) S3 bucket type D) CloudWatch metric
Answer: A — a container image runs the model serving code.
Q9: How to log predictions for analysis?
A) Emit logs or custom CloudWatch metrics B) Delete requests C) Store only in memory D) Disable logging
Answer: A — use logs or metrics to capture predictions for analysis.
Q10: What is a canary deployment?
A) Deploy to a subset of traffic and expand if healthy B) Roll out to all users instantly C) Delete previous version immediately D) Disable monitoring
Answer: A — canary rollouts reduce risk by gradual expansion.
Review Questions (previous lessons)
Q1: Where should model artifacts be stored for durability?
A) S3 B) tmp/ folder only C) local disk on laptop D) CloudWatch
Answer: A — S3 provides durable storage for artifacts.
Q2: Which service collects logs and metrics on AWS?
A) CloudWatch B) Route 53 C) IAM D) S3 Glacier
Answer: A — CloudWatch is used for observability.
Q3: When should you use Batch Transform instead of real-time endpoints?
A) For large offline jobs and bulk scoring B) For low-latency user requests C) For DNS routing D) For IAM policy updates
Answer: A — batch suits bulk processing needs.
Q4: How can you check IAM permissions before running a job?
A) Use the Policy Simulator or assume-role tests B) Change instance types C) Delete the role D) Use public S3 buckets
Answer: A — policy simulation and assume-role validate permissions.
Q5: What is a basic smoke test for a deployed endpoint?
A) Send a sample request and validate response schema and status B) Delete logs C) Increase instance sizes D) Modify DNS
Answer: A — smoke tests verify basic functionality quickly.
Q6: Why measure latency for endpoints?
A) To meet SLAs and ensure user experience B) To reduce security C) To encrypt data D) To change regions
Answer: A — latency affects user experience and SLA adherence.
Q7: Where to save run metadata for experiments?
A) S3 or dedicated tracking systems B) Route 53 records C) IAM groups D) /etc/hosts
Answer: A — metadata in S3 or tracking tools aids reproducibility.
Q8: What is a minimal deployment flow for a model?
A) Package artifact, upload to S3, create endpoint config and endpoint B) Commit code only C) Update DNS D) Use CloudFront
Answer: A — packaging and creating endpoint are the core steps.
Q9: How to validate model outputs programmatically?
A) Use unit tests and sample test cases to check outputs B) Rely on manual eyeballing only C) Disable logging D) Use random data
Answer: A — automated tests ensure outputs match expectations.
Q10: What practices reduce blast radius in deployments?
A) Feature flags, small PRs, and incremental rollouts B) Large monolithic releases only C) No testing D) Deleting backups
Answer: A — small changes and flags minimize risk.