← Back to all posts
Build in Public

AWS AI Practitioner — Day 4: Deploy Inference

Learn: differences between real-time endpoints and serverless/Lambda patterns; health checks and basic monitoring.

Hands-on: deploy a minimal endpoint and send test requests.

Practice question:

Q1: What is a key tradeoff between SageMaker real-time endpoints and batch transform?

A) Real-time endpoints are always cheaper B) Batch transform suits low-latency needs C) Real-time endpoints provide low-latency predictions; batch is cost-efficient for bulk jobs D) There is no difference

Answer: C — real-time endpoints are for low-latency, batch for bulk cost-efficiency.

Daily Practice Questions (new)

Q1: How do you test a SageMaker endpoint from Python?

A) boto3.invoke_endpoint B) aws s3 ls C) ping endpoint D) aws iam create-role

Answer: A — use boto3.invoke_endpoint to call real-time endpoints.

Q2: What is a cold start for model endpoints?

A) Initial latency when container spins up B) A security breach C) A type of S3 storage class D) A logging level

Answer: A — cold start is the initial startup latency.

Q3: How can you reduce cold starts?

A) Use provisioned capacity or warm containers B) Disable monitoring C) Use smaller models only D) Delete model artifacts

Answer: A — provisioned concurrency or warm pools reduce cold starts.

Q4: What does a feature flag approach enable?

A) Toggle rollout by config or flag service B) Encrypt data at rest C) Increase S3 costs D) Disable CI

Answer: A — feature flags control rollout scope safely.

Q5: Why add health checks to endpoints?

A) Auto-detect failures and improve reliability B) Increase costs C) Expose secrets D) Reduce logs

Answer: A — health checks help detect issues and trigger remediation.

Q6: Which HTTP code normally indicates success?

A) 200 OK B) 404 Not Found C) 500 Server Error D) 302 Redirect

Answer: A — 200 indicates a successful response.

Q7: How to secure an endpoint?

A) VPC endpoints, IAM policies, or API Gateway auth B) Open it publicly with no auth C) Store keys in code D) Disable TLS

Answer: A — use network/IAM/API controls to secure endpoints.

Q8: What is a model container?

A) The runtime image serving the model B) A physical shipping container C) S3 bucket type D) CloudWatch metric

Answer: A — a container image runs the model serving code.

Q9: How to log predictions for analysis?

A) Emit logs or custom CloudWatch metrics B) Delete requests C) Store only in memory D) Disable logging

Answer: A — use logs or metrics to capture predictions for analysis.

Q10: What is a canary deployment?

A) Deploy to a subset of traffic and expand if healthy B) Roll out to all users instantly C) Delete previous version immediately D) Disable monitoring

Answer: A — canary rollouts reduce risk by gradual expansion.

Review Questions (previous lessons)

Q1: Where should model artifacts be stored for durability?

A) S3 B) tmp/ folder only C) local disk on laptop D) CloudWatch

Answer: A — S3 provides durable storage for artifacts.

Q2: Which service collects logs and metrics on AWS?

A) CloudWatch B) Route 53 C) IAM D) S3 Glacier

Answer: A — CloudWatch is used for observability.

Q3: When should you use Batch Transform instead of real-time endpoints?

A) For large offline jobs and bulk scoring B) For low-latency user requests C) For DNS routing D) For IAM policy updates

Answer: A — batch suits bulk processing needs.

Q4: How can you check IAM permissions before running a job?

A) Use the Policy Simulator or assume-role tests B) Change instance types C) Delete the role D) Use public S3 buckets

Answer: A — policy simulation and assume-role validate permissions.

Q5: What is a basic smoke test for a deployed endpoint?

A) Send a sample request and validate response schema and status B) Delete logs C) Increase instance sizes D) Modify DNS

Answer: A — smoke tests verify basic functionality quickly.

Q6: Why measure latency for endpoints?

A) To meet SLAs and ensure user experience B) To reduce security C) To encrypt data D) To change regions

Answer: A — latency affects user experience and SLA adherence.

Q7: Where to save run metadata for experiments?

A) S3 or dedicated tracking systems B) Route 53 records C) IAM groups D) /etc/hosts

Answer: A — metadata in S3 or tracking tools aids reproducibility.

Q8: What is a minimal deployment flow for a model?

A) Package artifact, upload to S3, create endpoint config and endpoint B) Commit code only C) Update DNS D) Use CloudFront

Answer: A — packaging and creating endpoint are the core steps.

Q9: How to validate model outputs programmatically?

A) Use unit tests and sample test cases to check outputs B) Rely on manual eyeballing only C) Disable logging D) Use random data

Answer: A — automated tests ensure outputs match expectations.

Q10: What practices reduce blast radius in deployments?

A) Feature flags, small PRs, and incremental rollouts B) Large monolithic releases only C) No testing D) Deleting backups

Answer: A — small changes and flags minimize risk.