← Back to all posts
Build in Public

AWS AI Practitioner — Day 5: Tests & Monitoring

Learn: why observability matters for ML; simple metrics to track (error rate, latency, invocation count).

Hands-on: emit a custom CloudWatch metric for failed inferences and create an alarm.

Practice question:

Q1: Which service is commonly used to collect logs and metrics for SageMaker endpoints?

A) CloudWatch B) S3 C) Route 53 D) Secrets Manager

Answer: A — CloudWatch collects logs and metrics; use it for monitoring.

Daily Practice Questions (new)

Q1: How do you create a custom CloudWatch metric?

A) Use PutMetricData API or SDK B) Create an S3 bucket C) Modify IAM policies D) Use Route 53

Answer: A — PutMetricData sends custom metrics to CloudWatch.

Q2: Which metric is helpful for endpoints?

A) Invocation latency and error count B) S3 object size C) Route 53 DNS queries D) IAM password age

Answer: A — latency and error rate indicate endpoint health.

Q3: How do you create an alarm in CloudWatch?

A) Define a threshold and actions for the metric B) Change S3 lifecycle rules C) Modify EC2 instance type D) Update DNS records

Answer: A — alarms are configured with metric thresholds and actions.

Q4: Why validate inputs at the service boundary?

A) To prevent downstream failures and security issues B) To increase costs C) To bypass IAM D) To disable logging

Answer: A — input validation protects downstream services.

Q5: What is trace sampling?

A) Collecting a subset of traces for analysis B) A storage class in S3 C) A logging level in IAM D) A deployment strategy

Answer: A — sampling reduces overhead while keeping representative traces.

Q6: How do you view logs for a Lambda function?

A) CloudWatch Logs group for the function B) aws s3 ls C) aws iam list-users D) CloudTrail only

Answer: A — Lambda logs are in CloudWatch Logs groups.

Q7: What is an SLO?

A) Service Level Objective (target for reliability) B) Simple Logging Option C) Security Login OATH D) Storage Level Order

Answer: A — SLOs define reliability targets.

Q8: How to simulate a failure for testing alarms?

A) Send malformed input or high error rate to trigger alarms B) Delete the account C) Increase instance size D) Disable monitoring

Answer: A — intentionally inject failures to validate alarms.

Q9: What is log retention in CloudWatch?

A) How long logs are kept in CloudWatch B) The log file size limit C) The number of log streams D) The IAM policy controlling logs

Answer: A — retention settings control how long logs are stored.

Q10: How to reduce logging costs?

A) Sample, aggregate, or filter logs B) Store logs in IAM C) Disable TLS D) Increase verbosity

Answer: A — sampling and aggregation reduce storage costs.

Review Questions (previous lessons)

Q1: How do you deploy a SageMaker endpoint?

A) Create an endpoint config and then create the endpoint B) Upload to Route 53 C) Modify IAM user only D) Use CloudFront

Answer: A — endpoint config plus endpoint creation deploys the model.

Q2: What indicates a successful health check for an endpoint?

A) HTTP 200 and expected payload/schema B) 500 internal error C) AccessDenied D) Empty response

Answer: A — success is shown by a 200 and correct output.

Q3: Where are artifact files typically stored for deployment?

A) S3 B) Local laptop only C) /dev/null D) CloudWatch metrics

Answer: A — S3 is standard for artifact storage.

Q4: Why use SSE (server-side encryption) in S3?

A) To encrypt data at rest and meet security requirements B) To speed up downloads C) To reduce costs D) To disable access logging

Answer: A — SSE protects stored data.

Q5: Which factor most affects provisioning cost for endpoints?

A) Instance type and uptime B) DNS settings C) Bucket naming D) IAM passwords

Answer: A — instance size and run time drive cost.

Q6: How to validate IAM role permissions before running jobs?

A) Attempt actions or use the IAM Policy Simulator B) Change S3 lifecycle rules C) Delete the role D) Use public access

Answer: A — test permissions with simulation or actual role assumption.

Q7: What is a smoke test in deployment context?

A) A quick verification that basic functionality works after deploy B) A full regression suite C) A billing report D) A DNS update

Answer: A — smoke tests check core behavior quickly.

Q8: Why log predictions from models?

A) For debugging and drift detection B) For increasing latency intentionally C) For storing secrets D) For billing only

Answer: A — logs help detect issues and drift over time.

Q9: How to store run metadata for experiments?

A) Save a run manifest in S3 or use a tracking tool B) Post to social media C) Store in /tmp only D) Ignore metadata

Answer: A — manifests or tracking systems keep run context.

Q10: Which monitoring metric is commonly watched?

A) Error rate (and latency) B) Number of IAM users C) Bucket name length D) DNS TTL

Answer: A — error rate and latency are key observability metrics.