AWS AI Practitioner — Day 3: Train a Baseline Model
Learn: basics of training loop, model serialization, and artifact storage.
Hands-on: train a small classifier, save model.tar.gz to S3, and run a local inference script.
Practice question:
Q1: Where should you store immutable training artifacts for reproducibility?
A) In a local temp folder B) Overwrite the same S3 key each run C) Use a versioned S3 key or a run-specific prefix D) Email the artifact to yourself
Answer: C — store artifacts under versioned or run-specific keys for reproducibility.
Daily Practice Questions (new)
Q1: Which file format is common for small tabular datasets?
A) CSV B) MP4 C) PNG D) EXE
Answer: A — CSV is a common simple tabular format.
Q2: How do you package a model for SageMaker deployment?
A) tar gzip (model.tar.gz) containing model files B) Upload raw .py directly to endpoint C) Use a Dockerfile only with no artifact D) Put model in Route 53
Answer: A — SageMaker expects model artifacts packaged as model.tar.gz.
Q3: Which Python library is common for quick ML experiments?
A) scikit-learn (or pandas for data) B) react C) nginx D) terraform
Answer: A — scikit-learn and pandas are common for quick experiments.
Q4: How do you save a scikit-learn model locally?
A) joblib.dump(model, ‘model.joblib’) B) aws s3 cp model /bucket C) echo model > file D) use docker save
Answer: A — joblib.dump is commonly used for scikit-learn models.
Q5: Where should model artifacts be stored for deployment?
A) S3 B) local /tmp only C) Route 53 D) CloudWatch
Answer: A — S3 is used to store model artifacts for deployment.
Q6: What does a minimal inference script do?
A) Load model, accept input, return prediction B) Start a DB server C) Send emails D) Delete S3 buckets
Answer: A — minimal inference loads the model and serves predictions.
Q7: How to test inference locally?
A) Run the script with sample input and check outputs B) Change DNS records C) Create a new IAM user D) Disable versioning
Answer: A — run local inference with sample inputs to validate outputs.
Q8: Why set random seeds during training?
A) For reproducibility of results B) To speed up training C) To increase accuracy automatically D) To disable logging
Answer: A — seeds help reproducibility across runs.
Q9: What is a simple baseline model?
A) Logistic regression or decision tree B) GPU cluster C) Complex ensemble only D) DNS server
Answer: A — logistic regression/decision tree serve as simple baselines.
Q10: How to track experiments quickly?
A) Save a run manifest (hyperparams, metrics) to S3 or use tracking tool B) Use random filenames only C) Rely on memory D) Delete logs
Answer: A — run manifests or tracking tools record experiments.
Review Questions (previous lessons)
Q1: How do you list S3 buckets from the CLI?
A) aws s3 ls B) aws ec2 describe-instances C) aws iam get-role D) aws lambda list-functions
Answer: A — aws s3 ls lists buckets and objects.
Q2: What does least-privilege mean for IAM roles?
A) Grant only the minimal permissions required B) Grant admin to all users C) Share root credentials D) Remove all policies
Answer: A — least-privilege limits permissions to necessary actions.
Q3: Where can you run SageMaker notebooks?
A) SageMaker Studio or Notebook Instances B) Route 53 C) IAM only D) CloudFront
Answer: A — Studio and notebook instances provide managed notebooks.
Q4: What constitutes a basic endpoint health check?
A) Sending a test input and receiving HTTP 200 and expected schema B) Deleting the endpoint C) Changing IAM policies D) Listing S3 buckets
Answer: A — a test input with expected response shows health.
Q5: Why version model artifacts?
A) For reproducibility and rollback B) To increase costs C) To hide them D) To avoid storage
Answer: A — versioning ensures you can reproduce or roll back.
Q6: Which S3 storage class reduces cost for infrequent reads?
A) Standard-IA B) S3 Standard C) EC2 Instance Store D) Lambda
Answer: A — Standard-IA is optimized for infrequent access.
Q7: What is Batch Transform used for in SageMaker?
A) Bulk/offline scoring of datasets B) Real-time low-latency serving C) DNS management D) IAM evaluation
Answer: A — Batch Transform handles large-scale offline inference.
Q8: How can you test an IAM role’s permissions?
A) Assume the role or use IAM Policy Simulator B) Restart EC2 instances C) Change S3 lifecycle rules D) Delete the role
Answer: A — role assumption or the simulator validates permissions.
Q9: Why store run metadata for experiments?
A) For traceability and reproducibility B) To increase latency C) To reduce accuracy D) To disable logging
Answer: A — metadata helps trace experiments and compare runs.
Q10: Which result indicates a successful local inference test?
A) Correct output format and expected prediction B) 500 server error C) Empty response D) AccessDenied
Answer: A — expected predictions and correct format denote success.