Career pathway / Production systems

Production MLOps & AI Platform Engineering

Build the pipelines, release paths, and reliability practices that make AI useful after the demo.

Target roles
ML platform engineer, MLOps engineer
Learner level
Intermediate–advanced
Format
10 business days, online-first cohort
Commitment
6–8 hours each weekday; live sessions in CT
Prerequisites
Python, Git, command line, and basic software delivery experience
Certificate
Certificate of pathway completion
A platform engineer inspecting a reliable machine learning deployment pathway

The study plan

Six layers of operational confidence.

Days 01–02

Production ML systems

Understand the lifecycle and constraints before choosing platform machinery.

  • Lifecycle stages, architecture, reproducibility, and environments
  • System constraints, ownership, interfaces, and failure domains
  • Design docs, tradeoffs, and a practical platform vocabulary

Deliverable: a platform design document with boundaries, ownership, and operating assumptions.

Day 03

Data and model pipelines

Make training inputs and outputs inspectable, repeatable, and attributable.

  • Validation, feature pipelines, orchestration, and lineage
  • Artifact and version management across data and models
  • Reproducible runs, checks, and useful metadata

Deliverable: a repeatable training pipeline with validation checks and lineage notes.

Days 04–05

Packaging and deployment

Turn a model artifact into a release path that another engineer can operate.

  • Containers, batch and online serving, and service APIs
  • CI/CD, infrastructure configuration, and environment promotion
  • Release checks, rollback points, and documentation

Deliverable: an automated model release with deployment notes and a rollback procedure.

Days 06–07

Monitoring and reliability

Give operators enough signal to see when the system or its data has changed.

  • Service health, drift, data quality, and model quality
  • Alert thresholds, rollback, and incident playbooks
  • SLO thinking and the difference between noise and action

Deliverable: a monitoring dashboard and runbook for a defined production service.

Days 08–09

LLMOps and GPU-aware delivery

Apply platform judgment to newer model workloads and their cost profile.

  • Prompt and model registries, evaluation gates, and caching
  • Inference optimization, throughput, latency, and cost control
  • Capacity assumptions, privacy boundaries, and release evidence

Deliverable: a governed LLM release pipeline with evaluation and cost gates.

Day 10

Platform capstone

Design the golden path that lets a team move faster without erasing its boundaries.

  • Self-service interfaces, security boundaries, and documentation
  • SLOs, ownership, incident response, and platform adoption
  • Reference implementation walkthrough and case-study writing

Capstone: a production AI platform reference implementation with golden path, SLOs, security notes, and operator runbook.

By the end

Infrastructure that tells the truth.

You will be able to design a path from data to service, name the operational signals that matter, and make a release safer to change.

  • 01

    Design for repeatability

    Capture versions, lineage, environments, and checks so a run can be understood later.

  • 02

    Release with an exit

    Package deployments with gates, rollback paths, and enough context for an operator.

  • 03

    Monitor the whole system

    Pair service health with data quality, drift, model quality, and meaningful SLOs.

  • 04

    Build the golden path

    Make a secure, documented route to production that teams can use without guesswork.

“My capstone finally made the invisible work legible: ownership, rollback, and what happens on a bad day.”
Learner noteSamir D. · Platform engineer

Mentor profile

Rafael Ibarra

Platform engineering lead focused on reproducibility, developer experience, and the careful operational work between a model and a service.

4.9/5 learner rating760 cumulative learners taught

Before you begin

Quick answers

Do I need to be a cloud specialist?

No. The pathway introduces the relevant platform concepts while expecting comfort with Python, Git, and basic software delivery. You will practice reasoning about infrastructure without tying the work to one provider.

Does the course include a live production deployment?

The work produces a production-shaped reference implementation and release artifacts. The storefront does not promise a specific cloud environment, hosting account, or production workload.

Related pathways

Operate the right system.

Pathway 01 / Build

Generative AI Application Engineer

Start with the product, retrieval, evaluation, and application layer your platform will serve.

Explore generative AI →

Pathway 03 / Discover

Applied Machine Learning & Data Science

Understand the data and modeling workflows that production systems need to support.

Explore data science →