Role-based roadmap

SRE roadmap

An 8-course path built around production reliability — Linux, AWS, Kubernetes, and the observability skills that keep systems alive.

85+ hours 8 courses 106 lessons 56 hands-on

The path

In order, start to finish

Each course stands alone, but the sequence is the shortest route from where most people start to the role.

  1. Linux Fundamentals

    OS & Shell Mastery

    10 hrs · 13 lessons · 12 hands-on

    SRE starts with Linux — master the OS, shell scripting, and system administration from the ground up.

    Linux FundamentalsFile System & NavigationFile OperationsText ProcessingSystem Administration
    View course
  2. Cloud Fundamentals

    AWS Core Services

    15 hrs · 31 lessons · 22 hands-on

    Understand the cloud infrastructure your production systems run on — IAM, networking, and core services.

    Cloud IntroductionIAM & PermissionsVPC NetworkingEC2 ComputeS3 & DatabasesMonitoring & Costs
    View course
  3. AWS Deep Dive

    Production AWS Infrastructure

    20 hrs · 18 lessons · 13 hands-on

    Deep knowledge of the AWS services SREs manage — VPC, compute scaling, storage, and CloudWatch.

    VPC NetworkingCompute ScalingS3 & EBS StorageCloudWatchLoad BalancersAuto Scaling
    View course
  4. Docker

    Containerization

    6 hrs · 5 lessons · 3 hands-on

    Understand the container runtime powering production workloads you'll be responsible for keeping alive.

    Images & ContainersDockerfilesNetworking & VolumesAdvanced Docker
    View course
  5. Kubernetes

    Container Orchestration

    12 hrs · 17 lessons

    The environment you'll be on-call for — understand every layer from pods to cluster operations and RBAC.

    Architecture & ObjectsPods & DeploymentsServices & ConfigMapsStorage & NetworkingRBAC & Operations
    View course
  6. EKS

    Managed Kubernetes on AWS

    12 hrs · 12 lessons

    Where production workloads run — cluster management, autoscaling, and operational patterns on AWS.

    EKS FundamentalsCluster SetupDeploying AppsALB IngressIRSAFargate & Autoscaling
    View course
  7. GitHub Actions

    Automation & Toil Reduction

    5 hrs · 5 lessons · 3 hands-on

    Automate repetitive operational work — the SRE mandate is to eliminate toil with reliable pipelines.

    Workflows & TriggersJobs & StepsAdvanced PatternsSecrets Management
    View course
  8. Monitoring

    Observability & Alerting

    5 hrs · 5 lessons · 3 hands-on

    The most critical SRE skill — Prometheus metrics, Grafana dashboards, and alerts that page before users notice.

    Monitoring FundamentalsPrometheus SetupGrafana DashboardsAlerting
    View course
  9. SRE

    Job ready

Tools covered15

LinuxBashAWSEC2VPCCloudWatchDockerKuberneteskubectlEKSGitHub ActionsPrometheusGrafanaIAMAuto Scaling

Where the time goes

Linux Fundamentals10h
Cloud Fundamentals15h
AWS Deep Dive20h
Docker6h
Kubernetes12h
EKS12h
GitHub Actions5h
Monitoring5h

Total 85+ hours · 106 lessons · 56 hands-on

What you'll be able to do

  • Keep production systems reliable with deep Linux and cloud knowledge
  • Understand every layer of the stack you'll be on-call for
  • Set up Prometheus and Grafana to catch incidents before users do
  • Automate toil with GitHub Actions to free time for reliability work
  • Operate Kubernetes and EKS clusters with confidence
  • Interview-ready for SRE and production engineering roles