Role-based roadmap
SRE roadmap
An 8-course path built around production reliability — Linux, AWS, Kubernetes, and the observability skills that keep systems alive.
The path
In order, start to finish
Each course stands alone, but the sequence is the shortest route from where most people start to the role.
- 10 hrs · 13 lessons · 12 hands-on
Linux Fundamentals
OS & Shell Mastery
SRE starts with Linux — master the OS, shell scripting, and system administration from the ground up.
Linux FundamentalsFile System & NavigationFile OperationsText ProcessingSystem AdministrationView course - 15 hrs · 31 lessons · 22 hands-on
Cloud Fundamentals
AWS Core Services
Understand the cloud infrastructure your production systems run on — IAM, networking, and core services.
Cloud IntroductionIAM & PermissionsVPC NetworkingEC2 ComputeS3 & DatabasesMonitoring & CostsView course - 20 hrs · 18 lessons · 13 hands-on
AWS Deep Dive
Production AWS Infrastructure
Deep knowledge of the AWS services SREs manage — VPC, compute scaling, storage, and CloudWatch.
VPC NetworkingCompute ScalingS3 & EBS StorageCloudWatchLoad BalancersAuto ScalingView course - 6 hrs · 5 lessons · 3 hands-on
Docker
Containerization
Understand the container runtime powering production workloads you'll be responsible for keeping alive.
Images & ContainersDockerfilesNetworking & VolumesAdvanced DockerView course - 12 hrs · 17 lessons
Kubernetes
Container Orchestration
The environment you'll be on-call for — understand every layer from pods to cluster operations and RBAC.
Architecture & ObjectsPods & DeploymentsServices & ConfigMapsStorage & NetworkingRBAC & OperationsView course - 12 hrs · 12 lessons
EKS
Managed Kubernetes on AWS
Where production workloads run — cluster management, autoscaling, and operational patterns on AWS.
EKS FundamentalsCluster SetupDeploying AppsALB IngressIRSAFargate & AutoscalingView course - 5 hrs · 5 lessons · 3 hands-on
GitHub Actions
Automation & Toil Reduction
Automate repetitive operational work — the SRE mandate is to eliminate toil with reliable pipelines.
Workflows & TriggersJobs & StepsAdvanced PatternsSecrets ManagementView course - 5 hrs · 5 lessons · 3 hands-on
Monitoring
Observability & Alerting
The most critical SRE skill — Prometheus metrics, Grafana dashboards, and alerts that page before users notice.
Monitoring FundamentalsPrometheus SetupGrafana DashboardsAlertingView course SRE
Job ready
Tools covered15
Where the time goes
Total 85+ hours · 106 lessons · 56 hands-on
What you'll be able to do
- Keep production systems reliable with deep Linux and cloud knowledge
- Understand every layer of the stack you'll be on-call for
- Set up Prometheus and Grafana to catch incidents before users do
- Automate toil with GitHub Actions to free time for reliability work
- Operate Kubernetes and EKS clusters with confidence
- Interview-ready for SRE and production engineering roles