Get in Touch

Course Outline

Foundations of Cloud Operations on AWS

  • Operational roles and responsibilities in the cloud environment.
  • AWS account structure, AWS Organizations, and multi-account strategies.
  • Core operational services: CloudWatch, CloudTrail, and AWS Config.

Infrastructure as Code and Provisioning

  • Principles of IaC and immutable infrastructure.
  • Provisioning resources with Terraform and AWS CloudFormation.
  • Managing state files, modules, and environment promotion.

CI/CD and Deployment Strategies

  • Designing CI/CD pipelines for cloud-native applications.
  • Implementing blue/green, canary, and rolling deployment strategies.
  • Automating rollback processes, health checks, and release validation.

Monitoring, Observability, and Alerting

  • Handling metrics, logs, and traces: shipping, storing, and analyzing data.
  • Utilizing CloudWatch, X-Ray, and third-party observability tools.
  • Defining Service Level Objectives (SLOs)/Service Level Indicators (SLIs), alerting policies, and on-call practices.

Security Operations and Identity Management

  • IAM best practices, least privilege principles, and cross-account access.
  • Secrets management, AWS Key Management Service (KMS), and secure parameter stores.
  • Operational security: patching strategies, vulnerability scanning, and audit trails.

Resilience, Backup, and Disaster Recovery

  • Designing for fault tolerance and high availability.
  • Backup strategies, snapshot automation, and restore procedures.
  • Disaster recovery planning and creating runbooks.

Cost Optimization and Governance

  • Cost visibility: billing, tagging, and cost allocation strategies.
  • Rightsizing, reserved instances/savings plans, and budgeting controls.
  • Governance: policies, guardrails, and automation for compliance.

Containers, Serverless, and Runtime Operations

  • Operational considerations for ECS, EKS, and Lambda.
  • Service discovery, autoscaling, and resource limits.
  • Logging, tracing, and debugging containerized workloads.

Incident Response, Playbooks, and Chaos Engineering

  • Runbook-driven incident response and postmortem practices.
  • Automating remediation and self-healing patterns.
  • Introduction to chaos experiments for validating resilience.

Hands-on Workshop: Operating a Sample Workload

  • Deploying a sample application using IaC and a CI/CD pipeline.
  • Implementing monitoring, alerts, and an automated remediation script.
  • Simulating incidents and practicing runbook-based response.

Summary and Next Steps

Requirements

  • A fundamental understanding of cloud concepts and networking principles.
  • Familiarity with the Linux command line and scripting.
  • Experience with source control systems (Git) and basic CI/CD concepts.

Target Audience

  • Cloud operations engineers.
  • Site Reliability Engineers (SREs) and platform engineers.
  • DevOps engineers and technical team leads.
 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 4800 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (2)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories