Get in Touch

Course Outline

Introduction to Agentic AI for Operations

  • Evolution of IT automation: moving from static runbooks to reasoning agents
  • Agent anatomy: understanding reasoning loops, tool utilization, memory, and planning
  • Determining when to automate processes versus when to maintain human oversight

Agent Frameworks and Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
  • Multi-agent architectures: supervisor, hierarchical, and swarm models
  • Comparing frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
  • Developing your first operational agent: querying monitoring, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
  • Log querying via agents: integrating with Elasticsearch, Loki, and Splunk
  • Utilizing infrastructure tools through agent actions: kubectl, Terraform, and Ansible
  • Designing secure tool interfaces with parameter validation and idempotency

Incident Response Automation

  • Automating incident triage: severity classification and routing
  • Generating root cause hypotheses and gathering supporting evidence
  • Automated remediation: executing restart, scale, rollback, and failover actions
  • Creating an incident runbook agent with progressive levels of autonomy

Safety, Guardrails, and Human-in-the-Loop

  • Action classification: distinguishing read-only, low-risk, high-risk, and destructive operations
  • Establishing approval gates and escalation policies for critical tasks
  • Implementing guardrail patterns: action allowlists, blast radius limits, and rollback guarantees
  • Maintaining audit trails and decision provenance for compliance

Multi-Agent Orchestration for Complex Incidents

  • Coordinating specialist agents: triage, diagnosis, and remediation roles
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents suggest contradictory actions
  • Simulating end-to-end major incident responses with multi-agent coordination

Observability and Evaluation

  • Tracing agent reasoning chains to facilitate debugging and auditing
  • Assessing agent decision quality: precision, recall, and time-to-resolution
  • Creating feedback loops: learning from operator overrides and operational outcomes
  • Tracking costs and managing token economics for operational agents

Production Deployment and Operations

  • Deploying agents as services: APIs, webhooks, and scheduled jobs
  • Rolling out gradual autonomy: progressing from shadow mode to full auto-remediation
  • Managing agent failures: protocols for when the agent itself encounters issues
  • Building the business case and measuring ROI for autonomous operations

Requirements

  • Practical experience with IT operations, DevOps, or SRE practices.
  • Proficiency with Python scripting and REST APIs.
  • A foundational understanding of LLM capabilities and prompt engineering.

Audience

  • SRE and DevOps engineers investigating AI-driven automation.
  • Platform engineers focused on building self-healing infrastructure.
  • IT operations leaders assessing agentic AI for incident management.
 14 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 3200 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories