Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- Evolution of IT automation: moving from static runbooks to reasoning agents
- Agent anatomy: understanding reasoning loops, tool utilization, memory, and planning
- Determining when to automate processes versus when to maintain human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Comparing frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Developing your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to APIs for Prometheus, Grafana, Datadog, and PagerDuty
- Log querying via agents: integrating with Elasticsearch, Loki, and Splunk
- Utilizing infrastructure tools through agent actions: kubectl, Terraform, and Ansible
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automating incident triage: severity classification and routing
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation: executing restart, scale, rollback, and failover actions
- Creating an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop
- Action classification: distinguishing read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical tasks
- Implementing guardrail patterns: action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents suggest contradictory actions
- Simulating end-to-end major incident responses with multi-agent coordination
Observability and Evaluation
- Tracing agent reasoning chains to facilitate debugging and auditing
- Assessing agent decision quality: precision, recall, and time-to-resolution
- Creating feedback loops: learning from operator overrides and operational outcomes
- Tracking costs and managing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled jobs
- Rolling out gradual autonomy: progressing from shadow mode to full auto-remediation
- Managing agent failures: protocols for when the agent itself encounters issues
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience with IT operations, DevOps, or SRE practices.
- Proficiency with Python scripting and REST APIs.
- A foundational understanding of LLM capabilities and prompt engineering.
Audience
- SRE and DevOps engineers investigating AI-driven automation.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders assessing agentic AI for incident management.
14 Hours
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 3200 € + VAT*
Contact us for an exact quote and to hear our latest promotions