Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
The AI Observability Landscape
- Shifting from dashboards to conversations: the move toward AI-augmented observability
- LLM capabilities relevant to observability: summarization, reasoning, and pattern matching
- Architecture patterns for embedding AI into existing observability stacks
Natural Language Telemetry Querying
- Text-to-PromQL: translating natural language into monitoring queries
- NL querying for Elasticsearch, OpenSearch, and Loki log stores
- SQL generation from natural language for structured telemetry
- Developing a query assistant agent with tool use and context awareness
LLM-Powered Log Analysis
- Automated log parsing and structuring using LLMs
- Anomaly detection in log streams via embedding similarity
- Log clustering and pattern discovery at scale
- Generating human-readable explanations from raw log sequences
Intelligent Alerting and Incident Enrichment
- Alert correlation and deduplication with semantic understanding
- Automated gathering of incident context from runbooks, past incidents, and documentation
- Smart alert routing based on content understanding and team expertise
- Mitigating alert fatigue through AI-driven noise reduction
AI-Assisted Root Cause Analysis
- Hypothesis generation from multi-source telemetry correlation
- Evidence chaining: linking symptoms across metrics, logs, and traces
- Guided troubleshooting through interactive AI diagnosis sessions
- Building a root cause analysis agent with progressive investigation capabilities
Automated Incident Response and Communication
- Generating incident summaries and status updates from telemetry data
- Automated postmortem drafting with timeline reconstruction
- Tailoring stakeholder communication for both technical and executive audiences
- Suggesting runbooks and providing automated remediation recommendations
Machine Learning for Observability
- Time-series forecasting for capacity planning and anomaly prediction
- Foundation models for zero-shot anomaly detection on metrics
- Embedding-based service dependency mapping and topology discovery
- Training and deploying lightweight ML models alongside observability pipelines
Production Deployment and Ethics
- Latency and cost considerations for real-time AI observability
- Data privacy: ensuring LLMs do not leak sensitive telemetry data
- Human oversight: determining when AI diagnosis requires operator validation
- Measuring impact: tracking MTTD, MTTR, and on-call experience metrics
Requirements
- Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.
- Familiarity with log management and metrics concepts.
- Basic Python scripting skills for data processing.
Target Audience
- SRE and observability engineers adopting AI-enhanced tooling.
- Platform engineers developing next-generation monitoring pipelines.
- DevOps leads evaluating the integration of LLMs into incident workflows.
14 Hours
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 3200 € + VAT*
Contact us for an exact quote and to hear our latest promotions