Course Outline
Foundations of Agentic Systems in Production
- Agentic architectures: covering loops, tools, memory, and orchestration layers
- Agent lifecycle: encompassing development, deployment, and continuous operation
- Challenges associated with managing agents at production scale
Infrastructure and Deployment Models
- Implementing agents in containerized and cloud-based environments
- Scaling patterns: distinguishing between horizontal and vertical scaling, concurrency, and throttling
- Orchestration of multi-agent systems and workload distribution
Monitoring and Observability
- Essential metrics: tracking latency, success rates, memory consumption, and agent call depth
- Monitoring agent activities and mapping call graphs
- Establishing observability through Prometheus, OpenTelemetry, and Grafana
Logging, Auditing, and Compliance
- Implementing centralized logging and structured event collection
- Ensuring compliance and auditability within agentic workflows
- Creating audit trails and replay mechanisms to facilitate debugging
Performance Tuning and Resource Optimization
- Minimizing inference overhead and refining agent orchestration cycles
- Utilizing model caching and lightweight embeddings to accelerate retrieval
- Conducting load testing and stress scenario analysis for AI pipelines
Cost Control and Governance
- Identifying agent cost factors: API calls, memory usage, compute resources, and external integrations
- Monitoring agent-specific costs and establishing chargeback models
- Defining automation policies to curb agent sprawl and reduce idle resource consumption
CI/CD and Rollout Strategies for Agents
- Incorporating agent pipelines into CI/CD ecosystems
- Adopting testing, versioning, and rollback strategies for iterative agent updates
- Executing progressive rollouts and ensuring safe deployment mechanisms
Failure Recovery and Reliability Engineering
- Architecting for fault tolerance and graceful degradation
- Applying retry, timeout, and circuit breaker patterns to enhance agent reliability
- Developing incident response and post-mortem frameworks for AI operations
Capstone Project
- Construct and deploy an agentic AI system equipped with comprehensive monitoring and cost tracking
- Simulate load conditions, assess performance, and refine resource utilization
- Present the final architecture and monitoring dashboard to peers
Summary and Next Steps
Requirements
- Comprehensive knowledge of MLOps and production machine learning systems
- Proficiency in containerized deployments (Docker/Kubernetes)
- Acquaintance with cloud cost optimization and observability tools
Target Audience
- MLOps engineers
- Site Reliability Engineers (SREs)
- Engineering managers supervising AI infrastructure
Custom Corporate Training
Training solutions designed exclusively for businesses.
- Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
- Flexible Schedule: Dates and times adapted to your team's agenda.
- Format: Online (live), In-company (at your offices), or Hybrid.
Price per private group, online live training, starting from 4800 € + VAT*
Contact us for an exact quote and to hear our latest promotions
Testimonials (3)
The trainer is patient and very helpful. He knows the topic well.
CLIFFORD TABARES - Universal Leaf Philippines, Inc.
Course - Agentic AI for Business Automation: Use Cases & Integration
Good mixvof knowledge and practice
Ion Mironescu - Facultatea S.A.I.A.P.M.
Course - Agentic AI for Enterprise Applications
The mix of theory and practice and of high level and low level perspectives