Get in Touch

Course Outline

Foundations of Agentic Systems in Production

  • Agentic architectures: covering loops, tools, memory, and orchestration layers
  • Agent lifecycle: encompassing development, deployment, and continuous operation
  • Challenges associated with managing agents at production scale

Infrastructure and Deployment Models

  • Implementing agents in containerized and cloud-based environments
  • Scaling patterns: distinguishing between horizontal and vertical scaling, concurrency, and throttling
  • Orchestration of multi-agent systems and workload distribution

Monitoring and Observability

  • Essential metrics: tracking latency, success rates, memory consumption, and agent call depth
  • Monitoring agent activities and mapping call graphs
  • Establishing observability through Prometheus, OpenTelemetry, and Grafana

Logging, Auditing, and Compliance

  • Implementing centralized logging and structured event collection
  • Ensuring compliance and auditability within agentic workflows
  • Creating audit trails and replay mechanisms to facilitate debugging

Performance Tuning and Resource Optimization

  • Minimizing inference overhead and refining agent orchestration cycles
  • Utilizing model caching and lightweight embeddings to accelerate retrieval
  • Conducting load testing and stress scenario analysis for AI pipelines

Cost Control and Governance

  • Identifying agent cost factors: API calls, memory usage, compute resources, and external integrations
  • Monitoring agent-specific costs and establishing chargeback models
  • Defining automation policies to curb agent sprawl and reduce idle resource consumption

CI/CD and Rollout Strategies for Agents

  • Incorporating agent pipelines into CI/CD ecosystems
  • Adopting testing, versioning, and rollback strategies for iterative agent updates
  • Executing progressive rollouts and ensuring safe deployment mechanisms

Failure Recovery and Reliability Engineering

  • Architecting for fault tolerance and graceful degradation
  • Applying retry, timeout, and circuit breaker patterns to enhance agent reliability
  • Developing incident response and post-mortem frameworks for AI operations

Capstone Project

  • Construct and deploy an agentic AI system equipped with comprehensive monitoring and cost tracking
  • Simulate load conditions, assess performance, and refine resource utilization
  • Present the final architecture and monitoring dashboard to peers

Summary and Next Steps

Requirements

  • Comprehensive knowledge of MLOps and production machine learning systems
  • Proficiency in containerized deployments (Docker/Kubernetes)
  • Acquaintance with cloud cost optimization and observability tools

Target Audience

  • MLOps engineers
  • Site Reliability Engineers (SREs)
  • Engineering managers supervising AI infrastructure
 21 Hours

Custom Corporate Training

Training solutions designed exclusively for businesses.

  • Customized Content: We adapt the syllabus and practical exercises to the real goals and needs of your project.
  • Flexible Schedule: Dates and times adapted to your team's agenda.
  • Format: Online (live), In-company (at your offices), or Hybrid.
Investment

Price per private group, online live training, starting from 4800 € + VAT*

Contact us for an exact quote and to hear our latest promotions

Testimonials (3)

Provisional Upcoming Courses (Contact Us For More Information)

Related Categories