Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 21 hours
Course Outline
Foundations of Production-Ready Agentic Systems
- Agentic architectures: examining loops, tools, memory, and orchestration layers
- Agent lifecycle: covering development, deployment, and continuous operation
- Challenges associated with managing agents at production scale
Infrastructure and Deployment Models
- Deploying agents within containerized and cloud-based environments
- Scaling strategies: comparing horizontal and vertical scaling, concurrency, and throttling
- Orchestrating multi-agent systems and balancing workloads
Monitoring and Observability
- Essential metrics: tracking latency, success rates, memory consumption, and agent call depth
- Tracing agent activities and visualizing call graphs
- Implementing observability using Prometheus, OpenTelemetry, and Grafana
Logging, Auditing, and Compliance
- Setting up centralized logging and structured event collection
- Ensuring compliance and auditability within agentic workflows
- Creating audit trails and replay mechanisms to aid debugging
Performance Tuning and Resource Optimization
- Minimizing inference overhead and streamlining agent orchestration cycles
- Utilizing model caching and lightweight embeddings for enhanced retrieval speeds
- Conducting load testing and stress scenarios for AI pipelines
Cost Control and Governance
- Identifying agent cost drivers: API calls, memory usage, compute resources, and external integrations
- Monitoring agent-level costs and establishing chargeback models
- Implementing automation policies to prevent agent sprawl and reduce idle resource consumption
CI/CD and Rollout Strategies for Agents
- Integrating agent pipelines into CI/CD systems
- Employing testing, versioning, and rollback strategies for iterative agent updates
- Executing progressive rollouts and ensuring safe deployment mechanisms
Failure Recovery and Reliability Engineering
- Designing systems for fault tolerance and graceful degradation
- Applying retry, timeout, and circuit breaker patterns to ensure agent reliability
- Implementing incident response and post-mortem frameworks for AI operations
Capstone Project
- Develop and deploy an agentic AI system with comprehensive monitoring and cost tracking
- Simulate load, assess performance, and refine resource usage
- Present the final architecture and monitoring dashboard to peers
Summary and Next Steps
Requirements
- Proficient understanding of MLOps and production-grade machine learning systems
- Practical experience with containerized deployments using Docker/Kubernetes
- Working knowledge of cloud cost optimization and observability tooling
Target Audience
- MLOps engineers
- Site Reliability Engineers (SREs)
- Engineering managers responsible for AI infrastructure
Testimonials (3)
The trainer is patient and very helpful. He knows the topic well.
CLIFFORD TABARES - Universal Leaf Philippines, Inc.
Course - Agentic AI for Business Automation: Use Cases & Integration
Good mixvof knowledge and practice
Ion Mironescu - Facultatea S.A.I.A.P.M.
Course - Agentic AI for Enterprise Applications
The mix of theory and practice and of high level and low level perspectives