Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- Evolution of IT automation: moving from static runbooks to reasoning agents
- Agent anatomy: understanding the reasoning loop, tool utilisation, memory, and planning
- Determining when to automate versus when to retain human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent structures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom-built agents
- Constructing your first operational agent: querying, diagnosing, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-driven log queries: integration with Elasticsearch, Loki, and Splunk
- Infrastructure automation: using kubectl, Terraform, and Ansible via agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated triage: severity classification and intelligent routing
- Generating root cause hypotheses and gathering supporting evidence
- Automated remediation: executing restart, scale, rollback, and failover actions
- Creating an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Establishing approval gates and escalation policies for critical operations
- Implementing guardrails: action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating end-to-end major incident responses using multiple agents
Observability and Evaluation
- Tracing agent reasoning chains to facilitate debugging and auditing
- Evaluating decision quality: precision, recall, and time-to-resolution
- Establishing feedback loops: learning from operator overrides and final outcomes
- Tracking costs and analysing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled jobs
- Gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
- Handling agent failures: runbooks for when the agent itself malfunctions
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience with IT operations, DevOps, or SRE methodologies.
- Proficiency in Python scripting and REST APIs.
- Foundational knowledge of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers focused on building self-healing infrastructure.
- IT operations leaders assessing Agentic AI for incident management.
14 Hours