Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Deployment

  • Encapsulating models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Best practices for versioning and storage

Serving Models on Kubernetes

  • Introduction to inference servers
  • Deployment of TensorFlow Serving and TorchServe
  • Configuration of model endpoints

Optimizing Inference Performance

  • Implementing batching strategies
  • Managing concurrent requests effectively
  • Tuning for optimal latency and throughput

Autoscaling Machine Learning Workloads

  • Horizontal Pod Autoscaler (HPA)
  • Vertical Pod Autoscaler (VPA)
  • Kubernetes Event-Driven Autoscaling (KEDA)

GPU Allocation and Resource Management

  • Setting up GPU-enabled nodes
  • Overview of the NVIDIA device plugin
  • Defining resource requests and limits for ML workloads

Model Rollout and Release Methodologies

  • Blue/green deployment patterns
  • Canary rollout techniques
  • A/B testing for rigorous model evaluation

Monitoring and Observability for Production ML

  • Key metrics for inference workloads
  • Standard logging and tracing procedures
  • Creation of dashboards and alerting systems

Security and Reliability Enhancements

  • Protecting model endpoints
  • Implementing network policies and access controls
  • Maintaining high availability

Summary and Future Directions

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Working knowledge of Kubernetes fundamentals

Target Audience

  • ML Engineers
  • DevOps Engineers
  • Platform Engineering Teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories