Get in Touch

Course Outline

Overview of Biren GPU Architecture

  • Biren introduction and key use cases
  • Hardware configuration: cores, memory, and compute clusters
  • Comparative analysis with NVIDIA and AMD GPUs

Configuring the Biren Programming Environment

  • Installation of the Biren SDK and runtime components
  • Understanding the toolchain and compiler models
  • Basic project organization and build workflows

GPU Programming Using the Biren Stack

  • Thread and block modeling
  • Memory management and data transfer mechanisms
  • Kernel development and launch strategies

Migration from CUDA to Biren

  • Techniques for translating CUDA code
  • Mapping and adapting common APIs
  • Hands-on labs for code conversion and practice

Debugging and Profiling

  • Utilizing Biren’s debugger and profiler tools
  • Pinpointing performance bottlenecks
  • Optimizing memory access patterns

Optimization Strategies

  • Thread scheduling and instruction pipelining
  • Loop unrolling and effective shared memory utilization
  • Advanced kernel tuning to maximize throughput

Case Studies and Application Examples

  • Training models using Biren accelerators
  • Porting and profiling vision or NLP models
  • Performance comparison against CUDA/NVIDIA platforms

Conclusion and Next Steps

Requirements

  • A solid grasp of GPU architecture and parallel processing concepts
  • Prior experience with CUDA, OpenCL, or comparable GPU programming frameworks
  • Familiarity with deep learning frameworks like PyTorch or TensorFlow

Target Audience

  • HPC developers
  • AI infrastructure engineers
  • Performance optimization specialists
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories