CuAsmRL: optimizing GPU SASS schedules via deep reinforcement learning (CGO 2025 - Main Conference)

Who

Guoliang He, Eiko Yoneki

Track

CGO 2025 Main Conference

Time Zone

The program is currently displayed in (GMT-08:00) Pacific Time (US & Canada).

Use conference time zone: (GMT-08:00) Pacific Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Tue 4 Mar 2025 14:20 - 14:40 at Casuarina Ballroom (Level 2) - GPU & Parallelism Chair(s): Bastian Hagedorn

Abstract

Large language models (LLMs) are remarked by their sub- stantial computational requirements. To mitigate the cost, researchers develop specialized CUDA kernels, which often fuse several tensor operations to maximize the utilization of GPUs as much as possible. However, those specialized kernels may still leave performance on the table as CUDA assembly experts show that manual optimization of GPU SASS schedules can lead to better performance, and trial- and-error is largely employed to manually find the best GPU SASS schedules. In this work, we employ an automatic approach to op- timize GPU SASS schedules, which thus can be integrated into existing compiler frameworks. The key to automatic optimization is training an RL agent to mimic how human experts perform manual scheduling. To this end, we formu- late an assembly game, where RL agents can play to find the best GPU SASS schedules. The assembly game starts from a -O3 optimized SASS schedule, and the RL agents can itera- tively apply actions to mutate the current schedules. Positive rewards are generated if the mutated schedules get higher throughput by executing on GPUs. Experiments show that CuAsmRL can further improve the performance of existing specialized CUDA kernels transparently by up to 26%, and on average 9%. Moreover, it can be used as a tool to reveal potential optimization moves learned automatically.

Guoliang He

University of Cambridge

Eiko Yoneki

U. of Cambridge