Aalto computer scientists in SOSP 2024
In the annual ACM Symposium on Operating Systems Principles academic and industrial participants present research and experience papers that cover the full range of theory and practice of computer systems software.
The conference is organised on 4-6 November 2024 in Austin, Texas.
Accepted papers
Click the title to see the authors and the abstract. Link to the paper open on different website.
Authors
Marcel Wagenländer, Guo Li, Bo Zhao, Luo Mai, Peter Pietzuch
Abstract
Deep learning (DL) jobs use multi-dimensional parallelism, i.e. combining data, model, and pipeline parallelism, to use large GPU clusters efficiently. Long-running jobs may experience changes to their GPU allocation: (i) resource elasticity during training adds or removes GPUs; (ii) hardware maintenance may require redeployment on different GPUs; and (iii) GPU failures force jobs to run with fewer devices. Current DL frameworks tie jobs to a set of GPUs and thus lack support for these scenarios. In particular, they cannot change the multi-dimensional parallelism of an already-running job in an efficient and model-independent way. We describe Scalai, a state management library for DL systems that enables jobs to change their parallelism dynamically after the GPU allocation is updated at runtime. Scalai achieves this through a new abstraction, a parallelizable tensor collection (PTC), that externalizes the job state during training. After a GPU change, Scalai uses the PTC to transform the job state: the PTC repartitions the dataset state under data parallelism and exposes it to DL workers through a virtual file system; and the PTC obtains the model state as partitioned checkpoints and transforms them to reflect the new parallelization configuration. For efficiency, Scalai executes PTC transformations in parallel with minimum data movement between workers. Our experiments show that Scalai enables DL jobs to support dynamic parallelization with low overhead.
Department of Computer Science
We are an internationally-oriented community and home to world-class research in modern computer science.
School of Science
Science for tomorrow’s technology, innovations and businesses
Read more news
Major funding powers development of next-generation machine technology aimed at productivity leap in export sectors
The BEST research project is developing new types of sealing, bearing, and damping technology.
The TAIMI project builds an equal working life – a six-year consortium project seeks solutions to recruitment and skill challenges
Artificial intelligence (AI) is changing skill requirements, the population is aging, and the labor shortage is deepening. Meanwhile, the potential of international experts often remains unused in Finland. These challenges in working life are addressed by the six-year TAIMI project funded by the Strategic Research Council, and implemented by a broad consortium.
Unite! Seed Fund 2026: Call opens on 20 January 2026
Gain an early overview of the Unite! Seed Fund Call of Spring 2026. The call includes three funding lines: Student Activities, Teaching and Learning, and Research and PhD.