Transit-Layer Tensor Representation for Serverless Inference
Events section menu
Abstract: The scale of frontier machine learning (ML) systems has led to an increasing need to efficiently store and transfer model weights. This talk introduces Condense, an ultra-fast method of compressing tensor weights across ML applications to accelerate training and inference. Topics will include design objectives in tensor compression, parallel codec techniques and broader usefulness in applications such as serverless inference. In all, this talk will outline how and why compressing model tensors is a worthwhile systems optimization.
Bio: Ben Mechels is a second-year undergraduate studying computer science at the University of Minnesota. He is currently a research assistant in the Data Science Institute at the University of Chicago where he is advised by Professor Kyle Chard.
Series: See upcoming and previous presentations at CS Seminar Series.