Model compression
Model compression refers to a set of techniques aimed at reducing the size and computational requirements of machine learning models while maintaining their performance. It enables deployment of models on resource-constrained devices and improves inference efficiency.