C
Posted 55m agoVandoeuvre les nancy, Grand Est, France

PhD position H/F - Assessing and Explaining Deep Learning Model Compression

MiddleOn-site (Vandoeuvre les nancy)Salary undisclosed
Required Skills
PyTorchTensorFlow
Job Description

Research Work

Scientific Context

The success of deep neural networks is often constrained by their reliance on large amounts of labeled data, which are both costly and time-consuming to obtain. Prior work has demonstrated that models can be compressed by up to 84% without any loss in performance [4]. At the same time, self-supervised learning (SSL) has emerged as a promising alternative, enabling models to learn from unlabeled data. Moreover, SSL models offer the advantage of being adaptable to multiple downstream tasks thereby reducing the cost associated with training models for auxiliary tasks. In this context, we have proposed novel approaches for SSL model compression [15], particularly in the domains of speech recognition and emotion detection, achieving results that surpass the current state of the art.

However, the increasing complexity of deep neural networks raises significant challenges in terms of interpretability and trust, especially in sensitive domains such as medicine [5] and automotive systems. Explainable Artificial Intelligence (XAI) addresses these concerns by making model decisions more transparent and understandable [6]. The THM team has made notable contributions to XAI research, particularly through the RisKa project, which applies these methods to ECG analysis [5], and more recently through the SIGN method, designed to reduce bias in model explanations [7].

This thesis builds on and extends the work conducted by both teams. It is part of CESI-LINEACT’s Human-Machine Interaction research and aligns with the Future Industry and Future City domains. It also connects with THM-KITE’s research on explainable AI for signal and image processing, requiring strong collaboration and complementary expertise from both teams.Haut du formulaire

Bas du formulaire

 

 

Thesis abstract

Artificial intelligence (AI) has seen tremendous growth, becoming so omnipresent in our daily lives that intelligent applications are now integrated into our phones, vehicles, workplaces, and even our homes. These applications typically rely on large-scale deep neural network (DNN) models, as increasing model size often correlates with improved performance. However, the deployment of such models comes with significant computing and financial costs and contributes to a substantial carbon footprint. This not only challenges the inclusivity of AI [3] but also poses environmental concerns.

The success of DNNs is often limited by the need for vast amounts of labeled data, which can be both time-consuming and expensive to obtain. Self-supervised learning (SSL) emerges as a promising alternative, enabling models to learn from unlabeled data. SSL models have the advantage of being usable in multiple downstream tasks. Currently, few studies have focused on SSL compression [15]. Among the existing research on SSL compression, [9] applies knowledge distillation (KD) to the Wav2vec acoustic model, achieving a compression rate of 4.8 times. The authors report a word error rate (WER) that is 3.62 times higher than the original model, which is quite high for our applications. In [10], genetic algorithms are proposed for the structured pruning of Wav2vec2 XLSR53, and a slight increase in the word error rate of 0.21% (1.26% relative) is reported for a 40% pruning. The authors of [11] employ symmetric linear quantization to dynamically reduce the precision of weights and activations from FLOAT32 to INT8. They also explore quantization-aware training for the BERT language model. They find that post-training dynamic quantization slightly degrades performance, while quantization-aware training achieves performance comparable to the original model. To our knowledge, there is no research yet on the quantization of self-supervised speech models or on evaluating the effects of SSL pruning on auxiliary task performance.

Alternatively, XAI methods can significantly improve the transparency and understanding of speech processing models, whether for speech recognition, emotion detection, or speech-based language translation. Among these techniques are spectrogram heatmaps, which visualize the influential regions within the spectrogram; LIME (Local Interpretable Model-agnostic Explanations), which offers local explanations by simplifying a complex model around the prediction of an individual sample; and activation maximization, a method that visualizes the input features that most strongly activate a specific neuron within the network. In the existing literature [12] recommends using attention maps within complex transformer architectures to emphasize the audio signal segments most crucial for accurate speech recognition. The authors in [13] propose a method of Neuron Activation Profiles to explain model responses to certain groups of inputs. They investigate to what extent the model learns phonemes as an intermediate representation to predict graphemes, and show that phonemes are encoded in earlier layers than graphemes.

The PhD thesis addresses these challenges by investigating the combined potential of Green AI approaches and explainable AI (XAI) techniques. While Green AI provides methods to reduce computational costs and accelerate execution, XAI enables developers and users to better understand model decisions, identify biases, and debug unexpected behaviors. This work will systematically study this integration within the specific context of automatic speech recognition (ASR). In particular, it will explore the bidirectional relationship between model compression and explainability, with the goal of guiding both the selection of optimal compression methods and the tuning of associated strategies and parameters. Such systems will be more sustainable and suitable for deployment on edge devices, such as hearing aids. Furthermore, improved transparency in inference and on-device deployment will enhance trustworthiness and align with the requirements of the European AI Act.

 

 

Expected Scientific/Technical Output

●       The research results are expected to be published in top-tier international conferences and journals.

●       The thesis will lead to the development of an efficient and explainable solution for Human-Machine vocal interaction

Laboratory Presentation: CESI LINEACT

 

CESI LINEACT (UR 7527), Laboratory for Digital Innovation for Businesses and Learning to Support the Competitiveness of Territories, anticipates and accompanies the technological mutations of sectors and services related to industry and construction. The historical proximity of CESI with companies is a determining element for our research activities. It has led us to focus our efforts on applied research close to companies and in partnership with them. A human-centered approach coupled with the use of technologies, as well as territorial networking and links with training, have enabled the construction of cross-cutting research ; it puts humans, their needs and their uses, at the center of its issues and addresses the technological angle through these contributions. Its research is organized according to two interdisciplinary scientific teams and several application areas.

— Team 1 "Learning and Innovating" mainly concerns Cognitive Sciences, Social Sciences and Management Sciences, Training Techniques and those of Innovation. The main scientific objectives are the understanding of the effects of the environment, and more particularly of situations instrumented by technical objects (platforms, prototyping workshops, immersive systems...) on learning, creativity and innovation processes.

— Team 2 "Engineering and Digital Tools" mainly concerns Digital Sciences and Engineering. The main scientific objectives focus on modeling, simulation, optimization and data analysis of cyber physical systems. Research work also focuses on decision support tools and on the study of human-system interactions in particular through digital twins coupled with virtual or augmented environments.

These two teams develop and cross their research in application areas such as

— Industry 5.0,

— Construction 4.0 and Sustainable City,

— Digital Services.

Areas supported by research platforms, mainly those in Rouen dedicated to Factory 5.0 and those in Nanterre dedicated to Factory 5.0 and Construction 4.0.

 

Website of CESI LINEACT: https://lineact.cesi.fr/

 

Laboratory Presentation: Centre of Competence for Information Technology (KITE)

The interdisciplinary Centre of Competence for Information Technology (KITE) is involved in applied research into and development of artificial intelligence, machine learning and IT security. It combines expertise from the Departments of Health Sciences (GES), Mathematics, Natural Sciences and Data Processing (MND) and Mathematics, Natural Sciences and Informatics (MNI). The informational approaches and modelling procedures range from medicine (e-health) and biology (bio-informatics), through industry (smart factories), to digital humanities.

Website of KITE:  https://www.thm.de/kompetenzzentren/en/kite/profile.html

Skills

Scientific and technical skills in one or more of the following areas:

— Strong mathematical skills, particularly in convex and non-convex optimization,

matrix theory, and probability.

— Proficiency in key AI techniques, experience with PyTorch or TensorFlow would be

appreciated.

— Strong programming skills.

 

Soft skills:

— Good level of written and spoken English.

— Ability to work independently, with initiative and curiosity.

— Ability to work in a team and maintain good interpersonal relations.

— Attention to detail and rigor.

Organisation

It is a joint French-German PhD thesis that will be registered in ENSAM (France) and in THM (Germany). Double PhD degree will be delivered conditioned on a successful PhD thesis.

Location:   18 months in France (CESI LINEACT – Campus de Nancy, France)

                  18 months in Germany (KITE – Friedberg, Germany)

Starting date: A soon as possible

Duration: 3 years

 

Supervisors

Leila BEN LETAIFA. PhD co-supervisor. CESI LINEACT (France)

Yuehua DING. PhD Supervisor. CESI LINEACT (France)

Michael GUCKERT. PhD Supervisor. THM (Germany)
Christin SEIFERT.  PhD co-supervisor. THM (Germany)

Bibliography

[1] Roy Schwartz et al., “Green AI,” Communications of the ACM, vol. 63, no. 12, 2020, pp. 54–63.

[2] Jingjing Xu et al., “A Survey on Green Deep Learning,” arXiv preprint arXiv:2111.05193, 2021.

[3] Girish Sastry et al., “Computing Power and the Governance of Artificial Intelligence,” arXiv preprint arXiv:2402.08797, 2024.

[4] Leila Ben Letaifa and Jean-Luc Rouas, “Transformer Model Compression for End-to-End Speech Recognition on Mobile Devices,” in Proceedings of the 30th European Signal Processing Conference (EUSIPCO), IEEE, 2022, pp. 439–443.

[5] Nils Gumpfer et al., “Towards Trustworthy AI in Cardiology: A Comparative Analysis of Explainable AI Methods for Electrocardiogram Interpretation,” in International Conference on Artificial Intelligence in Medicine, Springer, 2024, pp. 350–361.

[6] Sajid Ali et al., “Explainable Artificial Intelligence (XAI): What We Know and What Is Left to Attain Trustworthy Artificial Intelligence,” Information Fusion, vol. 99, 2023, p. 101805. doi:10.1016/j.inffus.2023.101805.

[7] Nils Gumpfer et al., “SIGNed Explanations: Unveiling Relevant Features by Reducing Bias,” Information Fusion, vol. 99, 2023, p. 101883.

[8] Leila Ben Letaifa and Jean-Luc Rouas, “Variable Scale Pruning for Transformer Model Compression in End-to-End Speech Recognition,” Algorithms, Aug. 2023.

[9] Zilun Peng et al., “Shrinking BigFoot: Reducing Wav2Vec 2.0 Footprint,” arXiv preprint arXiv:2103.15760, 2021.

[10] Oswaldo Ludwig and Tom Claes, “Compressing Wav2Vec 2.0 for Embedded Applications,” in Proceedings of the 33rd IEEE International Workshop on Machine Learning for Signal Processing (MLSP), IEEE, 2023, pp. 1–6.

[11] Ofir Zafrir et al., “Q8BERT: Quantized 8-bit BERT,” in Proceedings of the Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing (EMC2-NIPS), IEEE, 2019, pp. 36–39.

[12] Tianli Sun et al., “Explainability of Speech Recognition Transformers via Gradient-Based Attention Visualization,” IEEE Transactions on Multimedia, vol. 26, 2024, pp. 1395–1406. doi:10.1109/TMM.2023.3282488.

[13] Andreas Krug, René Knaebel, and Sebastian Stober, “Neuron Activation Profiles for Interpreting Convolutional Speech Recognition Models,” in NeurIPS Workshop on Interpretability and Robustness in Audio, Speech, and Language (IRASL), 2018.

[14] J.-L. Rouas et Leila ben Letaifa., “Structured Pruning for Efficient Systolic Array Accelerated Cascade Speech-to-Text Translation,” in Proceedings of Interspeech, 2025.

[15] Mouaad Oujabour, Leila Ben Letaifa, Jean-François Dollinger, and Jean-Luc Rouas, “Adaptive Compression of Supervised and Self-Supervised Models for Green Speech Recognition,” in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025.

Similar Openings in AI & Machine Learning

View all in category