Distilling the knowledge in a neural network. Hinton et al. arXiv:1503.02531
Learning from Noisy Labels with Distillation. Li, Yuncheng et al. ICCV 2017
Training Deep Neural Networks in Generations:A More Tolerant Teacher Educates Better Students. arXiv:1805.05551
Learning Metrics from Teachers: Compact Networks for Image Embedding. Yu, Lu et al. CVPR 2019
Relational Knowledge Distillation. Park, Wonpyo et al. CVPR 2019
On Knowledge Distillation from Complex Networks for Response Prediction. Arora, Siddhartha et al. NAACL 2019
On the Efficacy of Knowledge Distillation. Cho, Jang Hyun & Hariharan, Bharath. arXiv:1910.01348. ICCV 2019
Revisit Knowledge Distillation: a Teacher-free Framework (Revisiting Knowledge Distillation via Label Smoothing Regularization). Yuan, Li et al. CVPR 2020 [code]
Improved Knowledge Distillation via Teacher Assistant: Bridging the Gap Between Student and Teacher. Mirzadeh et al. arXiv:1902.03393
Ensemble Distribution Distillation. ICLR 2020
Noisy Collaboration in Knowledge Distillation. ICLR 2020
On Compressing U-net Using Knowledge Distillation. arXiv:1812.00249
Self-training with Noisy Student improves ImageNet classification. Xie, Qizhe et al.(Google) CVPR 2020
Variational Student: Learning Compact and Sparser Networks in Knowledge Distillation Framework. AAAI 2020
Preparing Lessons: Improve Knowledge Distillation with Better Supervision. arXiv:1911.07471
Adaptive Regularization of Labels. arXiv:1908.05474
Positive-Unlabeled Compression on the Cloud. Xu, Yixing et al. (HUAWEI) NeurIPS 2019
Snapshot Distillation: Teacher-Student Optimization in One Generation. Yang, Chenglin et al. CVPR 2019
QUEST: Quantized embedding space for transferring knowledge. Jain, Himalaya et al. arXiv:2020
Conditional teacher-student learning. Z. Meng et al. ICASSP 2019
Subclass Distillation. Müller, Rafael et al. arXiv:2002.03936
MarginDistillation: distillation for margin-based softmax. Svitov, David & Alyamkin, Sergey. arXiv:2003.02586
An Embarrassingly Simple Approach for Knowledge Distillation. Gao, Mengya et al. MLR 2018
Sequence-Level Knowledge Distillation. Kim, Yoon & Rush, Alexander M. arXiv:1606.07947
Boosting Self-Supervised Learning via Knowledge Transfer. Noroozi, Mehdi et al. CVPR 2018
Meta Pseudo Labels. Pham, Hieu et al. ICML 2020 [code]
Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation from a Blackbox Model. CVPR 2020 [code]
Distilled Binary Neural Network for Monaural Speech Separation. Chen Xiuyi et al. IJCNN 2018
Teacher-Class Network: A Neural Network Compression Mechanism. Malik et al. arXiv:2004.03281
Deeply-supervised knowledge synergy. Sun, Dawei et al. CVPR 2019
What it Thinks is Important is Important: Robustness Transfers through Input Gradients. Chan, Alvin et al. CVPR 2020
Triplet Loss for Knowledge Distillation. Oki, Hideki et al. IJCNN 2020
Role-Wise Data Augmentation for Knowledge Distillation. ICLR 2020 [code]
Distilling Spikes: Knowledge Distillation in Spiking Neural Networks. arXiv:2005.00288
Improved Noisy Student Training for Automatic Speech Recognition. Park et al. arXiv:2005.09629
Learning from a Lightweight Teacher for Efficient Knowledge Distillation. Yuang Liu et al. arXiv:2005.09163
ResKD: Residual-Guided Knowledge Distillation. Li, Xuewei et al. arXiv:2006.04719
Distilling Effective Supervision from Severe Label Noise. Zhang, Zizhao. et al. CVPR 2020 [code]
Local Correlation Consistency for Knowledge Distillation. ECCV 2020
Few-Shot Class-Incremental Learning. Tao, Xiaoyu et al. CVPR 2020
Semantic Relation Preserving Knowledge Distillation for Image-to-Image Translation. ECCV 2020
Interpretable Foreground Object Search As Knowledge Distillation. ECCV 2020
Improving Knowledge Distillation via Category Structure. ECCV 2020
Few-Shot Class-Incremental Learning via Relation Knowledge Distillation. Dong, Songlin et al. AAAI 2021
Complementary Relation Contrastive Distillation. Zhu, Jinguo et al. CVPR 2021
Information Theoretic Representation Distillation. Miles et al. BMVC 2022 [code]
Privileged Information
Learning using privileged information: similarity control and knowledge transfer. Vapnik, Vladimir and Rauf, Izmailov. MLR 2015
Unifying distillation and privileged information. Lopez-Paz, David et al. ICLR 2016
Model compression via distillation and quantization. Polino, Antonio et al. ICLR 2018
KDGAN:Knowledge Distillation with Generative Adversarial Networks. Wang, Xiaojie. NeurIPS 2018
Efficient Video Classification Using Fewer Frames. Bhardwaj, Shweta et al. CVPR 2019
Retaining privileged information for multi-task learning. Tang, Fengyi et al. KDD 2019
A Generalized Meta-loss function for regression and classification using privileged information. Asif, Amina et al. arXiv:1811.06885
Private Knowledge Transfer via Model Distillation with Generative Adversarial Networks. Gao, Di & Zhuo, Cheng. AAAI 2020
Privileged Knowledge Distillation for Online Action Detection. Zhao, Peisen et al. cvpr 2021
Adversarial Distillation for Learning with Privileged Provisions. Wang, Xiaojie et al. TPAMI 2019
KD + GAN
Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks. Xu, Zheng et al. arXiv:1709.00513
KTAN: Knowledge Transfer Adversarial Network. Liu, Peiye et al. arXiv:1810.08126
KDGAN:Knowledge Distillation with Generative Adversarial Networks. Wang, Xiaojie. NeurIPS 2018
Adversarial Learning of Portable Student Networks. Wang, Yunhe et al. AAAI 2018
Adversarial Network Compression. Belagiannis et al. ECCV 2018
Cross-Modality Distillation: A case for Conditional Generative Adversarial Networks. ICASSP 2018
Adversarial Distillation for Efficient Recommendation with External Knowledge. TOIS 2018
Training student networks for acceleration with conditional adversarial networks. Xu, Zheng et al. BMVC 2018
DAFL:Data-Free Learning of Student Networks. Chen, Hanting et al. ICCV 2019
MEAL: Multi-Model Ensemble via Adversarial Learning. Shen, Zhiqiang et al. AAAI 2019
Knowledge Distillation with Adversarial Samples Supporting Decision Boundary. Heo, Byeongho et al. AAAI 2019
Exploiting the Ground-Truth: An Adversarial Imitation Based Knowledge Distillation Approach for Event Detection. Liu, Jian et al. AAAI 2019
Adversarially Robust Distillation. Goldblum, Micah et al. AAAI 2020
GAN-Knowledge Distillation for one-stage Object Detection. Hong, Wei et al. arXiv:1906.08467
Lifelong GAN: Continual Learning for Conditional Image Generation. Kundu et al. arXiv:1908.03884
Compressing GANs using Knowledge Distillation. Aguinaldo, Angeline et al. arXiv:1902.00159
Bootstrap Your Own Latent: A New Approach to Self-Supervised Learning. Grill et al. arXiv:2006.07733 [code]
Unpaired Learning of Deep Image Denoising. Wu, Xiaohe et al. arXiv:2008.13711 [code]
SSKD: Self-Supervised Knowledge Distillation for Cross Domain Adaptive Person Re-Identification. Yin, Junhui et al. arXiv:2009.05972
Introspective Learning by Distilling Knowledge from Online Self-explanation. Gu, Jindong et al. ACCV 2020
Robust Pre-Training by Adversarial Contrastive Learning. Jiang, Ziyu et al. NeurIPS 2020 [code]
CompRess: Self-Supervised Learning by Compressing Representations. Koohpayegani et al. NeurIPS 2020 [code]
Big Self-Supervised Models are Strong Semi-Supervised Learners. Che, Ting et al. NeurIPS 2020 [code]
Rethinking Pre-training and Self-training. Zoph, Barret et al. NeurIPS 2020 [code]
ISD: Self-Supervised Learning by Iterative Similarity Distillation. Tejankar et al. cvpr 2021 [code]
Momentum^2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning. Li, Zeming et al. arXiv:2101.07525
Beyond Self-Supervision: A Simple Yet Effective Network Distillation Alternative to Improve Backbones. Cui, Cheng et al. arXiv:2103.05959
Distilling Audio-Visual Knowledge by Compositional Contrastive Learning. Chen, Yanbei et al. CVPR 2021
DisCo: Remedy Self-supervised Learning on Lightweight Models with Distilled Contrastive Learning. Gao, Yuting et al. arXiv:2104.09124
Self-Ensembling Contrastive Learning for Semi-Supervised Medical Image Segmentation. Xiang, Jinxi et al. arXiv:2105.12924
Semi-Supervised Semantic Segmentation with Cross Pseudo Supervision. Chen, Xiaokang et al. CPVR 2021
Adversarial Knowledge Transfer from Unlabeled Data. Gupta et al. ACM-MM 2020 code
Multi-teacher and Ensemble KD
Learning from Multiple Teacher Networks. You, Shan et al. KDD 2017
Learning with single-teacher multi-student. You, Shan et al. AAAI 2018
Knowledge distillation by on-the-fly native ensemble. Lan, Xu et al. NeurIPS 2018
Semi-Supervised Knowledge Transfer for Deep Learning from Private Training Data. ICLR 2017
Knowledge Adaptation: Teaching to Adapt. Arxiv:1702.02052
Deep Model Compression: Distilling Knowledge from Noisy Teachers. Sau, Bharat Bhusan et al. arXiv:1610.09650
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Tarvainen, Antti and Valpola, Harri. NeurIPS 2017
Born-Again Neural Networks. Furlanello, Tommaso et al. ICML 2018
Deep Mutual Learning. Zhang, Ying et al. CVPR 2018
Collaborative learning for deep neural networks. Song, Guocong and Chai, Wei. NeurIPS 2018
Data Distillation: Towards Omni-Supervised Learning. Radosavovic, Ilija et al. CVPR 2018
Multilingual Neural Machine Translation with Knowledge Distillation. ICLR 2019
Unifying Heterogeneous Classifiers with Distillation. Vongkulbhisal et al. CVPR 2019
Distilled Person Re-Identification: Towards a More Scalable System. Wu, Ancong et al. CVPR 2019
Diversity with Cooperation: Ensemble Methods for Few-Shot Classification. Dvornik, Nikita et al. ICCV 2019
Model Compression with Two-stage Multi-teacher Knowledge Distillation for Web Question Answering System. Yang, Ze et al. WSDM 2020
FEED: Feature-level Ensemble for Knowledge Distillation. Park, SeongUk and Kwak, Nojun. AAAI 2020
Stochasticity and Skip Connection Improve Knowledge Transfer. Lee, Kwangjin et al. ICLR 2020
Online Knowledge Distillation with Diverse Peers. Chen, Defang et al. AAAI 2020
Hydra: Preserving Ensemble Diversity for Model Distillation. Tran, Linh et al. arXiv:2001.04694
Distilled Hierarchical Neural Ensembles with Adaptive Inference Cost. Ruiz, Adria et al. arXv:2003.01474
Distilling Knowledge from Ensembles of Acoustic Models for Joint CTC-Attention End-to-End Speech Recognition. Gao, Yan et al. arXiv:2005.09310
Large-Scale Few-Shot Learning via Multi-Modal Knowledge Discovery. ECCV 2020
Collaborative Learning for Faster StyleGAN Embedding. Guan, Shanyan et al. arXiv:2007.01758
Temporal Self-Ensembling Teacher for Semi-Supervised Object Detection. Chen, Cong et al. IEEE 2020 [code]
Dual-Teacher: Integrating Intra-domain and Inter-domain Teachers for Annotation-efficient Cardiac Segmentation. MICCAI 2020
Joint Progressive Knowledge Distillation and Unsupervised Domain Adaptation. Nguyen-Meidine et al. WACV 2020
Semi-supervised Learning with Teacher-student Network for Generalized Attribute Prediction. Shin, Minchul et al. ECCV 2020
Knowledge Distillation for Multi-task Learning. Li, WeiHong & Bilen, Hakan. arXiv:2007.06889 [project]
Amalgamating Knowledge towards Comprehensive Classification. Shen, Chengchao et al. AAAI 2019
Amalgamating Filtered Knowledge : Learning Task-customized Student from Multi-task Teachers. Ye, Jingwen et al. IJCAI 2019
Knowledge Amalgamation from Heterogeneous Networks by Common Feature Learning. Luo, Sihui et al. IJCAI 2019
Student Becoming the Master: Knowledge Amalgamation for Joint Scene Parsing, Depth Estimation, and More. Ye, Jingwen et al. CVPR 2019
Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation. ICCV 2019
Data-Free Knowledge Amalgamation via Group-Stack Dual-GAN. CVPR 2020
Cross-modal / DA / Incremental Learning
SoundNet: Learning Sound Representations from Unlabeled Video SoundNet Architecture. Aytar, Yusuf et al. NeurIPS 2016
Cross Modal Distillation for Supervision Transfer. Gupta, Saurabh et al. CVPR 2016
Emotion recognition in speech using cross-modal transfer in the wild. Albanie, Samuel et al. ACM MM 2018
Through-Wall Human Pose Estimation Using Radio Signals. Zhao, Mingmin et al. CVPR 2018
Compact Trilinear Interaction for Visual Question Answering. Do, Tuong et al. ICCV 2019
Cross-Modal Knowledge Distillation for Action Recognition. Thoker, Fida Mohammad and Gall, Juerge. ICIP 2019
Learning to Map Nearly Anything. Salem, Tawfiq et al. arXiv:1909.06928
Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval. Liu, Qing et al. ICCV 2019
UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task Distillation. Kundu et al. ICCV 2019
CrDoCo: Pixel-level Domain Transfer with Cross-Domain Consistency. Chen, Yun-Chun et al. CVPR 2019
XD:Cross lingual Knowledge Distillation for Polyglot Sentence Embeddings. ICLR 2020
Effective Domain Knowledge Transfer with Soft Fine-tuning. Zhao, Zhichen et al. arXiv:1909.02236
ASR is all you need: cross-modal distillation for lip reading. Afouras et al. arXiv:1911.12747v1
Knowledge distillation for semi-supervised domain adaptation. arXiv:1908.07355
Domain Adaptation via Teacher-Student Learning for End-to-End Speech Recognition. Meng, Zhong et al. arXiv:2001.01798
Cluster Alignment with a Teacher for Unsupervised Domain Adaptation. ICCV 2019
Attention Bridging Network for Knowledge Transfer. Li, Kunpeng et al. ICCV 2019
Unpaired Multi-modal Segmentation via Knowledge Distillation. Dou, Qi et al. arXiv:2001.03111
Multi-source Distilling Domain Adaptation. Zhao, Sicheng et al. arXiv:1911.11554
Creating Something from Nothing: Unsupervised Knowledge Distillation for Cross-Modal Hashing. Hu, Hengtong et al. CVPR 2020
Improving Semantic Segmentation via Self-Training. Zhu, Yi et al. arXiv:2004.14960
Speech to Text Adaptation: Towards an Efficient Cross-Modal Distillation. arXiv:2005.08213
Joint Progressive Knowledge Distillation and Unsupervised Domain Adaptation. arXiv:2005.07839
Knowledge as Priors: Cross-Modal Knowledge Generalization for Datasets without Superior Knowledge. Zhao, Long et al. CVPR 2020
Large-Scale Domain Adaptation via Teacher-Student Learning. Li, Jinyu et al. arXiv:1708.05466
Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data. Fayek, Haytham M. & Kumar, Anurag. IJCAI 2020
Distilling Cross-Task Knowledge via Relationship Matching. Ye, Han-Jia. et al. CVPR 2020 [code]
Modality distillation with multiple stream networks for action recognition. Garcia, Nuno C. et al. ECCV 2018
Domain Adaptation through Task Distillation. Zhou, Brady et al. ECCV 2020 [code]
Dual Super-Resolution Learning for Semantic Segmentation. Wang, Li et al. CVPR 2020 [code]
Adaptively-Accumulated Knowledge Transfer for Partial Domain Adaptation. Jing, Taotao et al. ACM MM 2020
Domain2Vec: Domain Embedding for Unsupervised Domain Adaptation. Peng, Xingchao et al. ECCV 2020 [code]
Unsupervised Domain Adaptive Knowledge Distillation for Semantic Segmentation. Kothandaraman et al. arXiv:2011.08007
A Student‐Teacher Architecture for Dialog Domain Adaptation under the Meta‐Learning Setting. Qian, Kun et al. AAAI 2021
Multimodal Fusion via Teacher‐Student Network for Indoor Action Recognition. Bruce et al. AAAI 2021
Dual-Teacher++: Exploiting Intra-domain and Inter-domain Knowledge with Reliable Transfer for Cardiac Segmentation. Li, Kang et al. TMI 2021
Knowledge Distillation Methods for Efficient Unsupervised Adaptation Across Multiple Domains. Nguyen et al. IVC 2021
Feature-Supervised Action Modality Transfer. Thoker, Fida Mohammad and Snoek, Cees. ICPR 2020.
There is More than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking with Sound by Distilling Multimodal Knowledge. Francisco et al. CVPR 2021
Adaptive Consistency Regularization for Semi-Supervised Transfer Learning
Abulikemu. Abulikemu et al. CVPR 2021 [code]
Semantic-aware Knowledge Distillation for Few-Shot Class-Incremental Learning. Cheraghian et al. CVPR 2021
Distilling Causal Effect of Data in Class-Incremental Learning. Hu, Xinting et al. CVPR 2021 [code]
Semi-supervised Domain Adaptation based on Dual-level Domain Mixing for Semantic Segmentation. Chen, Shuaijun et al. CVPR 2021
PLOP: Learning without Forgetting for Continual Semantic Segmentation. Arthur et al. CVPR 2021
Continual Semantic Segmentation via Repulsion-Attraction of Sparse and Disentangled Latent Representations. Umberto & Pietro. CVPR 2021
Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution. Sun, Baoli et al. CVPR 2021 [code]
CReST: A Class-Rebalancing Self-Training Framework for Imbalanced Semi-Supervised Learning. Wei, Chen et al. CVPR 2021
Adaptive Boosting for Domain Adaptation: Towards Robust Predictions in Scene Segmentation. Zheng, Zhedong & Yang, Yi. CVPR 2021
Image Classification in the Dark Using Quanta Image Sensors. Gnanasambandam, Abhiram & Chan, Stanley H. ECCV 2020
Dynamic Low-Light Imaging with Quanta Image Sensors. Chi, Yiheng et al. ECCV 2020
Boosting Weakly Supervised Object Detection with Progressive Knowledge Transfer. Zhong, Yuanyi et al. ECCV 2020 [code]
Weight Decay Scheduling and Knowledge Distillation for Active Learning. ECCV 2020
Circumventing Outliers of AutoAugment with Knowledge Distillation. ECCV 2020
Improving Face Recognition from Hard Samples via Distribution Distillation Loss. ECCV 2020
Exclusivity-Consistency Regularized Knowledge Distillation for Face Recognition. ECCV 2020
Self-similarity Student for Partial Label Histopathology Image Segmentation. Cheng, Hsien-Tzu et al. ECCV 2020
Deep Semi-supervised Knowledge Distillation for Overlapping Cervical Cell Instance Segmentation. Zhou, Yanning et al. arXiv:2007.10787 [code]
Two-Level Residual Distillation based Triple Network for Incremental Object Detection. Yang, Dongbao et al. arXiv:2007.13428
Towards Unsupervised Crowd Counting via Regression-Detection Bi-knowledge Transfer. Liu, Yuting et al. ACM MM 2020
Teacher-Critical Training Strategies for Image Captioning. Huang, Yiqing & Chen, Jiansheng. arXiv:2009.14405
Object Relational Graph with Teacher-Recommended Learning for Video Captioning. Zhang, Ziqi et al. CVPR 2020
Multi-Frame to Single-Frame: Knowledge Distillation for 3D Object Detection. Wang Yue et al. ECCV 2020
Residual Feature Distillation Network for Lightweight Image Super-Resolution. Liu, Jie et al. ECCV 2020
Intra-Utterance Similarity Preserving Knowledge Distillation for Audio Tagging. Interspeech 2020
Federated Model Distillation with Noise-Free Differential Privacy. arXiv:2009.05537
Long-tailed Recognition by Routing Diverse Distribution-Aware Experts. Wang, Xudong et al. arXiv:2010.01809
Fast Video Salient Object Detection via Spatiotemporal Knowledge Distillation. Yi, Tang & Yuan, Li. arXiv:2010.10027
Multiresolution Knowledge Distillation for Anomaly Detection. Salehi et al. cvpr 2021
Channel-wise Distillation for Semantic Segmentation. Shu, Changyong et al. arXiv: 2011.13256
Teach me to segment with mixed supervision: Confident students become masters. Dolz, Jose et al. arXiv:2012.08051
Invariant Teacher and Equivariant Student for Unsupervised 3D Human Pose Estimation. Xu, Chenxin et al. AAAI 2021 [code]
Training data-efficient image transformers & distillation through attention. Touvron, Hugo et al. arXiv:2012.12877 [code]
SID: Incremental Learning for Anchor-Free Object Detection via Selective and Inter-Related Distillation. Peng, Can et al. arXiv:2012.15439
PSSM-Distil: Protein Secondary Structure Prediction (PSSP) on Low-Quality PSSM by Knowledge Distillation with Contrastive Learning. Wang, Qin et al. AAAI 2021
Diverse Knowledge Distillation for End-to‐End Person Search. Zhang, Xinyu et al. AAAI 2021
Enhanced Audio Tagging via Multi‐ to Single‐Modal Teacher‐Student Mutual Learning. Yin, Yifang et al. AAAI 2021
Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks. Li, Yige et al. ICLR 2021 [code]
Unbiased Teacher for Semi-Supervised Object Detection. Liu, Yen-Cheng et al. ICLR 2021 [code]
Localization Distillation for Object Detection. Zheng, Zhaohui et al. cvpr 2021 [code]
Distilling Knowledge via Intermediate Classifier Heads. Aryan & Amirali. arXiv:2103.00497
Distilling Object Detectors via Decoupled Features. (HUAWEI-Noah). CVPR 2021
General Instance Distillation for Object Detection. Dai, Xing et al. CVPR 2021
Multiresolution Knowledge Distillation for Anomaly Detection. Mohammadreza et al. CVPR 2021
Student-Teacher Feature Pyramid Matching for Unsupervised Anomaly Detection. Wang, Guodong et al. arXiv:2103.04257
Teacher-Explorer-Student Learning: A Novel Learning Method for Open Set Recognition. Jaeyeon Jang & Chang Ouk Kim. IEEE 2021
Dense Relation Distillation with Context-aware Aggregation for Few-Shot Object Detection. Hu, Hanzhe et al. CVPR 2021 [code]
Compressing Visual-linguistic Model via Knowledge Distillation. Fang, Zhiyuan et al. arXiv:2104.02096
Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification. Tian, Xudong et al. CVPR 2021
Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation. Wang, Liwei et al. CVPR 2021
Orderly Dual-Teacher Knowledge Distillation for Lightweight Human Pose Estimation. Zhao, Zhongqiu et al. arXiv:2104.10414
Boosting Light-Weight Depth Estimation Via Knowledge Distillation. Hu, Junjie et al. arXiv:2105.06143
Weakly Supervised Dense Video Captioning via Jointly Usage of Knowledge
Distillation and Cross-modal Matching. Wu, Bofeng et al. arViv:2105.08252
Revisiting Knowledge Distillation for Object Detection. Banitalebi-Dehkordi, Amin. arXiv: 2105.10633
Towards Compact Single Image Super-Resolution via Contrastive Self-distillation. Yanbo, Wang et al. IJCAI 2021
How many Observations are Enough? Knowledge Distillation for Trajectory Forecasting. Monti, Alessio et al. CVPR 2022
for NLP & Data-Mining
Patient Knowledge Distillation for BERT Model Compression. Sun, Siqi et al. arXiv:1908.09355
TinyBERT: Distilling BERT for Natural Language Understanding. Jiao, Xiaoqi et al. arXiv:1909.10351
Learning to Specialize with Knowledge Distillation for Visual Question Answering. NeurIPS 2018
Knowledge Distillation for Bilingual Dictionary Induction. EMNLP 2017
A Teacher-Student Framework for Maintainable Dialog Manager. EMNLP 2018
Understanding Knowledge Distillation in Non-Autoregressive Machine Translation. arxiv 2019
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. Sanh, Victor et al. arXiv:1910.01108
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models. Turc, Iulia et al. arXiv:1908.08962
On Knowledge distillation from complex networks for response prediction. Arora, Siddhartha et al. NAACL 2019
Distilling the Knowledge of BERT for Text Generation. arXiv:1911.03829v1
Understanding Knowledge Distillation in Non-autoregressive Machine Translation. arXiv:1911.02727
MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices. Sun, Zhiqing et al. ACL 2020
Acquiring Knowledge from Pre-trained Model to Neural Machine Translation. Weng, Rongxiang et al. AAAI 2020
TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval. Lu, Wenhao et al. KDD 2020
Improving BERT Fine-Tuning via Self-Ensemble and Self-Distillation. Xu, Yige et al. arXiv:2002.10345
FastBERT: a Self-distilling BERT with Adaptive Inference Time. Liu, Weijie et al. ACL 2020
LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression. Mao, Yihuan et al. arXiv:2004.04124
DynaBERT: Dynamic BERT with Adaptive Width and Depth. Hou, Lu et al. NeurIPS 2020
Structure-Level Knowledge Distillation For Multilingual Sequence Labeling. Wang, Xinyu et al. ACL 2020
Distilled embedding: non-linear embedding factorization using knowledge distillation. Lioutas, Vasileios et al. arXiv:1910.06720
Knowledge Distillation for Multilingual Unsupervised Neural Machine Translation. Sun, Haipeng et al. arXiv:2004.10171
Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation. Reimers, Nils & Gurevych, Iryna arXiv:2004.09813
Distilling Knowledge for Fast Retrieval-based Chat-bots. Tahami et al. arXiv:2004.11045
Single-/Multi-Source Cross-Lingual NER via Teacher-Student Learning on Unlabeled Data in Target Language. ACL 2020
Local Clustering with Mean Teacher for Semi-supervised Learning. arXiv:2004.09665
Time Series Data Augmentation for Neural Networks by Time Warping with a Discriminative Teacher. arXiv:2004.08780
Syntactic Structure Distillation Pretraining For Bidirectional Encoders. arXiv: 2005.13482
Distill, Adapt, Distill: Training Small, In-Domain Models for Neural Machine Translation. arXiv:2003.02877
Distilling Neural Networks for Faster and Greener Dependency Parsing. arXiv:2006.00844
Distilling Knowledge from Well-informed Soft Labels for Neural Relation Extraction. AAAI 2020 [code]
More Grounded Image Captioning by Distilling Image-Text Matching Model. Zhou, Yuanen et al. CVPR 2020
Multimodal Learning with Incomplete Modalities by Knowledge Distillation. Wang, Qi et al. KDD 2020
Distilling the Knowledge of BERT for Sequence-to-Sequence ASR. Futami, Hayato et al. arXiv:2008.03822
Contrastive Distillation on Intermediate Representations for Language Model Compression. Sun, Siqi et al. EMNLP 2020 [code]
Noisy Self-Knowledge Distillation for Text Summarization. arXiv:2009.07032
Simplified TinyBERT: Knowledge Distillation for Document Retrieval. arXiv:2009.07531
Autoregressive Knowledge Distillation through Imitation Learning. arXiv:2009.07253
BERT-EMD: Many-to-Many Layer Mapping for BERT Compression with Earth Mover’s Distance. EMNLP 2020 [code]
Interpretable Embedding Procedure Knowledge Transfer. Seunghyun Lee et al. AAAI 2021 [code]
LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding. Fu, Hao et al. AAAI 2021
Towards Zero-Shot Knowledge Distillation for Natural Language Processing. Ahmad et al. arXiv:2012.15495
Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across Domains. Pan, Haojie et al. AAAI 2021
Learning to Augment for Data-Scarce Domain BERT Knowledge Distillation. Feng, Lingyun et al. AAAI 2021
Label Confusion Learning to Enhance Text Classification Models. Guo, Biyang et al. AAAI 2021
NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application. Wu, Chuhan et al. kdd 2021
for RecSys
Developing Multi-Task Recommendations with Long-Term Rewards via Policy Distilled Reinforcement Learning. Liu, Xi et al. arXiv:2001.09595
A General Knowledge Distillation Framework for Counterfactual Recommendation via Uniform Data. Liu, Dugang et al. SIGIR 2020 [Sildes][code]
LightRec: a Memory and Search-Efficient Recommender System. Lian, Defu et al. WWW 2020
Privileged Features Distillation at Taobao Recommendations. Xu, Chen et al. KDD 2020
Next Point-of-Interest Recommendation on Resource-Constrained Mobile Devices. WWW 2020
Adversarial Distillation for Efficient Recommendation with External Knowledge. Chen, Xu et al. ACM Trans, 2018
Ranking Distillation: Learning Compact Ranking Models With High Performance for Recommender System. Tang, Jiaxi et al. SIGKDD 2018
A novel Enhanced Collaborative Autoencoder with knowledge distillation for top-N recommender systems. Pan, Yiteng et al. Neurocomputing 2019 [code]
ADER: Adaptively Distilled Exemplar Replay Towards Continual Learning for Session-based Recommendation. Mi, Fei et al. ACM RecSys 2020
Ensembled CTR Prediction via Knowledge Distillation. Zhu, Jieming et al.(Huawei) CIKM 2020
DE-RRD: A Knowledge Distillation Framework for Recommender System. Kang, Seongku et al. CIKM 2020 [code]
Neural Compatibility Modeling with Attentive Knowledge Distillation. Song, Xuemeng et al. SIGIR 2018
Binarized Collaborative Filtering with Distilling Graph Convolutional Networks. Wang, Haoyu et al. IJCAI 2019
Collaborative Distillation for Top-N Recommendation. Jae-woong Lee, et al. CIKM 2019
Distilling Structured Knowledge into Embeddings for Explainable and Accurate Recommendation. Zhang Yuan et al. WSDM 2020
UMEC:Unified Model and Embedding Compression for Efficient Recommendation Systems. ICLR 2021
Bidirectional Distillation for Top-K Recommender System. WWW 2021
Privileged Graph Distillation for Cold-start Recommendation. SIGIR 2021
Topology Distillation for Recommender System [KDD 2021]
Conditional Attention Networks for Distilling Knowledge Graphs in Recommendation [CIKM 2021]
Explore, Filter and Distill: Distilled Reinforcement Learning in Recommendation [CIKM 2021] [Video][Code]
Graph Structure Aware Contrastive Knowledge Distillation for Incremental Learning in Recommender Systems[CIKM 2021]
Conditional Graph Attention Networks for Distilling and Refining Knowledge Graphs in Recommendation[CIKM 2021]
Target Interest Distillation for Multi-Interest Recommendation [CIKM 2022] [Video][Code]
KDCRec: Knowledge Distillation for Counterfactual Recommendation Via Uniform Data [TKDE 2022] [Code]
Revisiting Graph based Social Recommendation: A Distillation Enhanced Social Graph Network[WWW 2022] [Code]
False Negative Distillation and Contrastive Learning for Personalized Outfit Recommendation [Arxiv 2110.06483]
Dual Correction Strategy for Ranking Distillation in Top-N Recommender System[ArXiv 2109.03459v1]
Scene-adaptive Knowledge Distillation for Sequential Recommendation via Differentiable Architecture Search. Chen, Lei et al.[ArXiv 2107.07173v1]
Interpolative Distillation for Unifying Biased and Debiased Recommendation [SIGIR 2022] [Video][Code]
FedSPLIT: One-Shot Federated Recommendation System Based on Non-negative Joint Matrix Factorization and Knowledge Distillation[Arxiv 2205.02359v1]
On-Device Next-Item Recommendation with Self-Supervised Knowledge Distillation[SIGIR 2022] [Code]
Cross-Task Knowledge Distillation in Multi-Task Recommendation[AAAI 2022]
Toward Understanding Privileged Features Distillation in Learning-to-Rank [NIPS 2022]
Debias the Black-box: A Fair Ranking Framework via Knowledge Distillation [WISE 2022]
Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense Embeddings[SIGIR 2022] [Code]
AutoFAS: Automatic Feature and Architecture Selection for Pre-Ranking System [KDD 2022]
An Incremental Learning framework for Large-scale CTR Prediction[RecSys 22]
Directed Acyclic Graph Factorization Machines for CTR Prediction via Knowledge Distillation [WSDM 2023] [Code]
Unbiased Knowledge Distillation for Recommendation [WSDM 2023] [Code]
DistilledCTR: Accurate and scalable CTR prediction model through model distillation [ESWA 2022]
Top-aware recommender distillation with deep reinforcement learning [Information Sciences 2021]
Model Pruning or Quantization
Accelerating Convolutional Neural Networks with Dominant Convolutional Kernel and Knowledge Pre-regression. ECCV 2016
N2N Learning: Network to Network Compression via Policy Gradient Reinforcement Learning. Ashok, Anubhav et al. ICLR 2018
Slimmable Neural Networks. Yu, Jiahui et al. ICLR 2018
Co-Evolutionary Compression for Unpaired Image Translation. Shu, Han et al. ICCV 2019
MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning. Liu, Zechun et al. ICCV 2019
LightPAFF: A Two-Stage Distillation Framework for Pre-training and Fine-tuning. ICLR 2020
Pruning with hints: an efficient framework for model acceleration. ICLR 2020
Training convolutional neural networks with cheap convolutions and online distillation. arXiv:1909.13063
Cooperative Pruning in Cross-Domain Deep Neural Network Compression. Chen, Shangyu et al. IJCAI 2019
QKD: Quantization-aware Knowledge Distillation. Kim, Jangho et al. arXiv:1911.12491v1
Neural Network Pruning with Residual-Connections and Limited-Data. Luo, Jian-Hao & Wu, Jianxin. CVPR 2020
Training Quantized Neural Networks with a Full-precision Auxiliary Module. Zhuang, Bohan et al. CVPR 2020
Towards Effective Low-bitwidth Convolutional Neural Networks. Zhuang, Bohan et al. CVPR 2018
Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations. Zhuang, Bohan et al. arXiv:1908.04680
Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation. Le et al. arXiv:2006.11487 [code]
Knowledge Distillation Beyond Model Compression. Choi, Arthur et al. arxiv:2007.01493
Distillation Guided Residual Learning for Binary Convolutional Neural Networks. Ye, Jianming et al. ECCV 2020
Cascaded channel pruning using hierarchical self-distillation. Miles & Mikolajczyk. BMVC 2020
TernaryBERT: Distillation-aware Ultra-low Bit BERT. Zhang, Wei et al. EMNLP 2020
Weight Distillation: Transferring the Knowledge in Neural Network Parameters. arXiv:2009.09152
Stochastic Precision Ensemble: Self-‐Knowledge Distillation for Quantized Deep Neural Networks. Boo, Yoonho et al. AAAI 2021
Binary Graph Neural Networks. Bahri, Mehdi et al. CVPR 2021
Self-Damaging Contrastive Learning. Jiang, Ziyu et al. ICML 2021
Information Theoretic Representation Distillation. Miles et al. BMVC 2022 [code]
Distillation Guided Residual Learning for Binary Convolutional Neural Networks. Ye, Jianming et al. ECCV 2020
Cascaded channel pruning using hierarchical self-distillation. Miles & Mikolajczyk. BMVC 2020
TernaryBERT: Distillation-aware Ultra-low Bit BERT. Zhang, Wei et al. EMNLP 2020
Weight Distillation: Transferring the Knowledge in Neural Network Parameters. arXiv:2009.09152
Stochastic Precision Ensemble: Self-‐Knowledge Distillation for Quantized Deep Neural Networks. Boo, Yoonho et al. AAAI 2021
Binary Graph Neural Networks. Bahri, Mehdi et al. CVPR 2021
Self-Damaging Contrastive Learning. Jiang, Ziyu et al. ICML 2021
Beyond
Do deep nets really need to be deep?. Ba,Jimmy, and Rich Caruana. NeurIPS 2014
When Does Label Smoothing Help? Müller, Rafael, Kornblith, and Hinton. NeurIPS 2019
Towards Understanding Knowledge Distillation. Phuong, Mary and Lampert, Christoph. ICML 2019
Harnessing deep neural networks with logical rules. ACL 2016
Adaptive Regularization of Labels. Ding, Qianggang et al. arXiv:1908.05474
Knowledge Isomorphism between Neural Networks. Liang, Ruofan et al. arXiv:1908.01581
(survey) Modeling Teacher-Student Techniques in Deep Neural Networks for Knowledge Distillation. arXiv:1912.13179
Understanding and Improving Knowledge Distillation. Tang, Jiaxi et al. arXiv:2002.03532
The State of Knowledge Distillation for Classification. Ruffy, Fabian and Chahal, Karanbir. arXiv:1912.10850 [code]
Explaining Knowledge Distillation by Quantifying the Knowledge. Zhang, Quanshi et al. CVPR 2020
DeepVID: deep visual interpretation and diagnosis for image classifiers via knowledge distillation. IEEE Trans, 2019.
On the Unreasonable Effectiveness of Knowledge Distillation: Analysis in the Kernel Regime. Rahbar, Arman et al. arXiv:2003.13438
(survey) Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks. Wang, Lin & Yoon, Kuk-Jin. arXiv:2004.05937
Why distillation helps: a statistical perspective. arXiv:2005.10419
Transferring Inductive Biases through Knowledge Distillation. Abnar, Samira et al. arXiv:2006.00555
Does label smoothing mitigate label noise? Lukasik, Michal et al. ICML 2020
An Empirical Analysis of the Impact of Data Augmentation on Knowledge Distillation. Das, Deepan et al. arXiv:2006.03810
(survey) Knowledge Distillation: A Survey. Gou, Jianping et al. IJCV 2021
Does Adversarial Transferability Indicate Knowledge Transferability? Liang, Kaizhao et al. arXiv:2006.14512
On the Demystification of Knowledge Distillation: A Residual Network Perspective. Jha et al. arXiv:2006.16589
Enhancing Simple Models by Exploiting What They Already Know. Dhurandhar et al. ICML 2020
Feature-Extracting Functions for Neural Logic Rule Learning. Gupta & Robles-Kelly.arXiv:2008.06326
On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective. SeongUk et al. arXiv:2009.04120
Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect Teacher. Ji, Guangda & Zhu, Zhanxing. NeurIPS 2020
In Defense of Feature Mimicking for Knowledge Distillation. Wang, Guo-Hua et al. arXiv:2011.0142
Solvable Model for Inheriting the Regularization through Knowledge Distillation. Luca Saglietti & Lenka Zdeborova. arXiv:2012.00194
Undistillable: Making A Nasty Teacher That CANNOT Teach Students. ICLR 2021
Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning. Allen-Zhu, Zeyuan & Li, Yuanzhi.(Microsoft) arXiv:2012.09816
Student-Teacher Learning from Clean Inputs to Noisy Inputs. Hong, Guanzhe et al. CVPR 2021
Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study. ICLR 2021 [project]
Model Distillation for Revenue Optimization: Interpretable Personalized Pricing. Biggs, Max et al. ICML 2021
A statistical perspective on distillation. Aditya et al(Google). ICML 2021
(survey) Data-Free Knowledge Transfer: A Survey. Liu, Yuang et al. arXiv:2112.15278
Knowledge Distillation Beyond Model Compression. Choi, Sarfraz et. al. arxiv:2007.01493