Zhejiang University of Science and Technology, China
A – Conceptualization; B – Methodology; C – Software; D – Validation; E – Formal analysis; F – Investigation; G – Resources; H – Data curation; I – Writing – original draft; J – Writing – review & editing; K – Visualization; L – Supervision; M – Project administration; N – Funding acquisition
In dynamic, occlusion-prone environments like football stadiums, reliable human pose estimation is essential for mobile robots, where conventional systems often fail due to partial visibility and rapid motion. This paper presents a lightweight pose estimation framework for real-time operation on resource-constrained edge platforms and the X5 robot. It employs selective knowledge distillation to transfer occlusion-robust and motion-aware features from a pre-trained teacher model to a compact student model, preserving efficiency while enhancing reliability. Three components are integrated by synthetic occlusion pattern embeddings for visibility, temporal motion cue extraction for movements and cross-modal attention alignment to focus on visible body regions. Validated on the D-Robotics mono2d_body_detection benchmark, the system achieves accuracy improvements under heavy occlusion while maintaining ≥30 FPS on the target platform. Experimental results confirm the framework balances high accuracy with low complexity, which makes it suitable for reliable deployment in football stadiums.
REFERENCES(36)
1.
Cao Z, Hidalgo G, Simon T, Wei SE, Sheikh Y. OpenPose: realtime multi-person 2D pose estimation using part affinity fields. IEEE Transactions on Pattern Analysis and Machine Intelligence 2021; 43(1): 172-186. https://doi:10.1109/TPAMI.2019....
Cheng Y, Yang B, Wang B, Yan W, Tan RT. Occlusion-aware networks for 3D human pose estimation in video. In: Proceedings of the IEEE/CVF International Conference on Computer Vision 2019; 723-732.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I. Attention is all you need. In: Advances in Neural Information Processing Systems 2017; 30: 5998-6008.
Hamdan S, Ayyash M, Almajali S. Edge-computing architectures for Internet of Things applications: a survey. Sensors 2020; 20(22): 6441. https:// doi:10.3390/s20226441.
He K, Gkioxari G, Dollár P, Girshick R. Mask R-CNN. In: Proceedings of the IEEE International Conference on Computer Vision 2017; 2961-2969. https:// doi:10.1109/ICCV.2017.322.
Cho NG, Yuille AL, Lee SW. Adaptive occlusion state estimation for human pose tracking under self-occlusions. Pattern Recognition 2013; 46(3): 649-661. https://doi:10.1016/j.patcog.2....
Li Z, Ye J, Song M, Huang Y, Pan Z. Online knowledge distillation for efficient pose estimation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision 2021; 11740-11750. https:// doi:10.1109/ICCV48922.2021.01153.
Liu K, Yang Z, Zhang J, Wang J, Wang S, Yuan C, Guo R. Boosting pose estimators via cross-representation distillation. In: Computer Vision–ECCV 2024 Workshops, 2024.
Choi K, Kersner M, Morton J, Chang B. Temporal knowledge distillation for on-device audio classification. In: IEEE International Conference on Acoustics, Speech and Signal Processing 2022. https:// doi:10.1109/ICASSP43922.2022.9747633.
Edriss S, Romagnoli C, Caprioli L, Zanela A, Panichi E, Campoli F, Padua E. The role of emergent technologies in the dynamic and kinematic assessment of human movement in sport and clinical applications. Applied Sciences 2024; 14(3): 1012. https://doi:10.3390/app1403101....
Hafner FM, Bhuyian A, Kooij JFP, Granger E. Cross-modal distillation for RGB-depth person re-identification. Computer Vision and Image Understanding 2022; 216: 103352. https://doi:10.1016/j.cviu.202....
Wang T, Hu G, Fu B, Wang H. Uncertainty-guided cross-modal distillation for category-level object pose estimation. IEEE Signal Processing Letters 2026; 33: 26-30. https://doi:10.1109/LSP.2025.3....
Sun K, Xiao B, Liu D, Wang J. Deep high-resolution representation learning for human pose estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2019; 5693-5703.
Pepik B, Stark M, Gehler P, Schiele B. Occlusion patterns for object class detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2013. https:// doi:10.1109/CVPR.2013.422.
Khirodkar R, Tripathi S, Kitani K. Occluded human mesh recovery. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2022.
Park W, Kim D, Lu Y, Cho M. Relational knowledge distillation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2019; 3967-3976.
Teed Z, Deng J. RAFT: recurrent all-pairs field transforms for optical flow. In: Computer Vision–ECCV 2020. Lecture Notes in Computer Science 2020; 12347: 402-419. https:// doi:10.1007/978-3-030-58536-5_24.
Xiu Y, Li J, Wang H, Fang Y, Lu C. Pose flow: efficient online pose tracking. arXiv preprint arXiv:1802.00977. 2018. https:// doi:10.48550/arXiv.1802.00977.
Howard A, Sandler M, Chen B, Wang W, Chen LC, Tan M, Chu G, Vasudevan V, Zhu Y, Pang R, Le QV, Adam H. Searching for MobileNetV3. In: Proceedings of the IEEE/CVF International Conference on Computer Vision 2019; 1314-1324.
Shi X, Chen Z, Wang H, Yeung DY, Wong WK, Woo WC. Convolutional LSTM network: a machine learning approach for precipitation nowcasting. In: Advances in Neural Information Processing Systems 2015; 28.
Andrews P, Borch N, Fjeld M. FootyVision: multi-object tracking, localisation, and augmentation of players and ball in football video. In: Proceedings of the 32nd ACM International Conference on Multimedia 2024. https://doi:10.1145/3664647.36....
Wen B, Mitash C, Soorian S, Kimmel A, Sintov A, Bekris KE. Robust, occlusion-aware pose estimation for objects grasped by adaptive hands. In: IEEE International Conference on Robotics and Automation 2020; 6210-6217.
Balasubramanian S, Melendez-Calderon A, Roby-Brami A, Burdet E. On the analysis of movement smoothness. Journal of NeuroEngineering and Rehabilitation 2015; 12:112. https:// doi:10.1186/s12984-015-0090-9.
Sárándi I, Linder T, Arras KO, Leibe B. How robust is 3D human pose estimation to occlusion? arXiv preprint arXiv:1808.09316. 2018. https:// doi:10.48550/arXiv.1808.09316.
Bertasius G, Feichtenhofer C, Tran D, Shi J, Torresani L. Learning temporal pose estimation from sparsely-labeled videos. In: Advances in Neural Information Processing Systems 2019; 32.
Xiao B, Wu H, Wei Y. Simple baselines for human pose estimation and tracking. In: Computer Vision–ECCV 2018. Lecture Notes in Computer Science 2018; 466-481.
We process personal data collected when visiting the website. The function of obtaining information about users and their behavior is carried out by voluntarily entered information in forms and saving cookies in end devices. Data, including cookies, are used to provide services, improve the user experience and to analyze the traffic in accordance with the Privacy policy. Data are also collected and processed by Google Analytics tool (more).
You can change cookies settings in your browser. Restricted use of cookies in the browser configuration may affect some functionalities of the website.