![]() Flowchart of the proposed SORA framework.
|
[SORA:結合骨架引導視覺語言模型之嚴重度感知序位風險評估方法於嬰幼兒照護之研究] Infant-care monitoring requires graded risk decisions rather than binary anomaly labels because unsafe situations differ in severity and response urgency. This thesis proposes SORA, a severity-aware ordinal risk assessment framework that integrates Contrastive Language-Image Pre-training (CLIP)-based visual encoding, skeleton-guided caregiver-infant interaction cues, Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (BLIP-2) captions, and keyword-enhanced text encoding. Instead of treating the three labels as independent nominal classes, SORA predicts a scalar risk score so that low-, medium-, and high-risk states remain ordered along one response axis and are decoded by validation-selected thresholds. It also stores captions, interaction cues, timestamps, and risk labels as structured event records for later review. On a three-level proxy benchmark, SORA achieves 89.22% accuracy, 88.70% Macro-F1, and 99.38% high-risk recall, improving over AnomalyCLIP by 4.71, 4.88, and 11.11 percentage points, respectively. The results show that ordered risk modeling and interaction-aware multimodal evidence reduce safety-critical underestimation while preserving interpretable review records.
SUMMARY (中文總結): RESULTS:
The full ordinal SORA model achieves the best CES, high-risk recall, Macro-F1, accuracy, and QWK among the baseline risk assessment systems compared above. Among 162 high-risk test samples, SORA correctly identifies 161, leaving only one high-risk underestimation.
REFERENCES: [1] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, "Learning transferable visual models from natural language supervision," in Proceedings of the 38th International Conference on Machine Learning, vol. 139, pp. 8748-8763, 2021. [2] Q. Zhou, G. Pang, Y. Tian, S. He, and J. Chen, "AnomalyCLIP: Object-agnostic prompt learning for zero-shot anomaly detection," in The Twelfth International Conference on Learning Representations, 2024. [3] W. Cao, V. Mirjalili, and S. Raschka, "Rank consistent ordinal regression for neural networks with application to age estimation," Pattern Recognition Letters, vol. 140, pp. 325-331, 2020. [4] J. Li, D. Li, S. Savarese, and S. Hoi, "BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models," in Proceedings of the 40th International Conference on Machine Learning, vol. 202, pp. 19730-19742, 2023. [5] T. Jiang, X. Xie, and Y. Li, "RTMW: Real-time multi-person 2D and 3D whole-body pose estimation," arXiv preprint arXiv:2407.08634, 2024. [6] C. Lyu, W. Zhang, H. Huang, Y. Zhou, Y. Wang, Y. Liu, S. Zhang, and K. Chen, "RTMDet: An empirical study of designing real-time object detectors," arXiv preprint arXiv:2212.07784, 2022. [7] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems, vol. 30, pp. 6000-6010, 2017. |