MELT-FSIL: Adaptive Hybrid-Loss Few-Shot Learning for Inclusive Sign Language Control in Autonomous Vehicles
IEEE Access, cilt.14, ss.98046-98060, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14
- Basım Tarihi: 2026
- Doi Numarası: 10.1109/access.2026.3702867
- Dergi Adı: IEEE Access
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
- Sayfa Sayıları: ss.98046-98060
- Anahtar Kelimeler: American sign language (ASL), assistive technology, Autonomous vehicles, catastrophic forgetting, hybrid loss function, incremental learning, multimodal interaction
- Van Yüzüncü Yıl Üniversitesi Adresli: Evet
Özet
Despite significant advances in autonomous vehicle (AV) technologies, current interaction modalities remain insufficiently accessible to people with physical and communication impairments. Conventional AV interfaces, particularly voice-activated and touchscreen-based systems, may create usability barriers for users who are deaf, non-verbal, or have limited motor control, thereby limiting effective human–vehicle interaction. To address this gap, this study proposes a vision-based ASL recognition framework as an initial step toward accessible AV interaction. The proposed framework uses Few-Shot Incremental Learning (FSIL) to recognize American Sign Language (ASL) hand gestures from static RGB hand-sign images, enabling the model to learn new gesture classes from limited examples while maintaining performance on previously learned classes.The core contribution of this work is a hybrid loss function, termed Metric–Entropy Learning Tradeoff (MELT), which combines prototypical metric learning with cross-entropy-based discriminative learning within an FSIL training regime. By balancing prototype-based class separation with entropy-guided classification, MELT improves the model’s ability to adapt to newly introduced ASL gestures while reducing catastrophic forgetting. In addition, prototype normalization and episodic memory replay are incorporated to stabilize incremental learning performance. Empirical results demonstrate that the proposed FSIL framework with MELT loss improves gesture recognition accuracy from 96.0% using the baseline CNN to 97.5% in the FSL setting and 97.9% in the FSIL setting after the integration of novel gesture classes. These findings demonstrate the potential of hybrid loss-based FSIL for ASL gesture recognition in assistive AV interaction scenarios. While the broader long-term goal is to develop a fully multimodal AV interface, the present study is limited to static RGB image-based ASL recognition; future work will extend the system by incorporating additional modalities such as depth, audio, wearable signals, facial expressions, and vehicle-state information. Thus, this work represents an early step toward more inclusive and accessible AV interaction systems.