Abstract:
To address the limitations of existing robotic grasping slip detection methods in capturing slip features and suppressing redundant noise, which lead to low detection accuracy and limited generalization ability, this paper proposes a vision-tactile slip detection method based on bi-temporal dynamic capture and grouped coordinate attention. A parallel feature extraction pipeline for vision and tactile modalities is constructed, along with an adaptive fusion architecture for vision-tactile and cross-modal features. According to the temporal characteristics of slip events, a Bi-temporal slip perception fusion module is designed to enhance the network's ability to extract features precisely at the moment of slip onset. To improve the network's capability in suppressing interference from redundant information, a Global grouped coordinate attention module is developed. Adaptive dynamic weight assignment is applied to vision-tactile and cross-modal features to fully exploit the complementary nature of different modalities. Experimental results demonstrate that the proposed method outperforms existing methods by average margins of 10.04 and 5.68 percentage points, respectively. In physical robot experiments, the proposed method achieves a slip detection accuracy of 98.43% for 16 unseen objects, validating its effectiveness and application potential in practical robotic grasping tasks.