基于双时序动态捕捉与分组坐标注意力的视触觉滑动检测

    Vision-tactile Slip Detection Based on Bi-temporal Dynamic Capture and Grouped Coordinate Attention

    • 摘要: 针对当前机器人抓取滑动检测方法在滑动特征捕捉、冗余噪声抑制方面存在不足,导致检测精度不高、模型泛化能力受限的问题。本文提出一种基于双时序动态捕捉与分组坐标注意力的视触觉滑动检测方法,并构建了视觉与触觉并行的特征提取链路和视触觉、跨模态特征的自适应融合架构。根据滑动事件的时序特征,本文设计了双时滑动感知融合模块,以增强网络对滑动发生瞬间的提取能力。为提升网络抑制冗余信息干扰的能力,本文设计了全局分组坐标注意力模块,对视触觉、跨模态特征进行自适应动态权重分配,以挖掘跨模态间的互补特性。实验结果表明,所提方法相比已有方法平均提升了10.04个百分点和5.68个百分点。在实体机器人实验中,对于16种未见物体,滑动检测的准确率为98.43%,由此验证了该方法在实际抓取过程中的有效性和应用潜力。

       

      Abstract: To address the limitations of existing robotic grasping slip detection methods in capturing slip features and suppressing redundant noise, which lead to low detection accuracy and limited generalization ability, this paper proposes a vision-tactile slip detection method based on bi-temporal dynamic capture and grouped coordinate attention. A parallel feature extraction pipeline for vision and tactile modalities is constructed, along with an adaptive fusion architecture for vision-tactile and cross-modal features. According to the temporal characteristics of slip events, a Bi-temporal slip perception fusion module is designed to enhance the network's ability to extract features precisely at the moment of slip onset. To improve the network's capability in suppressing interference from redundant information, a Global grouped coordinate attention module is developed. Adaptive dynamic weight assignment is applied to vision-tactile and cross-modal features to fully exploit the complementary nature of different modalities. Experimental results demonstrate that the proposed method outperforms existing methods by average margins of 10.04 and 5.68 percentage points, respectively. In physical robot experiments, the proposed method achieves a slip detection accuracy of 98.43% for 16 unseen objects, validating its effectiveness and application potential in practical robotic grasping tasks.

       

    /

    返回文章
    返回