Journals & Magazines >IEEE Transactions on Cognitiv... >Volume: 17 Issue: 1

SpikingViT: A Multiscale Spiking Vision Transformer Model for Event-Based Object Detection

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

Event cameras have unique advantages in object detection, capturing asynchronous events without continuous frames. They excel in dynamic range, low latency, and high-spee...Show More

Metadata

Abstract:

Event cameras have unique advantages in object detection, capturing asynchronous events without continuous frames. They excel in dynamic range, low latency, and high-speed motion scenarios, with lower power consumption. However, aggregating event data into image frames leads to information loss and reduced detection performance. Applying traditional neural networks to event camera outputs is challenging due to event data's distinct characteristics. In this study, we present a novel spiking neural networks (SNNs)-based object detection model, the spiking vision transformer (SpikingViT) to address these issues. First, we design a dedicated event data converting module that effectively captures the unique characteristics of event data, mitigating the risk of information loss while preserving its spatiotemporal features. Second, we introduce SpikingViT, a novel object detection model that leverages SNNs capable of extracting spatiotemporal information among events data. SpikingViT combines the advantages of SNNs and transformer models, incorporating mechanisms such as attention and residual voltage memory to further enhance detection performance. Extensive experiments have substantiated the remarkable proficiency of SpikingViT in event-based object detection, positioning it as a formidable contender. Our proposed approach adeptly retains spatiotemporal information inherent in event data, leading to a substantial enhancement in detection performance.

Published in: IEEE Transactions on Cognitive and Developmental Systems ( Volume: 17, Issue: 1, February 2025)

Page(s): 130 - 146

Date of Publication: 04 July 2024

ISSN Information:

DOI: 10.1109/TCDS.2024.3422873

Funding Agency:

Contents

References is not available for this document.

SpikingViT: A Multiscale Spiking Vision Transformer Model for Event-Based Object Detection

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

References

IEEE Account

Purchase Details

Profile Information

Need Help?

SpikingViT: A Multiscale Spiking Vision Transformer Model for Event-Based Object Detection

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

References

IEEE Account

Purchase Details

Profile Information

Need Help?