Explore open access research and scholarly works from STORE - University of Staffordshire Online Repository

Advanced Search

Self-Supervised Monocular Depth Estimation With Hybrid CNN-Transformer Architecture Using Skip Attention Mechanism

KHAN, Afrasiab, NAWAZ, Tahir, ASADUZZAMAN, Md, HASAN, Mohammad, ZEB, Ayesha, MALIK, Anjum, AQEEL, Anas and SHAFAIT, Faisal (2026) Self-Supervised Monocular Depth Estimation With Hybrid CNN-Transformer Architecture Using Skip Attention Mechanism. Signal, Image and Video Processing. ISSN 1863-1703

[thumbnail of Depth_Estimation_RP_Clean_Copy.pdf] Text
Depth_Estimation_RP_Clean_Copy.pdf - AUTHOR'S ACCEPTED Version (default)
Restricted to Repository staff only until 12 September 2027.
Available under License Type All Rights Reserved.

Download (5MB) | Request a copy
Official URL: https://link.springer.com/article/10.1007/s11760-0...

Abstract or description

Monocular depth estimation is vital for robotics, self-driving vehicles, and virtual reality, where accurate spatial understanding and real-time performance are essential. Existing methods, however, face challenges such as reliance on large annotated datasets, multi-camera views, complex architectures unsuitable for real-time use, limited generalization, semantic inconsistencies, and over-smoothing. We address these challenges with a hybrid self-supervised model incorporating a Skip Attention Mechanism (SAM) to minimize semantic inconsistencies between encoder and decoder, and a mixed pooling method combining max and average pooling to address over-smoothing and contextual information loss. To demonstrate the effectiveness of the proposed method, we performed extensive evaluation and comparisons on publicly available datasets, including KITTI, Make3D, and NYU-Depth V2. On the KITTI dataset, the proposed model achieves an absolute relative error of 0.108 and an RMSE of 4.600, delivering favorable accuracy compared to established lightweight methods such as Monodepth2, R-MSFM3, R-MSFM6, RTIA-Mono, and Lite-Mono-Small, while using approximately one-fifth of the parameters of Monodepth2 and achieving 14.62 FPS on the Jetson Nano with only 2.478M parameters. Our proposed model also achieves the most encouraging accuracy and depth error metrics among all baseline methods on both Make3D (Abs-Rel of 0.301, RMSE of 6.849) and NYU-Depth V2 (Abs-Rel of 0.317, RMSE of 0.989) datasets, while maintaining competitive real-time inference speed at 14.6 FPS and 15 FPS on the Jetson Nano board, respectively, showing a favorable accuracy and efficiency trade-off suitable for deployment on resource-constrained devices.

Item Type: Article
Faculty: School of Digital, Technologies and Arts > Engineering
Depositing User: Md ASADUZZAMAN
Date Deposited: 28 Sep 2026 14:04
Last Modified: 28 Sep 2026 14:04
URI: https://eprints.staffs.ac.uk/id/eprint/9782

Actions (login required)

View Item
View Item