KHAN, Afrasiab, NAWAZ, Tahir, ASADUZZAMAN, Md, HASAN, Mohammad, ZEB, Ayesha, MALIK, Anjum, AQEEL, Anas and SHAFAIT, Faisal (2026) Self-Supervised Monocular Depth Estimation With Hybrid CNN-Transformer Architecture Using Skip Attention Mechanism. Signal, Image and Video Processing. ISSN 1863-1703
Depth_Estimation_RP_Clean_Copy.pdf - AUTHOR'S ACCEPTED Version (default)
Restricted to Repository staff only until 12 September 2027.
Available under License Type All Rights Reserved.
Download (5MB) | Request a copy
Abstract or description
Monocular depth estimation is vital for robotics, self-driving vehicles, and virtual reality, where accurate spatial understanding and real-time performance are essential. Existing methods, however, face challenges such as reliance on large annotated datasets, multi-camera views, complex architectures unsuitable for real-time use, limited generalization, semantic inconsistencies, and over-smoothing. We address these challenges with a hybrid self-supervised model incorporating a Skip Attention Mechanism (SAM) to minimize semantic inconsistencies between encoder and decoder, and a mixed pooling method combining max and average pooling to address over-smoothing and contextual information loss. To demonstrate the effectiveness of the proposed method, we performed extensive evaluation and comparisons on publicly available datasets, including KITTI, Make3D, and NYU-Depth V2. On the KITTI dataset, the proposed model achieves an absolute relative error of 0.108 and an RMSE of 4.600, delivering favorable accuracy compared to established lightweight methods such as Monodepth2, R-MSFM3, R-MSFM6, RTIA-Mono, and Lite-Mono-Small, while using approximately one-fifth of the parameters of Monodepth2 and achieving 14.62 FPS on the Jetson Nano with only 2.478M parameters. Our proposed model also achieves the most encouraging accuracy and depth error metrics among all baseline methods on both Make3D (Abs-Rel of 0.301, RMSE of 6.849) and NYU-Depth V2 (Abs-Rel of 0.317, RMSE of 0.989) datasets, while maintaining competitive real-time inference speed at 14.6 FPS and 15 FPS on the Jetson Nano board, respectively, showing a favorable accuracy and efficiency trade-off suitable for deployment on resource-constrained devices.
| Item Type: | Article |
|---|---|
| Faculty: | School of Digital, Technologies and Arts > Engineering |
| Depositing User: | Md ASADUZZAMAN |
| Date Deposited: | 28 Sep 2026 14:04 |
| Last Modified: | 28 Sep 2026 14:04 |
| URI: | https://eprints.staffs.ac.uk/id/eprint/9782 |
Lists
Lists