A Hybrid Deep Learning and Computer Vision Approach for Automated Sea Debris Detection

Authors

  • Naroob Nadeem University of Agriculture, fsd. Author
  • Milhan Afzal Khan University of Agriculture, fsd. Author

DOI:

https://doi.org/10.63163/jpehss.v4i1.1641

Keywords:

Marine Debris, Deep Learning, Computer Vision, CNN, Transformer, YOLO, Underwater Detection

Abstract

Identification of marine pollutants at an early stage is important to effective waste management and environmental protection as marine debris is now a serious threat to marine ecosystems, biodiversity, fisheries, tourism, coastal economies and human health, particularly plastic waste. The traditional techniques of monitoring, such as man-to-man shore inspection, diver surveys, sonar and boats, tend to be expensive to operate, slow to respond, limited in scale and reliant on man's capability, in particular in turbid water, where the concentration of suspended solids changes, in depth and where light penetration is uneven. To address these restrictions, this study proposes an automated sea debris detecting framework based on computer vision and deep learning, which can detect the sea debris, locate and classify the sea debris with different types in underwater and floating water environments. To include environmental diversity and generalization of the model, the public benchmark datasets, including TrashCan, SeaClear, and other open-source underwater data repositories were added. The preprocessing pipeline includes the following steps: Verification of annotations, resizing, normalization, removal of corrupted images, underwater processing, such as histogram equalization, CLAHE, white balance correction, dehazing and contrast stretching for improving visibility of debris. Albumentations were used to strengthen the robustness and minimize overfitting in data augmentation, which involved using rotation, flipping, scaling, and blurring with Gaussian noise. The main contribution of the work is the new Hybrid CNN-Transformer Architecture (HCTA) that combines the ResNet50 network to extract local texture features with the Transformer Encoder network to learn global contextual features to identify debris objects when they are blurred, occluded, and/or low in contrast. The model was coded in Python and trained on Google Colab GPU machines, and its results were compared to the state-of-the-art models, such as Vision Transformer, DETR, Swin Transformer, and EfficientFormer. Precision, recall, F1-score, IoU, ROC and mAP were used to assess performance. The results of these experiments reveal that the Autonomous Underwater Vehicles, marine drones, and the robotic ocean clean-up systems are more powerful and have great potential for use in real-time.

Downloads

Published

2026-03-26

Issue

Section

Computer Science and Information Technology