Paper Title: Intelligent visual perception beyond classification: a survey of deep learning and foundation models
Authors: Sunita Dixit
Corresponding Author: Sunita Dixit (bhardwajsunita23@gmail.com)/India
Abstract
Intelligent visual perception has advanced from traditional image classification to thorough comprehension of objects, scenes, relations, spatial context, and semantic information. This paper surveys the development of intelligent visual perception, particularly focusing on DL and foundation models. This section provides a comprehensive overview of visual recognition, detection, segmentation, and scene understanding, including key methods such as convolutional neural networks, recurrent neural networks, attention mechanisms, encoder-decoder architectures, generative models, and Vision Transformers. The study also explores feature attribution, visual attention, visual reasoning, and multimodal vision-language models, which combine visual and text information for tasks such as visual question answering, captioning, and visual generation. This study examines recent foundation models incorporating vision-language architectures for their transferability, contextual understanding, and generalized perception capabilities. The study also discusses challenges related to computational efficiency, environmental variability, data requirements, robustness, explainability, and real-time deployment. Finally, it contrasts current developments to identify new approaches to scalable, interpretable, multimodal, and human-centric visual intelligence systems.