ENGLISH
La vitrine de diffusion des publications et contributions des chercheurs(-euses) de l'ÉTS
RECHERCHER

Foundation models for autonomous driving: A comprehensive survey

Fourati, Sonda, Jaafar, Wael, Baccar, Noura, Alfattani, Safwan et Langar, Rami. 2026. « Foundation models for autonomous driving: A comprehensive survey ». Engineering Applications of Artificial Intelligence, vol. 176.

[thumbnail of Jaafar-W-2026-33728.pdf]
Prévisualisation
PDF
Jaafar-W-2026-33728.pdf - Version publiée
Licence d'utilisation : Creative Commons CC BY.

Télécharger (5MB) | Prévisualisation

Résumé

Large Language Models (LLMs) have showcased remarkable proficiency in various information-processing tasks. They excel at data extraction, literature summarization, content generation, predictive modeling, decision-making, and system control. Moreover, Vision-Language Models (VLMs) and Multimodal LLMs (MLLMs), collectively referred to in this work as Cross-modal Language Models (XLMs), integrate multiple data modalities with language understanding, thereby advancing Autonomous Driving Systems (ADS). On the implemented Artificial Intelligence (AI) side, we analyze core techniques such as prompt engineering, supervised fine-tuning, reinforcement learning from human feedback, knowledge distillation, quantization and pruning, and safety alignment/verification, together with edge-aware deployment strategies. On the application of AI side, we map XLMs capabilities to the driving stack, including perception, prediction, planning, control, and human–machine interaction/vehicle-to-everything, and summarize how XLMs improve scene understanding, intent forecasting, decision-making, and closed-loop control by coupling natural-language reasoning with multimodal sensory inputs, such as panoramic images, Light Detection and Ranging (LiDAR), and radar. In this survey, we synthesize the state of XLMs for ADS: we review the relevant literature on ADS and XLMs, including their architectures, tools, and frameworks. We then compare deployment approaches across the driving stack and summarize datasets, simulators, and benchmarks for both open- and closed-loop evaluation. Finally, we analyze key challenges, such as grounding and hallucination, long-tail robustness, real-time and resource constraints, safety alignment and verification, and data governance and privacy, and outline research directions toward safe, efficient, and trustworthy XLM-enabled ADS.

Type de document: Article publié dans une revue, révisé par les pairs
Chercheur(-euse):
Chercheur(-euse)
Jaafar, Waël
Langar, Rami
Affiliation: Génie logiciel et des technologies de l'information, Génie logiciel et des technologies de l'information
Date de dépôt: 12 mai 2026 14:40
Dernière modification: 22 mai 2026 22:16
URI: https://espace2.etsmtl.ca/id/eprint/33728

Actions (Authentification requise)

Dernière vérification avant le dépôt Dernière vérification avant le dépôt