Fourati, Sonda, Jaafar, Wael, Baccar, Noura, Alfattani, Safwan et Langar, Rami.
2026.
« Foundation models for autonomous driving: A comprehensive survey ».
Engineering Applications of Artificial Intelligence, vol. 176.
Prévisualisation |
PDF
Jaafar-W-2026-33728.pdf - Version publiée Licence d'utilisation : Creative Commons CC BY. Télécharger (5MB) | Prévisualisation |
Résumé
Large Language Models (LLMs) have showcased remarkable proficiency in various information-processing tasks. They excel at data extraction, literature summarization, content generation, predictive modeling, decision-making, and system control. Moreover, Vision-Language Models (VLMs) and Multimodal LLMs (MLLMs), collectively referred to in this work as Cross-modal Language Models (XLMs), integrate multiple data modalities with language understanding, thereby advancing Autonomous Driving Systems (ADS). On the implemented Artificial Intelligence (AI) side, we analyze core techniques such as prompt engineering, supervised fine-tuning, reinforcement learning from human feedback, knowledge distillation, quantization and pruning, and safety alignment/verification, together with edge-aware deployment strategies. On the application of AI side, we map XLMs capabilities to the driving stack, including perception, prediction, planning, control, and human–machine interaction/vehicle-to-everything, and summarize how XLMs improve scene understanding, intent forecasting, decision-making, and closed-loop control by coupling natural-language reasoning with multimodal sensory inputs, such as panoramic images, Light Detection and Ranging (LiDAR), and radar. In this survey, we synthesize the state of XLMs for ADS: we review the relevant literature on ADS and XLMs, including their architectures, tools, and frameworks. We then compare deployment approaches across the driving stack and summarize datasets, simulators, and benchmarks for both open- and closed-loop evaluation. Finally, we analyze key challenges, such as grounding and hallucination, long-tail robustness, real-time and resource constraints, safety alignment and verification, and data governance and privacy, and outline research directions toward safe, efficient, and trustworthy XLM-enabled ADS.
| Type de document: | Article publié dans une revue, révisé par les pairs |
|---|---|
| Chercheur(-euse): | Chercheur(-euse) Jaafar, Waël Langar, Rami |
| Affiliation: | Génie logiciel et des technologies de l'information, Génie logiciel et des technologies de l'information |
| Date de dépôt: | 12 mai 2026 14:40 |
| Dernière modification: | 22 mai 2026 22:16 |
| URI: | https://espace2.etsmtl.ca/id/eprint/33728 |
Actions (Authentification requise)
![]() |
Dernière vérification avant le dépôt |

