This position is within one of TRATON’s companies.

Master thesis - Vision Foundation Models enhanced Vision-Language Model 3D Spatiotemporal Reasoning

30 credits - Using Vision Foundation Models to Improve Vision-Language Model 3D Spatiotemporal Reasoning for Autonomous Driving 

 


Introduction
 

A Master's thesis is an excellent way to get closer to TRATON Group R&D and build relationships for the future. 

 


Background 

Vision-Language Models (VLMs) are increasingly being explored for autonomous driving and embodied AI due to their strong compositional and logical reasoning capabilities and generalizability. However, while these models excel at general visual understanding and reasoning, they often struggle with the fine-grained spatiotemporal perception required for safe and precise driving decisions. Conversely, modern Vision Foundation Models (VFMs) such as generative video world models (e.g., Wan2.1 [5], Cosmos [3]) and 3D geometry models (e.g., VGGT [6], DepthAnythingV2 [1]) have demonstrated exceptional ability in learning visual representations from large-scale data, capturing geometry and dynamics. This presents a highly promising opportunity by transferring the rich, low-level spatiotemporal priors from VFMs into the high-level reasoning frameworks of VLMs. However, the optimal strategies for combining these paradigms, without drastically increasing computational overhead and without deteriorating the language-aligned reasoning of the VLM (catastrophic forgetting), remain an open and exciting research challenge. 



Objective 

The objective of this thesis is to investigate methods for enhancing the spatiotemporal reasoning capabilities of VLMs by leveraging representations from state-of-the-art VFMs. The exact research questions will be formulated together with the student based on current literature, the student’s interests, and the specific project direction. Possible research directions can include: 

  • Exploring distillation methods to transfer spatiotemporal knowledge from VFMs to VLMs [2, 4]. 

  • Investigating architectural modifications to e.g., insert or fuse VFM features into the VLM pipeline [7, 8]. 

  • Evaluating the enhanced model on autonomous driving benchmarks that require complex temporal and spatial reasoning (e.g., spatial Vision-Question-Answering (VQA), 3D perception-based tasks, driving trajectory prediction). 



The Project Offers 

  • The opportunity to work with state-of-the-art deep learning architectures at the intersection of vision and language. 

  • Access to large-scale autonomous driving datasets and premium GPU computing infrastructure (e.g., A100 40-80Gb clusters). 

  • Close collaboration with researchers working on next-generation AI for autonomous vehicles. 

  • A chance to contribute to an active research area with the potential for scientific publication. 



Who are we looking for? 

We are looking for a Master’s student in Computer Science, Robotics, Engineering Physics, Electrical Engineering, Applied Mathematics, or a related field. 

Experience with one or more of the following is beneficial: 

  • Deep learning and machine learning 

  • Computer vision and/or natural language processing 

  • PyTorch or similar frameworks 

  • Autonomous driving or robotics 

The planned thesis start is January 2027. 


Number of students: 1 

Start date for the thesis work: [To be agreed] 

Estimated time required: 20 weeks, full time (30 credits) 



Contact persons and supervisors 

Jesper Eriksson, jesper.ericsson@scania.com; Thomas Gustafsson, thomas.gustafsson@scania.com; Mohammad Nazari, mohammad.nazari@scania.com; 

Hiring Manager: Maria Linnarsson, maria.linnarsson@scania.com  



Application 


Your application must include a CV, personal letter, and transcript of grades. 

A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified. 

 

References 

Requisition ID:  33772
Number of Openings:  1.0
Part-time / Full-time:  Full-time
Permanent / Temporary:  Temporary
Country/Region:  SE
Location(s): 

Södertälje, SE, 151 38

Required Travel:  0%
Workplace:  Hybrid