This position is within one of TRATON’s companies.

Thesis Worker 30 hp - Adaptive Hierarchical Inference with Foundation models

30 hp - Adaptive Hierarchical Inference with Foundation models for Commercial-Vehicle Edge-Cloud Systems 

 

 

Introduction

 

Thesis work is an excellent way to get closer to Scania and build relationships for the future. Many of today's employees began their Scania career with their degree project.

 

Background 

 

TRATON GROUP is one of the world’s leading commercial vehicle manufacturers. Its brands include Scania, MAN, International and Volkswagen Truck & Bus. The group offers light commercial vehicles, trucks and buses, supported by financing, charging and digital logistics services. Through its global operations, production sites, and sales and service networks, TRATON has access to diverse vehicle platforms, real-world data and fleet environments. This provides a strong basis for developing and validating solutions for more sustainable and efficient transportation. 

 

Within TRATON’s Cloud and Embedded Platform department, we develop solutions for connected vehicles and IoT platforms. This work supports TRATON’s growing focus on communication, digital services and smart transport. Advanced data analysis is an important part of these developments. Modern commercial vehicles generate diverse data from vehicle telemetry, time-series sensors, cameras, radar, LiDAR and operational logs. High-performance units and embedded platforms make it possible to process much of this data close to the vehicle, reducing latency, network dependence and unnecessary data transfer. Foundation models could support predictive maintenance, anomaly detection, diagnostic assistance, multimodal perception and reasoning. However, the most capable models require cloud-scale computing resources, while onboard systems face strict limits on computation, memory, energy and communication. This research therefore focuses on hierarchical inference. A compact model performs real-time analysis onboard, while a larger cloud-based model is used only when the local prediction is uncertain and the latency and resource conditions allow it. 

 

Objective

The objective is to design and evaluate an uncertainty-aware, adaptive inference cascade across sensing or embedded devices, a vehicle high-performance unit (HPU), and cloud resources. The focus of the work will be on efficient inference, confidence estimation, and orchestration rather than on development of foundation models. The problem definition of the thesis is: How can calibrated uncertainty and vehicle operating conditions be used to decide, per input, whether a compact onboard model should answer locally or defer to a larger model, while meeting task-quality and latency requirements and limiting communication and resource use? 

 

Together with the supervisors, the student will select one representative commercial-vehicle use case, such as predictive maintenance or anomaly detection, visual diagnostic assistance, time-series reasoning, or multimodal situational awareness. To keep the 20-week scope feasible, the thesis should prioritize one use case, one main data modality (or a clearly defined multimodal pair), and one primary algorithmic contribution: calibrated uncertainty and adaptive cloud escalation. The primary scope is the vehicle HPU-to-cloud cascade; sensor-level processing may be included only when it directly supports the chosen use case. The exact data, models, and target hardware will be defined based on availability and confidentiality requirements. 

 

Illustrative three-tier deployment concept 

(The focus will most likely be on Tier 2 and 3; Tier 1 is mentioned for context)

Execution tier 

Candidate approach 

Primary purpose 

Tier 1: Sensor / Embedded layer 

Lightweight preprocessing or tiny agent; signal-quality and uncertainty indicators 

Reduce data, detect degraded sensing, and prepare features for onboard inference 

Tier 2: Vehicle HPU layer 

Small specialized or multimodal foundation model 

Low-latency local inference with a  

or uncertainty score 

Tier 3: Cloud layer 

Large specialized or multimodal foundation model 

Enhanced reasoning for difficult cases when latency, connectivity, and policy permit 



Job description

  1. Conduct a literature study on compact foundation models, hierarchical inference, uncertainty estimation and calibration, learning to defer, and adaptive resource aware edge-cloud offloading. 

  1. Select and deploy a compact open-source model on a representative edge or HPU platform. Characterize model quality, latency, throughput, memory footprint, model size, and, where feasible, energy consumption. 

  1. Implement and evaluate a confidence or uncertainty-estimation method for the local model. 

  1. Develop an adaptive threshold-learning or cloud-escalation strategy that jointly considers uncertainty, task quality, and other relevant factors such as latency and operational cost. 

  1. Compare local-only, remote-only, static-threshold, adaptive-routing, and retrospective-oracle baselines. Report repeated trials and uncertainty for task quality, calibration, offload rate, latency, communication, memory, and energy where feasible.  

  1. Deliver reproducible code, experiment configurations, documented datasets and models, analysis scripts, and a clear account of limitations and deployment assumptions. 
     

An optional extension is a theoretical performance model, regret analysis, or a clearly scoped multimodal experiment. 

 

Expected outcome 


The expected result is a reproducible prototype and evaluation pipeline that combines an edge-deployable foundation model, calibrated uncertainty estimates, and an adaptive escalation policy. Subject to result quality and confidentiality constraints, the work may also contribute to a scientific publication. 

References 

[1] V. N. Moothedath, J. P. Champati, and J. Gross, "Getting the Best Out of Both Worlds: Algorithms for Hierarchical Inference at the Edge," IEEE Transactions on Machine Learning in Communications and Networking, 2024. 

[2] C.-H. Chang, A. P. Behera, S. Zhang Pettersson, and J. Gross, "A Cost-Aware Hierarchical Cascade for Anomaly Detection at the Edge in Connected Vehicles," ACM/IEEE Symposium on Edge Computing, 2025. 

[3] Z. Liu et al., "MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases," ICML 2024. https://arxiv.org/abs/2402.14905

Education/line/direction 


Master's programmes in Data Science, Machine Learning, Computer Science, Embedded Systems, Electrical Engineering, Engineering Physics, Applied Mathematics, or a similar field. Strong Python skills and experience with a deep-learning framework are expected. Familiarity with foundation-model inference, uncertainty estimation or calibration, embedded or edge platforms, multimodal or time-series learning, or performance profiling is beneficial. 

Number of students 

1 

Start date for the thesis project 

January 2027, or as agreed 

Estimated timescale 

20 weeks 

Location 

Within the TRATON GROUP; exact location and hybrid arrangement to be agreed 




Contact person and supervisor

Sophia Zhang Pettersson, Senior Data Scientist 
sophia.zhang.pettersson@scania.com 

Juan Carlos Andresen, Unit Manager, Research Advisor 

juan-carlos.andresen@scania.com 
 


Application

Your application should contain a CV, personal letter, and copies of grades. 

Application period and deadline: 2026-10-31, applicants may be assessed continuously until the position is filled. 

A background check might be conducted for this position. We are conducting interviews continuously and may close the recruitment earlier than the date specified.
Requisition ID:  33387
Number of Openings:  1.0
Part-time / Full-time:  Full-time
Permanent / Temporary:  Temporary
Country/Region:  SE
Location(s): 

Södertälje, SE, 151 38

Required Travel:  0%
Workplace:  On-site