Gå til indhold

Stage Steering LLM to inhibit biases- Saclay-H/F

Praktik 4 til 6 måneder

91400 Saclay (France)

Offentliggjort den 1. oktober 2026

  • Opslagstype

    Praktik 4 til 6 måneder

  • Sted

    91400 Saclay (France)

  • Startdato

    Så hurtigt som muligt

  • Løn

    Oplysninger ikke angivet

  • Hjemmearbejde

    Ikke specificeret

CEA illustration
Position description

Category

Engineering science

Contract

Internship

Job title

Stage Steering LLM to inhibit biases- Saclay-H/F

Subject

Multimodal Large Language Models (LLMs) are increasingly capable of processing and integrating information from multiple modalities, including text, images, audio, and speech. However, these models can inherit and amplify biases across modalities, potentially affecting the fairness, reliability, and robustness of their outputs. This internship will investigate steering-based approaches to identify and inhibit such biases at the model level, with the goal of developing more controllable and robust multimodal LLMs.

Contract duration (months)

6 mois

Job description

As an intern at the CEA, you will have the opportunity to work in a world-renowned research environment. Our teams consist of passionate and dedicated experts, providing an environment conducive to learning and collaboration. You will have access to state-of-the-art equipment and top-tier research resources to carry out your assignments. The work performed may potentially lead to a scientific publication.

Context

Through the thesis of Clément Cornet, the team has already developped several approaches of steering and other works in mechanistic interpretability [1,2]. A large part of the work is integrated into a light python library that can serve as basis for the work.

What do we expect from you ?

The intern will work on the following tasks :
  • Conduct a literature review on methods for bias in inhibition in multimodal LLMs with steering
  • Conduct experiments with available steering approaches to inhibate biases, including a rigorous quantitative evaluation on well chosen models and modalities
  • Develop novel approaches to inhibate biases with steering, in particular to determine its strength automatically
  • Develop a demonstrator to showcase the work carried out

Depending on the profile and motivation of the intern, the work may lead to a scientific publication and may be pursued with a PhD focused on a similar topic. The person will work in collaboration with Clement Cornet, Hervé Le Borgne, Romaric Besançon and possibly other researchers of the lab, depending on the direction of the work.

[1] Cornet et al (2025) Explaining How Visual, Textual and Multimodal Encoders Share Concepts, CoRR:2507.18512
[2] Cornet et al (2026) The Deleuzian Representation Hypothesis, ICLR

#Cea List

Methods / Means

Python - PyTorch

Applicant Profile

Profil :
  • Students in their final year of studies (M2 or last year of engineering school)
  • Strong foundations in machine learning and deep learning
  • Interest in mechanistic interpretability and bias of AI models
  • Python proficiency in pytorch

Position location

Site

Saclay

Job location

France, Ile-de-France, Essonne (91)

Location

Saclay

Candidate criteria

Prepared diploma

Bac+5 - Master 2

Requester

Position start date

01/12/2026

Ansøgningsfrist

Så længe stillingen er online

Uddannelsesniveau

Kandidatuddannelsesniveau eller tilsvarende

Funktion

Teknologi

Flere oplysninger om virksomheden