Publiée 25 juillet 2026
PhD Position F/M Autotelic Exploration for Automatic Evaluation of Large Generative AI Models: Adaptive Discovery and Mapping of Capabilities and Vulnerabilities
Inria
Talence, Nouvelle-Aquitaine 33400, France
CDI
Contexte et atouts du poste
Context: The Flowers team (Inria) has pioneered fundamental research in curiosity-driven autotelic learning and open-ended AI (Oudeyer et al., 2007; Colas et al., 2022; Gaven et al., 2025), and applications in domains ranging from robotics (Forestier et al., 2022) to educational technologies, with publications at NeurIPS, ICLR, and Nature Reviews. This project aims to extend the fundamental research in autotelic and open-ended AI and apply it to the societally important challenge of AI safety and discovery of emergent failures and capabilities.
Evaluating large language models (LLMs) represents a major challenge given their rapid evolution. Traditional benchmarks, often saturated and present in training data, struggle to reveal the true diversity of behaviors. This project proposes using autotelic curiosity algorithms to drive exploration by genAI models (Pourcel et al., 2024), inspired by human curiosity (Gottlieb & Oudeyer, 2018), to automatically generate adaptive evaluations that co-evolve with model capabilities.
Mission confiée
Objectives and Methods. We aim to develop a unified framework combining automatic benchmark generation and vulnerability discovery. The approach relies on autotelic curiosity algorithms, of which quality-diversity (QD) algorithms form a particular family. These algorithms systematically explore a space of goals/problems by simultaneously optimizing local quality and global diversity according to semantic descriptors. We will generalize the ACES (Pourcel et al., 2024) and ACD (Lu et al., 2025) approaches to create adaptive benchmarks covering multiple domains and generating problems that are both diverse and unlikely to appear in training data.
The iterative process maintains an evolving archive of problems/attacks: at each iteration, the system samples a target niche, selects examples from the archive, then uses a generator LLM to create new tests conditioned on those examples. We will adapt these principles to redteaming approaches (Samvelyan et al., 2024; Lee et al., 2024), e.g. by using the recently developed methods for characterizing biases in large models (Kovac et al., 2024; Perez et al., 2025).
A key contribution is the use of LLMs to automatically cluster vulnerability findings into human-interpretable categories and produce synthetic reports detailing the strengths and weaknesses of the evaluated models. We will integrate a meta-learning dimension allowing the system to learn to predict the most informative niches based on past history, thereby optimizing its exploration strategy. On the methodological side, this project extends ACES and ACD to new domains, unifies benchmark generation and red teaming, and develops self-adaptive methods that evolve through meta-learning. Practical impact includes reducing evaluation costs, automatically discovering vulnerabilities before deployment, and proactively identifying dangerous emergent behaviors. This project thus proposes a paradigm shift: evaluation systems that continuously explore the space of possible behaviors (and their diversity), remaining relevant in the face of the rapid evolution of generative AI.
The project affords national and international collaborations with other research teams (both academic and industrial).
Expected output: publications in major AI conferences (Neurips, ICLR, ICML, etc), impactful open-source code, and societal impact through collaboration with (inter)national public agencies on AI safety.
Principales activités
See above
Compétences
Required background:
To be eligible, candidate must hold a master-level (or equivalent) diploma.
A plus: Familiarity with LLMs (transformers, RLHF, prompting strategies), quality-diversity algorithms, or red teaming literature; prior research experience (internships).
Avantages
Rémunération
2300€ per month before taxs
Context: The Flowers team (Inria) has pioneered fundamental research in curiosity-driven autotelic learning and open-ended AI (Oudeyer et al., 2007; Colas et al., 2022; Gaven et al., 2025), and applications in domains ranging from robotics (Forestier et al., 2022) to educational technologies, with publications at NeurIPS, ICLR, and Nature Reviews. This project aims to extend the fundamental research in autotelic and open-ended AI and apply it to the societally important challenge of AI safety and discovery of emergent failures and capabilities.
Evaluating large language models (LLMs) represents a major challenge given their rapid evolution. Traditional benchmarks, often saturated and present in training data, struggle to reveal the true diversity of behaviors. This project proposes using autotelic curiosity algorithms to drive exploration by genAI models (Pourcel et al., 2024), inspired by human curiosity (Gottlieb & Oudeyer, 2018), to automatically generate adaptive evaluations that co-evolve with model capabilities.
Mission confiée
Objectives and Methods. We aim to develop a unified framework combining automatic benchmark generation and vulnerability discovery. The approach relies on autotelic curiosity algorithms, of which quality-diversity (QD) algorithms form a particular family. These algorithms systematically explore a space of goals/problems by simultaneously optimizing local quality and global diversity according to semantic descriptors. We will generalize the ACES (Pourcel et al., 2024) and ACD (Lu et al., 2025) approaches to create adaptive benchmarks covering multiple domains and generating problems that are both diverse and unlikely to appear in training data.
The iterative process maintains an evolving archive of problems/attacks: at each iteration, the system samples a target niche, selects examples from the archive, then uses a generator LLM to create new tests conditioned on those examples. We will adapt these principles to redteaming approaches (Samvelyan et al., 2024; Lee et al., 2024), e.g. by using the recently developed methods for characterizing biases in large models (Kovac et al., 2024; Perez et al., 2025).
A key contribution is the use of LLMs to automatically cluster vulnerability findings into human-interpretable categories and produce synthetic reports detailing the strengths and weaknesses of the evaluated models. We will integrate a meta-learning dimension allowing the system to learn to predict the most informative niches based on past history, thereby optimizing its exploration strategy. On the methodological side, this project extends ACES and ACD to new domains, unifies benchmark generation and red teaming, and develops self-adaptive methods that evolve through meta-learning. Practical impact includes reducing evaluation costs, automatically discovering vulnerabilities before deployment, and proactively identifying dangerous emergent behaviors. This project thus proposes a paradigm shift: evaluation systems that continuously explore the space of possible behaviors (and their diversity), remaining relevant in the face of the rapid evolution of generative AI.
The project affords national and international collaborations with other research teams (both academic and industrial).
Expected output: publications in major AI conferences (Neurips, ICLR, ICML, etc), impactful open-source code, and societal impact through collaboration with (inter)national public agencies on AI safety.
Principales activités
See above
Compétences
Required background:
- Strong mathematical foundations (probability, optimization, linear algebra);
- Solid Python programming and experience training/fine-tuning neural networks/genAI;
- Solid working knowledge in post-training, reinforcement learning, NLP.
To be eligible, candidate must hold a master-level (or equivalent) diploma.
A plus: Familiarity with LLMs (transformers, RLHF, prompting strategies), quality-diversity algorithms, or red teaming literature; prior research experience (internships).
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
Rémunération
2300€ per month before taxs