Publiée 13 août 2026
Open-sourcing and Privacy--Utility Evaluation of a Foundation Model for the French National Health Data System (SNDS)
Inria
Palaiseau, Île-de-France 91120, France
CDI
A propos du centre ou de la direction fonctionnelle
Created in 2008, the Inria Saclay Center is located at the heart of the Paris-Saclay scientific
and technological excellence cluster, which alone accounts for 15% of French research.
Serving the development of the Université Paris-Saclay and the Institut Polytechnique de
Paris, the Inria Saclay center employs 80 people in research support services and 500
scientists of 54 nationalities.
Benefiting from continuous growth, the center now has a total of 42 project-teams and two in
the process of being created, including 21 jointly with the Institut Polytechnique de Paris, 16
with the Université Paris-Saclay, as well as 7 Inria EPs, including one in collaboration with
Onera and one with the Pôle Universitaire Centre Val de Loire. These research teams are
spread over more than ten sites.
[Sources Numin : https://numin.inria.fr/portal/g/:spaces:yt54gs/centre_inria_de_saclay/lecentre ]
Contexte et atouts du poste
Inria's SODA team (Statistics, Optimization, and Data Science for Health), in partnership with the French
Directorate for Research, Studies, Evaluation and Statistics (Drees, Ministry of Health), is offering a
15 month position (research engineer or postdoctoral researcher, depending on the candidate's
profile and experience) to open-source and evaluate the privacy--utility trade-off of openSNDS2vec,
a foundation model trained on the healthcare trajectories of the entire French population, as part of the
SeqNDS project.
We are looking for a candidate with a strong background in machine learning, statistics, or applied
mathematics, with hands-on experience in data science and software engineering. Experience with privacy-preserving
machine learning, differential privacy, or health data is a strong asset but not a strict requirement. Excellent
Python skills are essential, together with good software engineering practices (testing, packaging, documentation, version control). A genuine interest in making advanced methods accessible to non-specialists, and in working with people from very different backgrounds such as statisticians, epidemiologists, legal experts, physicians, is central to this role. Fluency in written and spoken French and English is required.
The position will be split between Inria's SODA team at Inria Saclay and the Drees "Innovation and Evaluation in Health" Lab, 78 rue Olivier de Serres, Paris 15 extsuperscript{e}. Applications (CV and cover letter) should be sent by email to Audrey Bergès ( [email protected] ) and Judith Abécassis ( [email protected] ). The position is open until filled, with a target start date in October 2026.
The SNDS (Système national des données de santé) is one of the richest medico-administrative databases in
the world, covering the healthcare trajectories of more than 70 million French residents over a 5-year depth
of history. Since 2024, the SeqNDS project has applied state-of-the-art natural language processing methods (Word2Vec, BERT, GPT-style architectures) to these sequences of care, producing several foundation models inspired by Life2Vec [3], BEHRT [2], and Delphi-2M [4]. These models achieve strong predictive performance across mortality and 184 ICD-10-based diagnostic targets, and have been described in a recent Drees publication [1].
The next step is to share the resulting representations (ie, vector embeddings for each medical code
(ICD-10, ATC, CCAM, NABM, CIP, etc.)) as an open-source foundation model, openSNDS2vec, so that any
team working on health data structured around standard nomenclatures (including hospital data warehouses)
can benefit from knowledge learned on the exhaustive French population, without needing direct access to
individual-level SNDS data. Because these representations are derived from highly sensitive health data, a
rigorous, quantified privacy audit is required before any public release.
References
[1] Tristan Haugomat, Aurélia Manns, Gladys Baudet, Judith Abécassis, Gaël Varoquaux, and Milena Suarez
Castillo. Prédire la suite d'un parcours de soins dans le système national des données de santé. Drees
Méthodes, (25):1-65, 2026.
[2] Yikuan Li, Shishir Rao, José Roberto Ayala Solares, Abdelaali Hassaine, Rema Ramakrishnan, Dexter Canoy,
Yajie Zhu, Kazem Rahimi, and Gholamreza Salimi-Khorshidi. Behrt: transformer for electronic health records.
Scientific reports, 10(1):7155, 2020.
[3] Germans Savcisens, Tina Eliassi-Rad, Lars Kai Hansen, Laust Hvas Mortensen, Lau Lilleholt, Anna Rogers,
Ingo Zettler, and Sune Lehmann. Using sequences of life-events to predict human lives. Nature computational
science, 4(1):43-56, 2024.
[4] Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan
Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative
transformers. Nature, 647(8088):248-256, 2025.
Mission confiée
Missions :
- Open-source packaging and code sharing. Prepare, document, and release the SeqNDS codebase, including openSNDS2vec and the companion data-preparation tool SNDS2table, as clean, tested, reusable open-source software (Apache-2 license), suitable for reuse by teams without deep learning or AI expertise, both on Drees infrastructure and on the French Health Data Hub Platform (Plateforme des données de santé, PDS), or the CNAM platform.
- Privacy-utility trade-off of the shared model. Design and implement a k-anonymization procedure for the co-occurrence matrix underlying openSNDS2vec, so that a co-occurrence between medical codes is only released if it appears in the trajectories of at least k distinct individuals. Systematically characterize the trade-off between the choice of k, the risk of re-identification, and the retained predictive utility of the embeddings (measured via the AUC on the mortality and diagnosis-prediction benchmarks already established in SeqNDS).
- Privacy attacks and robustness evaluation using CNIL's PANAME framework. Apply and adapt privacy attacks (membership inference) using the PANAME framework developed with Inria's Premedical team and the CNIL, to empirically evaluate the residual re-identification risk of the de-identified representations before any public release, in direct collaboration with the Premedical team.
- Pedagogical example gallery for epidemiologists. Design and write a gallery of concrete, well-documented usage examples (in the spirit of scikit-learn's example gallery), in R and Python, showing how to query and reuse the vector representations in classical epidemiological workflows (e.g. enriching a small hospital cohort, building predictive models, exploring patient trajectory clusters). This work will be carried out in close collaboration with the CEPHEPI epidemiology network to ensure the material is accessible to non-AI-specialist researchers and clinicians.
- Documentation and dissemination. Contribute to the scientific write-up of the de-identification methodology (submission to a peer-reviewed journal), and to presentations to Drees leadership, Inria teams, the CNIL, and the broader SNDS user community (e.g. PDS meetups, EMOIS, EPICLIN).
Principales activités
Profile
- PhD (for a postdoctoral position) or engineering/master's degree with significant applied experience (for a research-engineer position), in machine learning, statistics, applied mathematics, or a related field;
- Strong, demonstrable coding skills in Python, standard scientific libraries (scikit-learn, pandas), and good software-engineering practices;
- Genuine interest or willingness to build expertise in privacy-preserving machine learning and differential privacy; prior exposure to re-identification attacks or privacy audits is a plus;
- Strong pedagogical skills and enjoyment of writing clear, didactic documentation and examples for non-specialist audiences;
- Ease and enthusiasm for working with people from very different backgrounds (statisticians, epidemiologists, physicians, legal and regulatory experts);
- Good written and oral communication skills in French and English;
- Experience with health data (SNDS, PMSI, EHR, or similar) is appreciated but not required.
Environment
The successful candidate will join Inria's Soda team ( https://team.inria.fr/soda/ ), led by Gaël Varoquaux. Soda is doing computational and statistical research, both fundamental and applied, to harness large databases on health and society, and is known for developing scikit-learn (the third most downloaded machine learning library worldwide), skrub, and the tabular foundation model TabICL.
Part of the work will take place within the Drees extbf{"Innovation and Evaluation in Health'' Lab}, a multidisciplinary team of data scientists, data engineers, and statisticians, at 78 rue Olivier de Serres, Paris 15. The candidate will also interact with the CEPHEPI epidemiology team (AP-HP) for external validation of the shared representations, and with the CNIL in the context of the project's privacy-by-design review process.
The salary range is EUR 2,234-3,090 net per month, depending on experience (Inria salary grid), i.e. EUR 32,299 44,675 gross annually.
The successful candidate will benefit from a standard French employment package, including partial reimbursement of commuting costs, 7 weeks of paid annual leave, RTT days, and comprehensive social security coverage.
Avantages
Rémunération
Salary is based on the candidate's profile,
experience, and the salary scales
Created in 2008, the Inria Saclay Center is located at the heart of the Paris-Saclay scientific
and technological excellence cluster, which alone accounts for 15% of French research.
Serving the development of the Université Paris-Saclay and the Institut Polytechnique de
Paris, the Inria Saclay center employs 80 people in research support services and 500
scientists of 54 nationalities.
Benefiting from continuous growth, the center now has a total of 42 project-teams and two in
the process of being created, including 21 jointly with the Institut Polytechnique de Paris, 16
with the Université Paris-Saclay, as well as 7 Inria EPs, including one in collaboration with
Onera and one with the Pôle Universitaire Centre Val de Loire. These research teams are
spread over more than ten sites.
[Sources Numin : https://numin.inria.fr/portal/g/:spaces:yt54gs/centre_inria_de_saclay/lecentre ]
Contexte et atouts du poste
Inria's SODA team (Statistics, Optimization, and Data Science for Health), in partnership with the French
Directorate for Research, Studies, Evaluation and Statistics (Drees, Ministry of Health), is offering a
15 month position (research engineer or postdoctoral researcher, depending on the candidate's
profile and experience) to open-source and evaluate the privacy--utility trade-off of openSNDS2vec,
a foundation model trained on the healthcare trajectories of the entire French population, as part of the
SeqNDS project.
We are looking for a candidate with a strong background in machine learning, statistics, or applied
mathematics, with hands-on experience in data science and software engineering. Experience with privacy-preserving
machine learning, differential privacy, or health data is a strong asset but not a strict requirement. Excellent
Python skills are essential, together with good software engineering practices (testing, packaging, documentation, version control). A genuine interest in making advanced methods accessible to non-specialists, and in working with people from very different backgrounds such as statisticians, epidemiologists, legal experts, physicians, is central to this role. Fluency in written and spoken French and English is required.
The position will be split between Inria's SODA team at Inria Saclay and the Drees "Innovation and Evaluation in Health" Lab, 78 rue Olivier de Serres, Paris 15 extsuperscript{e}. Applications (CV and cover letter) should be sent by email to Audrey Bergès ( [email protected] ) and Judith Abécassis ( [email protected] ). The position is open until filled, with a target start date in October 2026.
The SNDS (Système national des données de santé) is one of the richest medico-administrative databases in
the world, covering the healthcare trajectories of more than 70 million French residents over a 5-year depth
of history. Since 2024, the SeqNDS project has applied state-of-the-art natural language processing methods (Word2Vec, BERT, GPT-style architectures) to these sequences of care, producing several foundation models inspired by Life2Vec [3], BEHRT [2], and Delphi-2M [4]. These models achieve strong predictive performance across mortality and 184 ICD-10-based diagnostic targets, and have been described in a recent Drees publication [1].
The next step is to share the resulting representations (ie, vector embeddings for each medical code
(ICD-10, ATC, CCAM, NABM, CIP, etc.)) as an open-source foundation model, openSNDS2vec, so that any
team working on health data structured around standard nomenclatures (including hospital data warehouses)
can benefit from knowledge learned on the exhaustive French population, without needing direct access to
individual-level SNDS data. Because these representations are derived from highly sensitive health data, a
rigorous, quantified privacy audit is required before any public release.
References
[1] Tristan Haugomat, Aurélia Manns, Gladys Baudet, Judith Abécassis, Gaël Varoquaux, and Milena Suarez
Castillo. Prédire la suite d'un parcours de soins dans le système national des données de santé. Drees
Méthodes, (25):1-65, 2026.
[2] Yikuan Li, Shishir Rao, José Roberto Ayala Solares, Abdelaali Hassaine, Rema Ramakrishnan, Dexter Canoy,
Yajie Zhu, Kazem Rahimi, and Gholamreza Salimi-Khorshidi. Behrt: transformer for electronic health records.
Scientific reports, 10(1):7155, 2020.
[3] Germans Savcisens, Tina Eliassi-Rad, Lars Kai Hansen, Laust Hvas Mortensen, Lau Lilleholt, Anna Rogers,
Ingo Zettler, and Sune Lehmann. Using sequences of life-events to predict human lives. Nature computational
science, 4(1):43-56, 2024.
[4] Artem Shmatko, Alexander Wolfgang Jung, Kumar Gaurav, Søren Brunak, Laust Hvas Mortensen, Ewan
Birney, Tom Fitzgerald, and Moritz Gerstung. Learning the natural history of human disease with generative
transformers. Nature, 647(8088):248-256, 2025.
Mission confiée
Missions :
- Open-source packaging and code sharing. Prepare, document, and release the SeqNDS codebase, including openSNDS2vec and the companion data-preparation tool SNDS2table, as clean, tested, reusable open-source software (Apache-2 license), suitable for reuse by teams without deep learning or AI expertise, both on Drees infrastructure and on the French Health Data Hub Platform (Plateforme des données de santé, PDS), or the CNAM platform.
- Privacy-utility trade-off of the shared model. Design and implement a k-anonymization procedure for the co-occurrence matrix underlying openSNDS2vec, so that a co-occurrence between medical codes is only released if it appears in the trajectories of at least k distinct individuals. Systematically characterize the trade-off between the choice of k, the risk of re-identification, and the retained predictive utility of the embeddings (measured via the AUC on the mortality and diagnosis-prediction benchmarks already established in SeqNDS).
- Privacy attacks and robustness evaluation using CNIL's PANAME framework. Apply and adapt privacy attacks (membership inference) using the PANAME framework developed with Inria's Premedical team and the CNIL, to empirically evaluate the residual re-identification risk of the de-identified representations before any public release, in direct collaboration with the Premedical team.
- Pedagogical example gallery for epidemiologists. Design and write a gallery of concrete, well-documented usage examples (in the spirit of scikit-learn's example gallery), in R and Python, showing how to query and reuse the vector representations in classical epidemiological workflows (e.g. enriching a small hospital cohort, building predictive models, exploring patient trajectory clusters). This work will be carried out in close collaboration with the CEPHEPI epidemiology network to ensure the material is accessible to non-AI-specialist researchers and clinicians.
- Documentation and dissemination. Contribute to the scientific write-up of the de-identification methodology (submission to a peer-reviewed journal), and to presentations to Drees leadership, Inria teams, the CNIL, and the broader SNDS user community (e.g. PDS meetups, EMOIS, EPICLIN).
Principales activités
Profile
- PhD (for a postdoctoral position) or engineering/master's degree with significant applied experience (for a research-engineer position), in machine learning, statistics, applied mathematics, or a related field;
- Strong, demonstrable coding skills in Python, standard scientific libraries (scikit-learn, pandas), and good software-engineering practices;
- Genuine interest or willingness to build expertise in privacy-preserving machine learning and differential privacy; prior exposure to re-identification attacks or privacy audits is a plus;
- Strong pedagogical skills and enjoyment of writing clear, didactic documentation and examples for non-specialist audiences;
- Ease and enthusiasm for working with people from very different backgrounds (statisticians, epidemiologists, physicians, legal and regulatory experts);
- Good written and oral communication skills in French and English;
- Experience with health data (SNDS, PMSI, EHR, or similar) is appreciated but not required.
Environment
The successful candidate will join Inria's Soda team ( https://team.inria.fr/soda/ ), led by Gaël Varoquaux. Soda is doing computational and statistical research, both fundamental and applied, to harness large databases on health and society, and is known for developing scikit-learn (the third most downloaded machine learning library worldwide), skrub, and the tabular foundation model TabICL.
Part of the work will take place within the Drees extbf{"Innovation and Evaluation in Health'' Lab}, a multidisciplinary team of data scientists, data engineers, and statisticians, at 78 rue Olivier de Serres, Paris 15. The candidate will also interact with the CEPHEPI epidemiology team (AP-HP) for external validation of the shared representations, and with the CNIL in the context of the project's privacy-by-design review process.
The salary range is EUR 2,234-3,090 net per month, depending on experience (Inria salary grid), i.e. EUR 32,299 44,675 gross annually.
The successful candidate will benefit from a standard French employment package, including partial reimbursement of commuting costs, 7 weeks of paid annual leave, RTT days, and comprehensive social security coverage.
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in
working hours) + possibility of exceptional leave (sick children, moving home, etc.) - Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment,
etc.) - Social, cultural and sports events and activities
- Access to vocational training
Rémunération
Salary is based on the candidate's profile,
experience, and the salary scales