Publiée 22 août 2026
PhD Position F/M RobDNA:Robust data retrieval for DNA-based data storage
Inria
Rennes, Hauts-de-France 60420, France
CDI
A propos du centre ou de la direction fonctionnelle
The Inria Centre at Rennes University is one of Inria's nine centres and has more than thirty research teams. The Inria Centre is a major and recognized player in the field of digital sciences. It is at the heart of a rich R&D and innovation ecosystem: highly innovative PMEs, large industrial groups, competitiveness clusters, research and higher education players, laboratories of excellence, technological research institute, etc.
Contexte et atouts du poste
Context The volume of data generated worldwide is projected to approach 180 zettabytes (ZB) per year by 2025 [1]. However, current storage technologies face significant limitations in scaling sustainably to such volumes. One promising solution to address these challenges is DNA-based data storage, which offers several advantages, including extremely high data density, long-term retention, and low energy consumption [2].
From a density perspective, DNA can theoretically store up to 10 terabytes per mm³, which would allow all data generated throughout human history to be stored within a cube of approximately 30 cm per side [3]. In terms of retention, DNA can remain readable for centuries under suitable conditions, whereas conventional storage media typically degrade within decades [3]. Furthermore, DNA storage is energy-efficient, as it can be preserved at ambient temperature provided it is protected from light and humidity.
Mission confiée
Goal The goal of the project is to develop an algorithm to allow robust retrieval of data in the context of DNA-based data storage.
Principales activités
Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges:
(i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T).
(ii) DNA synthesis: writing data through the physical synthesis of DNA strands.
(iii) DNA sequencing: reading the stored data by sequencing DNA.
(iv) Data retrieval: reconstructing the original digital data from the sequenced symbols.
This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii).
Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known
e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction.
Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8].
In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9].
Bibliography
[1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. "The digitization of the world from edge to core.Framingham: International Data Corporation", 16:1-28, 2018.
[2] Luis Ceze, Jeff Nivala, and Karin Strauss. "Molecular digital data storage using DNA". NatureReviews Genetics, 20(8):456-466, 2019.
[3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. "Nucleic acid memory". Nature materials, 15(4):366-370, 2016.
[4] Delahaye, Clara, and Jacques Nicolas. "Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape," November 23, 2020. https://hal.inria.fr/hal-03123133.
[5] ---. "Sequencing DNA with Nanopores: Troubles and Biases." PLoS ONE, October 1, 202
[6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025.
[7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025.
[8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, "Embracing errors is more effective than avoiding them through constrained coding for DNA data storage," Allerton Conference on Communication, Control, and Computing 2023.
[9] F. De Moor, O. Boullé, D. Lavenier, De Bruijn Graph Partitioning for Scalable and Accurate DNA Storage Processing, BioRxiv, 2025.
[10] M. C. Davey and D. J. C. Mackay, "Low density parity check codes over GF(q)", Information Theory Workshop, 1998, pp. 70-71.
[11] D. Declercq and M. Fossorier, "Decoding algorithms for nonbinary LDPC codes over GF(q)", IEEE Trans. on Commun., vol. 55, pp. 633- 643, April 2007
Avantages
Rémunération
2 300€ per month
The Inria Centre at Rennes University is one of Inria's nine centres and has more than thirty research teams. The Inria Centre is a major and recognized player in the field of digital sciences. It is at the heart of a rich R&D and innovation ecosystem: highly innovative PMEs, large industrial groups, competitiveness clusters, research and higher education players, laboratories of excellence, technological research institute, etc.
Contexte et atouts du poste
Context The volume of data generated worldwide is projected to approach 180 zettabytes (ZB) per year by 2025 [1]. However, current storage technologies face significant limitations in scaling sustainably to such volumes. One promising solution to address these challenges is DNA-based data storage, which offers several advantages, including extremely high data density, long-term retention, and low energy consumption [2].
From a density perspective, DNA can theoretically store up to 10 terabytes per mm³, which would allow all data generated throughout human history to be stored within a cube of approximately 30 cm per side [3]. In terms of retention, DNA can remain readable for centuries under suitable conditions, whereas conventional storage media typically degrade within decades [3]. Furthermore, DNA storage is energy-efficient, as it can be preserved at ambient temperature provided it is protected from light and humidity.
Mission confiée
Goal The goal of the project is to develop an algorithm to allow robust retrieval of data in the context of DNA-based data storage.
Principales activités
Challenges and envisaged approach Despite its potential, making DNA a practical and efficient storage medium requires overcoming several key challenges:
(i) Data transformation: converting digital data into a quaternary alphabet (A, C, G, T).
(ii) DNA synthesis: writing data through the physical synthesis of DNA strands.
(iii) DNA sequencing: reading the stored data by sequencing DNA.
(iv) Data retrieval: reconstructing the original digital data from the sequenced symbols.
This PhD project focuses on the first and fourth challenges, by developing joint compression and error-correction algorithms that are robust to sequencing errors arising during step (iii).
Efficient DNA storage critically depends on fast sequencing technologies, which often come at the cost of increased error rates. For example, nanopore sequencing, developed by Oxford Nanopore Technologies (ONT), enables real-time analysis but introduces relatively high error rates [4,5]. Unlike traditional sequencing technologies, nanopore sequencing produces not only substitution errors but also insertion and deletion errors. Deletions are particularly challenging, as they differ from erasure errors where the position of missing data is known
e.g., packet losses in digital communications). In the case of deletions, neither the existence nor the location of the missing symbols is known, significantly complicating error correction.
Two main approaches have been proposed in the literature to address these errors. The first consists of avoiding error-prone patterns by designing constrained DNA sequences, such as limiting homopolymers or enforcing balanced GC content [6]. The second approach embraces sequencing errors and focuses on correcting them using coding techniques [7]. These strategies are typically considered contradictory, as one seeks to prevent errors while the other assumes their presence [8].
In this project, we propose to explore an alternative route that combines both strategies: avoiding the majority of sequencing errors while correcting the remaining ones. This will be achieved by jointly structuring the compressed DNA stream and designing error-correction mechanisms tailored to nanopore sequencing. In particular, we will exploit the properties of the quaternary alphabet and the interaction between compression and error-correction algorithms. The work will build upon concepts from non-binary channel coding [10,11] and DNA-specific transcoding techniques [6], while also integrating efficient post-sequencing data retrieval methods such as those proposed in [9].
Bibliography
[1] David Reinsel-John Gantz-John Rydning, John Reinsel, and John Gantz. "The digitization of the world from edge to core.Framingham: International Data Corporation", 16:1-28, 2018.
[2] Luis Ceze, Jeff Nivala, and Karin Strauss. "Molecular digital data storage using DNA". NatureReviews Genetics, 20(8):456-466, 2019.
[3] Victor Zhirnov, Reza M Zadegan, Gurtej S Sandhu, George M Church, and William L Hughes. "Nucleic acid memory". Nature materials, 15(4):366-370, 2016.
[4] Delahaye, Clara, and Jacques Nicolas. "Nanopore MinION Long Read Sequencer: An Overview of Its Error Landscape," November 23, 2020. https://hal.inria.fr/hal-03123133.
[5] ---. "Sequencing DNA with Nanopores: Troubles and Biases." PLoS ONE, October 1, 202
[6] S. Al Sayyed and A. Roumy and T. Maugey. ``Efficient constraining of transcoding for DNA-based data storage'', IEEE International Conference on Image Processing (ICIP), 2025.
[7] R. Khabbaz, M. Antonini, S. Kas Hanna, Marker Guess & Check Plus (MGC+): An Efficient Short Blocklength Code for Random Edit Errors, International Symposium on Topics in Coding (ISTC), 2025.
[8] F. Weindel, A. L. Gimpel, R. N. Grass and R. Heckel, "Embracing errors is more effective than avoiding them through constrained coding for DNA data storage," Allerton Conference on Communication, Control, and Computing 2023.
[9] F. De Moor, O. Boullé, D. Lavenier, De Bruijn Graph Partitioning for Scalable and Accurate DNA Storage Processing, BioRxiv, 2025.
[10] M. C. Davey and D. J. C. Mackay, "Low density parity check codes over GF(q)", Information Theory Workshop, 1998, pp. 70-71.
[11] D. Declercq and M. Fossorier, "Decoding algorithms for nonbinary LDPC codes over GF(q)", IEEE Trans. on Commun., vol. 55, pp. 633- 643, April 2007
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking (after 6 months of employment) and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
Rémunération
2 300€ per month