Publiée 22 août 2026
FR - Data Engineer - Complex Data Pipelines (On-Premise)
MARSS
Nice, Provence-Alpes-Côte d'Azur 06200, France
CDI
Job Title: Data Engineer - Complex Data Pipelines (On-Premise)
Reports to: VP AI
Location: France - Nice
Type: Full-time
The Role
We are looking for an experienced Data Engineer to join our AI team and play a key role in building our in-house data platform and pipelines.
We are looking for someone with strong professional experience who has already designed and built data pipelines from the ground up, rather than only maintaining or using existing pipelines.
This platform runs entirely on our own infrastructure, no AWS, GCP or Azure, and no managed services to fall back on. Deployed systems are genuinely air-gapped: they operate offline for extended periods and connect only briefly, on request, with limited bandwidth. Deciding what data is worth those windows, and designing pipelines that behave correctly when nodes reconnect with large backlogs, is central to the role. Your assumptions about cloud-native workflows will be tested regularly.
Our data environment is complex and varied. We work not only with traditional structured/tabular data, but also with images, video and temporal/time-series data, including data generated by sensors and tracking systems where the sequence and timing of information are critical.
You will design the architecture required to ingest, process, transform, store and make these different data sources available to our AI and software teams.
This role works alongside an MLOps Data Engineer who owns the infrastructure, environments and deployment layer. You own the data itself: architecture, models, pipelines, quality, and prioritisation. Both roles report directly to the VP AI.
Main Responsibilities
Requirements
Reports to: VP AI
Location: France - Nice
Type: Full-time
The Role
We are looking for an experienced Data Engineer to join our AI team and play a key role in building our in-house data platform and pipelines.
We are looking for someone with strong professional experience who has already designed and built data pipelines from the ground up, rather than only maintaining or using existing pipelines.
This platform runs entirely on our own infrastructure, no AWS, GCP or Azure, and no managed services to fall back on. Deployed systems are genuinely air-gapped: they operate offline for extended periods and connect only briefly, on request, with limited bandwidth. Deciding what data is worth those windows, and designing pipelines that behave correctly when nodes reconnect with large backlogs, is central to the role. Your assumptions about cloud-native workflows will be tested regularly.
Our data environment is complex and varied. We work not only with traditional structured/tabular data, but also with images, video and temporal/time-series data, including data generated by sensors and tracking systems where the sequence and timing of information are critical.
You will design the architecture required to ingest, process, transform, store and make these different data sources available to our AI and software teams.
This role works alongside an MLOps Data Engineer who owns the infrastructure, environments and deployment layer. You own the data itself: architecture, models, pipelines, quality, and prioritisation. Both roles report directly to the VP AI.
Main Responsibilities
- Design and build data pipelines from scratch, from data ingestion through processing, transformation, storage and consumption.
- Design ingestion for distributed recording nodes that are offline most of the time, including local buffering, resumable transfer, and reconciliation of late-arriving or out-of-order data on reconnection.
- Define, together with the ML team, which data is prioritised during short and bandwidth-limited connection windows.
- Design pipelines capable of handling multiple data types, including structured/tabular data, images, video and temporal/time-series data.
- Work with sensor-generated data, including sequential data where timestamps and the order of observations are important, and handle clock drift across nodes to ensure timing remains trustworthy downstream.
- Develop robust and scalable data processing solutions using Python and SQL.
- Design appropriate data models and storage approaches according to the nature and use of the data, including capacity planning and retention for large volumes of image and video data on self-managed storage.
- Own the workflow definitions in our orchestration layer: what runs, in what order, and with what retry, idempotency and backfill behaviour.
- Build processes for data ingestion, transformation, validation, quality control and traceability.
- Develop tools to support the preparation and availability of data for machine learning and AI applications.
- Work closely with Machine Learning Engineers, DevOps and Software Engineers to understand data requirements and provide appropriate solutions.
- Ensure data pipelines are reliable, maintainable and scalable as data volumes and use cases increase.
- Identify data-quality issues, including gaps and duplicates caused by node outages and retries, and define the data-quality and pipeline-health monitoring that the MLOps Engineer will implement in the deployment system.
- Define the overall architecture and technical standards for our in-house data platform.
- Document pipeline architecture, data flows and technical solutions.
Requirements
- Strong professional experience in a closely related data engineering role.
- Demonstrated experience designing and implementing data pipelines from zero (not cloud-based), including architectural and technical decisions.
- Good programming skills in Python & SQL
- Professional experience working with several different types of data
- Experience building pipelines involving at least some of the following: images, video, sensor data, time-series or other sequential data.
- Experience with systems that must tolerate unreliable or absent network connectivity and recover gracefully, offline-first, store-and-forward, edge collection or similar architectures.
- Experience running data infrastructure on bare metal or self-managed servers, rather than exclusively on managed cloud services.
- Good understanding of data ingestion, transformation, storage, validation and data-quality principles.
- Experience working with large or complex datasets.
- Strong Linux skills, including comfort with filesystems, storage, services and network troubleshooting.
- Good software engineering practices, including Git, testing, code review and documentation.
- Ability to independently investigate technical problems and propose appropriate architecture and solutions.
- Fluent English, written and spoken.