Skip to content
F
Field Ai

Senior Data Platform Engineer

AI Engineering
Irvine, CAOn-siteFull-time

About this role

FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data-driven approaches or pure transformer-only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field.

About the Role

We are building the data foundation that powers the full machine learning lifecycle for autonomous robotics. Our robots generate large-scale, multimodal datasets across real-world deployments, and turning that raw experience into reliable, discoverable, high-quality ML data is a core part of improving our autonomy systems.

As a Data Platform Engineer, you will help design and build the platform that manages data from ingestion through processing, validation, labeling, dataset generation, training, and evaluation.

This is not a traditional analytics data engineering role. You will work closely with ML engineers, researchers, labeling teams, robotics engineers, and infrastructure engineers to build scalable systems for robotics and ML data.

What You'll Do

  • Design and build scalable data architecture for large-scale multimodal robotics and ML datasets.
  • Build abstractions and services for ingestion, processing, datasets, metadata, lineage, and data quality.
  • Develop reliable pipelines for transforming raw robot data into versioned, ML-ready data products.
  • Define data models, schemas, contracts, and lifecycle states across data-processing workflows.
  • Build systems for tracking provenance and lineage across raw data, derived artifacts, labels, datasets, and downstream ML workloads.
  • Develop automated validation and data-quality frameworks that detect incomplete, corrupted, or unusable data early.
  • Design for incremental processing, reprocessing, backfills, and versioned transformations.
  • Improve observability and failure diagnosis across complex data workflows.
  • Partner with ML and robotics teams to understand domain-specific data requirements and turn recurring patterns into reusable platform capabilities.
  • Work closely with infrastructure/platform teams on storage, compute, orchestration, reliability, and scalability.

What We're Looking For

  • Strong experience building production data platforms or large-scale data-processing systems.
  • Strong software engineering skills, preferably Python and/or C++/Java/Go.
  • Experience with distributed data processing and workflow orchestration.
  • Experience with data lakes/lakehouses, object storage, metadata systems, schemas, and data versioning.
  • Strong understanding of data quality, lineage, reproducibility, and reliable pipeline design.
  • Experience with technologies such as S3, Airflow/Dagster, Spark/Ray, Kubernetes, Parquet, or similar systems.
  • Ability to work across ambiguous organizational and technical boundaries.
  • Strong systems-design and engineering judgment.

Nice to Have

  • Experience with ML datasets or ML infrastructure.
  • Robotics, autonomous vehicles, sensor, video, image, LiDAR, or other multimodal data.
  • Experience building internal developer/platform products.
  • Experience operating pipelines at TB/PB scale.

Explore related Data & AI jobs

Compare this role with AI engineer jobs and open roles at Field Ai.

You can also browse all Data & AI jobs to find similar openings by category, seniority, remote setup, and location.

Other jobs at Field Ai

Field Ai has no open Data & AI positions right now.

Browse all jobs
© Dataaxy. All rights reserved.Job data is gathered from publicly available sources or contributed by users.