↳Bachelor's Degree in AI, Computer Science, Data Science, Statistics, Physics, Engineering, or a related technical field.
↳Proficiency in Python.
↳Experience with distributed data-processing frameworks such as Spark, Ray and workflow-orchestration tools such as Airflow.
↳Experience with processing datasets at many petabyte scale and with trillion rows.
↳Experience building datasets for training foundation models, including language or multimodal models.
↳Strong experience in data analysis, data engineering, or both.
↳Ability to communicate technical findings clearly and effectively to research, engineering, and product teams.
↳Experience evaluating dataset quality and measuring the impact of data through controlled model-training experiments.