The role
03Your tasks:
- Design, develop, and maintain scalable batch and real-time data pipelines on the Databricks platform using PySpark, SQL, and Delta Lake.
- Build and optimize enterprise data architectures, including medallion frameworks (Bronze, Silver, Gold), ensuring data quality, consistency, and reliability across all layers.
- Integrate manufacturing and business-critical data sources such as SAP, MES, LIMS, eQMS, and process historians into a unified data ecosystem.
- Develop and manage automated data workflows, orchestration processes, and reusable data products to support analytics, reporting, and operational excellence.
- Implement data governance, security, lineage, and compliance standards, ensuring adherence to GMP, GxP, ALCOA+, and data integrity requirements.
- Monitor and enhance data platform performance through workload optimization, cost management, and adoption of engineering best practices, testing frameworks, and CI/CD processes.
- Collaborate closely with data scientists, process engineers, business stakeholders, QA, and validation teams to translate business needs into production-ready data solutions.
- Enable advanced analytics and machine learning initiatives by providing curated datasets, semantic data models, feature engineering capabilities, and self-service analytics foundations.
01Design, develop, and maintain scalable batch and real-time data pipelines on the Databricks platform using PySpark, SQL, and Delta Lake.
02Build and optimize enterprise data architectures, including medallion frameworks (Bronze, Silver, Gold), ensuring data quality, consistency, and reliability across all layers.
03Integrate manufacturing and business-critical data sources such as SAP, MES, LIMS, eQMS, and process historians into a unified data ecosystem.
04Develop and manage automated data workflows, orchestration processes, and reusable data products to support analytics, reporting, and operational excellence.
05Implement data governance, security, lineage, and compliance standards, ensuring adherence to GMP, GxP, ALCOA+, and data integrity requirements.
06Monitor and enhance data platform performance through workload optimization, cost management, and adoption of engineering best practices, testing frameworks, and CI/CD processes.
07Collaborate closely with data scientists, process engineers, business stakeholders, QA, and validation teams to translate business needs into production-ready data solutions.
08Enable advanced analytics and machine learning initiatives by providing curated datasets, semantic data models, feature engineering capabilities, and self-service analytics foundations.