Die Stelle
03As a Data Infrastructure & MLOps Engineer, you will own the reliability, scalability, and automation of Doodle's data and machine learning platforms.
01Design, build, and operate scalable data infrastructure for ingestion, transformation, storage, and serving.
02Develop reliable batch and streaming data pipelines that support product analytics, business intelligence, and machine learning use cases.
03Establish data platform standards for performance, availability, observability, documentation, and cost management.
04Improve data discoverability and usability through data cataloguing, lineage, ownership, and quality processes.
05Build and maintain MLOps workflows covering experimentation, data and model versioning, training, evaluation, deployment, and rollback.
06Operate machine learning workloads in production, including model serving, feature pipelines, scheduled retraining, and inference infrastructure.
07Partner with data scientists and software engineers to turn prototypes into reliable, maintainable production services.
08Introduce repeatable approaches for model validation, monitoring, drift detection, performance measurement, and incident response.
09Manage cloud-based data and machine learning infrastructure using infrastructure as code and automated deployment practices.
10Build secure, reproducible environments for development, testing, and production.
11Improve platform efficiency through automation, capacity planning, resource optimisation, and sensible cost controls.
12Contribute to platform architecture decisions and help evolve Doodle's technical foundations as the business grows.
13Define and maintain service level objectives, operational runbooks, alerts, dashboards, and on-call processes for critical data and ML systems.
14Protect sensitive data through appropriate access controls, encryption, secrets management, retention policies, and secure development practices.
15Support compliance, privacy, and responsible AI requirements by making data and model operations traceable, auditable, and well documented.
16Investigate incidents, lead root cause analysis, and implement preventative improvements across the platform.
17Work with product, engineering, analytics, data science, security, and operations teams to understand requirements and deliver practical platform solutions.
18Create clear documentation, reusable tooling, and self-service workflows that enable teams to work independently.
19Contribute to engineering standards, technical planning, code reviews, and knowledge sharing.