Own our AWS platform end to end across staging and production: EKS, RDS, MemoryDB/Redis, S3, DynamoDB, IAM, networking.
Understand and reduce our infrastructure costs, from the Spark data pipeline to backups, and keep unit economics improving as customer volume grows.
Own CI/CD and developer productivity: faster pipelines, fewer flaky tests, better internal tooling.
Keep our Kubernetes workloads stable and scalable, including the services behind our AI features.
Run identity and secrets management: Keycloak, SSO, OIDC.
Build AI-agent workflows for infrastructure work, such as recurring audits of AWS for cost, security, and performance findings.
Support external penetration tests and provide infrastructure evidence for ISO 27001 and client audits.
Document core systems and runbooks so the platform does not depend on any single person.
Responsibilities
Own our AWS platform end to end across staging and production: EKS, RDS, MemoryDB/Redis, S3, DynamoDB, IAM, networking.
Understand and reduce our infrastructure costs, from the Spark data pipeline to backups, and keep unit economics improving as customer volume grows.
Own CI/CD and developer productivity: faster pipelines, fewer flaky tests, better internal tooling.
Keep our Kubernetes workloads stable and scalable, including the services behind our AI features.
Run identity and secrets management: Keycloak, SSO, OIDC.
Build AI-agent workflows for infrastructure work, such as recurring audits of AWS for cost, security, and performance findings.
Support external penetration tests and provide infrastructure evidence for ISO 27001 and client audits.
Document core systems and runbooks so the platform does not depend on any single person.
Requirements
Experience as the sole or first infrastructure owner at a startup.
5+ years of hands-on infrastructure or DevOps experience with deep AWS expertise, including cost optimization work.
Production Kubernetes operations experience.
A proper software engineer: you write production-quality code and can diagnose and fix problems inside application codebases when needed (JVM experience a plus).
You have owned complex CI/CD end to end (GitHub Actions or equivalent).
An AI-native way of working: you use agents (Claude Code or similar) for real infrastructure work and can show concrete examples.
Comfortable being the only infrastructure specialist on the team: self-directed, you find and fix problems without being asked.
Bonus: Keycloak or similar identity provider operations (OIDC, SSO, secrets management).
Spark or large-scale data pipeline infrastructure experience.
Exposure to regulated environments (ISO 27001, audit support).