Die Stelle
03We are looking for a Senior Research Engineer or Research Scientist to own the training and optimisation side of our complete post-training stack. Starting from pretrained checkpoints, you will design, implement, scale, and operate the methods required to produce capable, reliable, and controllable production models. Your scope will include supervised fine-tuning, preference optimisation, reinforcement learning, reward and verifier integration, policy distillation or consolidation, and distributed training. This is not a single-GPU fine-tuning or adapter-only role. You should be comfortable operating training workloads where memory, communication, rollout generation, hardware topology, and fault recovery must be designed together.
01Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate.
02Design and execute full-parameter and parameter-efficient SFT.
03Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods.
04Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long-horizon agent tasks.
05Integrate reward models, verifiers, critics, graders, and process- or outcome-based rewards.
06Build scalable rollout-generation systems for iterative and on-policy training.
07Design multi-stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation.
08Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism.
09Select sharding, precision, checkpointing, optimiser, batch-size, sequence-length, and activation-recomputation strategies.
10Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs.
11Profile and improve accelerator utilisation, communication efficiency, data loading, and end-to-end training time.
12Diagnose numerical instability, communication failures, out-of-memory errors, stragglers, checkpoint issues, and convergence regressions.
13Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting.
14Build reliable checkpointing, recovery, monitoring, and reproducibility procedures.
15Collaborate closely with data, evaluation, infrastructure, and inference teams.
16Contribute clean, tested code, technical reports, and operational runbooks.