Responsibilities
-Lead the design and development of post‑training algorithms for large language models, including Supervised Fine‑Tuning (SFT), Reinforcement Learning from Human Feedback (RLHF), DPO, PPO, and other alignment techniques.
-Build and optimize data pipelines for instruction tuning, preference modeling, reward modeling, and safety alignment.
-Develop scalable training strategies to improve model helpfulness, safety, reasoning ability, and robustness across diverse tasks.
-Conduct experiments to evaluate model behavior, diagnose failure cases, and iterate on training methods to improve performance.
-Collaborate with data, infrastructure, and product teams to define post‑training objectives and integrate aligned models into production systems.
-Research and apply state‑of‑the‑art techniques in alignment, distillation, preference optimization, and model evaluation.
-Establish evaluation frameworks and benchmarks for reasoning, factuality, safety, and user experience.
-Mentor junior researchers and contribute to long‑term technical planning for model alignment and post‑training.