LLM Research Scientist (Post-training)
Post-training for domain models: SFT / RLHF / DPO, data mix, eval, plus RAG / agent orchestration. Papers and production in the same loop.
- MSc+ in CS or related; PhD preferred
- 3+ years NLP/LLM research; strong Python / PyTorch
- Has run full SFT / RLHF / DPO loops end to end
- First-author at ACL / EMNLP / NeurIPS / ICLR or equivalent open-source impact

