W
Posted Yesterday•Bengaluru, Karnataka, India

Research Scientist - RLHF, RLAIF & Reward Modeling

MiddleOn-site (Bengaluru)Salary undisclosed
Required Skills
PythonPyTorchLLMs
Job Description

This role is for one of Weekday’s clients
Salary range: Rs 5000000 - Rs 10000000 (ie INR 50 - 100 LPA)


Min Experience: 3+ years
Location: Bengaluru, Karnataka, India
JobType: full-time

We are looking for a highly skilled and research-oriented Research Scientist with 3–6 years of experience in machine learning, reinforcement learning, and large language model (LLM) alignment. The ideal candidate will have strong hands-on experience with Reinforcement Learning from Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), and Reward Modeling, and will contribute to developing and improving advanced AI systems.

You will work on research problems related to model alignment, preference learning, reward optimization, evaluation, and post-training. This role requires a strong understanding of modern machine learning techniques, the ability to translate research ideas into working systems, and experience conducting rigorous experiments on large-scale models.

Requirements

Key Responsibilities

  • Design, implement, and evaluate RLHF pipelines for training and aligning large language models with human preferences.
  • Develop and improve RLAIF methodologies using AI-generated feedback, preference signals, and automated evaluation frameworks.
  • Build, train, and validate reward models th

Similar Openings in AI & Machine Learning

View all in category➔