Research Intern – Reinforcement Learning for Large Foundation Models

Tencent·Singapore·Research / Applied Science

Tencent is hiring a Research Intern – Reinforcement Learning for Large Foundation Models in Singapore. Posted 2026-08-24; applications close 2026-10-23 (in 59 days).

Role details

Technology Engineering Group (TEG)

Technology Engineering Group (TEG) supports the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers. TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia, TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms, and supporting business innovation.

What the Role Entails

Research directions include but are not limited to:

  • RL Algorithms for Reasoning Models: Design robust RL training recipes (PPO/GRPO/GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.
  • RL for Autonomous Agents: Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies, and self-evolving agent behaviors in complex interactive environments.
  • Reward Modeling & Optimization: Develop reward signals and regularization techniques that go beyond outcome-based rewards. Explore token-level reward shaping, entropy-based regularization, and learned reward models that generalize across tasks.

Who We Look For

  • Enrolled in a PhD or Master’s program in computer science, machine learning, or a related field.
  • Solid understanding of RL fundamentals (policy gradients, PPO, GRPO, etc.) and hands-on experience applying them to LLM training.
  • Strong programming skills in Python and PyTorch; experience training models on multi-GPU setups; comfortable debugging training instability at scale.
  • Ability to independently read, critique, and build on recent research papers.
  • Published or submitted first-author papers at top ML/NLP venues (NeurIPS, ICML, ICLR, ACL, AAAI, EMNLP, etc.), or demonstrated equivalent research maturity through preprints and technical reports.
  • Familiarity with LLM post-training (RLHF, DPO, GRPO), model merging, or agent frameworks (ReAct, tool-use) is a strong plus.
  • Experience with large-scale distributed training (DeepSpeed, FSDP, Megatron) and open-source contributions are welcome.

More open roles at Tencent

Other open Research / Applied Science roles

Applying to this role

This Research Intern – Reinforcement Learning for Large Foundation Models role at Tencent runs through the firm's own careers portal and expects a CV and cover letter written specifically for the posting, not a portable submission carried across firms. Jorb AI's application agent tailors a CV and cover letter from your background to this posting and tracks the role alongside the rest of your applications.

Jorb AI tracks details for Research Intern – Reinforcement Learning for Large Foundation Models at Tencent. Postings refresh hourly from primary careers pages. Job details mirror the firm's posting; the apply link goes directly to the source. Last refreshed 2026-08-24.

Tencent careers

Save this role and tailor your cover letter with Jorb AI.