# Research Intern – Reinforcement Learning for Large Foundation Models

[Tencent](https://www.jorb.ai/firms/tencent.md) · Singapore · [Research / Applied Science](https://www.jorb.ai/jobs/research-applied-science.md)

Tencent is hiring a Research Intern – Reinforcement Learning for Large Foundation Models in Singapore. Posted 2026-08-24; applications close 2026-10-23.

**Apply**: https://tencent.wd1.myworkdayjobs.com/Tencent_Careers/job/Singapore-CapitaSky/Research-Intern---Reinforcement-Learning-for-Large-Foundation-Models_R108024

Posted 2d ago.

## Role details

## Technology Engineering Group (TEG)

Technology Engineering Group (TEG) supports the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers. TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia, TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms, and supporting business innovation.

## What the Role Entails

Research directions include but are not limited to:

  
- **RL Algorithms for Reasoning Models:** Design robust RL training recipes (PPO/GRPO/GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.
  
- **RL for Autonomous Agents:** Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies, and self-evolving agent behaviors in complex interactive environments.
  
- **Reward Modeling & Optimization:** Develop reward signals and regularization techniques that go beyond outcome-based rewards. Explore token-level reward shaping, entropy-based regularization, and learned reward models that generalize across tasks.

## Who We Look For

  
- Enrolled in a PhD or Master&rsquo;s program in computer science, machine learning, or a related field.
  
- Solid understanding of RL fundamentals (policy gradients, PPO, GRPO, etc.) and hands-on experience applying them to LLM training.
  
- Strong programming skills in Python and PyTorch; experience training models on multi-GPU setups; comfortable debugging training instability at scale.
  
- Ability to independently read, critique, and build on recent research papers.
  
- Published or submitted first-author papers at top ML/NLP venues (NeurIPS, ICML, ICLR, ACL, AAAI, EMNLP, etc.), or demonstrated equivalent research maturity through preprints and technical reports.
  
- Familiarity with LLM post-training (RLHF, DPO, GRPO), model merging, or agent frameworks (ReAct, tool-use) is a strong plus.
  
- Experience with large-scale distributed training (DeepSpeed, FSDP, Megatron) and open-source contributions are welcome.

## Applying to this role

This Research Intern – Reinforcement Learning for Large Foundation Models role at Tencent runs through the firm's own careers portal and expects a CV and cover letter written specifically for the posting, not a portable submission carried across firms. Jorb AI's application agent tailors a CV and cover letter from your background to this posting and tracks the role alongside the rest of your applications.

[Tailor this application](https://www.jorb.ai/signup?ref=job-atom&firm=tencent&job=6a8bc32643e332145879dea2)

## More open roles at Tencent

- [Global Talent Sourcing Intern](https://www.jorb.ai/jobs/6a8d14b461add68db96c7e01.md) – London, posted 1d ago
- [3D Artist Intern](https://www.jorb.ai/jobs/6a86ed750722ce20be6cd129.md) – London, posted 6d ago
- [Financial Accounting and Analysis Intern](https://www.jorb.ai/jobs/6a8288d7a36475e9c65fe46d.md) – Singapore, posted 9d ago
- [Product Operations Intern (Monetization)](https://www.jorb.ai/jobs/6a7a9f3cc57ca1e3e687c1d4.md) – Singapore, posted 15d ago
- [WeChat Pay - Business Development Intern](https://www.jorb.ai/jobs/6a7a9f3cc57ca1e3e687c1d0.md) – Singapore, posted 15d ago

## Other open Research / Applied Science roles

- [PhD Student Intern - AI Research](https://www.jorb.ai/jobs/69cba3357cfd48bf87e17edf.md) at [SAP](https://www.jorb.ai/firms/sap.md) – Singapore, posted 4mo ago
- [2027 Machine Learning Research Associate Program - PhD (New York)](https://www.jorb.ai/jobs/6a2c0d0370b3225d345daa2b.md) at [Morgan Stanley](https://www.jorb.ai/firms/morgan-stanley.md) – New York, posted 2mo ago
- [PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling](https://www.jorb.ai/jobs/6a7590fa25ecd425810a78c5.md) at [Canva](https://www.jorb.ai/firms/canva.md) – London, posted 18d ago
- [PhD Research Associate (Industry PhD Program)  - Artificial Intelligence, SAP Labs Singapore](https://www.jorb.ai/jobs/69503ab2340e619957b1718c.md) at [SAP](https://www.jorb.ai/firms/sap.md) – Singapore, posted 8mo ago
- [R&D Technician](https://www.jorb.ai/jobs/6a8d84f85e145c7b6c8d69f8.md) at [HP](https://www.jorb.ai/firms/hp.md) – Singapore, posted 1d ago

---

Updated: 2026-08-26
Canonical: https://www.jorb.ai/jobs/6a8bc32643e332145879dea2
