Research Intern – Reinforcement Learning for Large Foundation Models
Tencent·Singapore·Research / Applied Science
Tencent is hiring a Research Intern – Reinforcement Learning for Large Foundation Models in Singapore. Posted 2026-08-24; applications close 2026-10-23 (in 59 days).
Role details
Technology Engineering Group (TEG)
Technology Engineering Group (TEG) supports the company and its business groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers. TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia, TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms, and supporting business innovation.
What the Role Entails
Research directions include but are not limited to:
- RL Algorithms for Reasoning Models: Design robust RL training recipes (PPO/GRPO/GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.
- RL for Autonomous Agents: Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies, and self-evolving agent behaviors in complex interactive environments.
- Reward Modeling & Optimization: Develop reward signals and regularization techniques that go beyond outcome-based rewards. Explore token-level reward shaping, entropy-based regularization, and learned reward models that generalize across tasks.
Who We Look For
- Enrolled in a PhD or Master’s program in computer science, machine learning, or a related field.
- Solid understanding of RL fundamentals (policy gradients, PPO, GRPO, etc.) and hands-on experience applying them to LLM training.
- Strong programming skills in Python and PyTorch; experience training models on multi-GPU setups; comfortable debugging training instability at scale.
- Ability to independently read, critique, and build on recent research papers.
- Published or submitted first-author papers at top ML/NLP venues (NeurIPS, ICML, ICLR, ACL, AAAI, EMNLP, etc.), or demonstrated equivalent research maturity through preprints and technical reports.
- Familiarity with LLM post-training (RLHF, DPO, GRPO), model merging, or agent frameworks (ReAct, tool-use) is a strong plus.
- Experience with large-scale distributed training (DeepSpeed, FSDP, Megatron) and open-source contributions are welcome.
More open roles at Tencent
- 3D Artist Intern
London · 4d ago
- Financial Accounting and Analysis Intern
Singapore · 7d ago
- WeChat Pay - Business Development Intern
Singapore · 13d ago
- Product Operations Intern (Monetization)
Singapore · 13d ago
- WeChat - Backend Engineer Intern
Singapore · 18d ago
Other open Research / Applied Science roles
- PhD Student Intern - AI Research
SAP · Singapore · 4mo ago
- PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling
Canva · London · 17d ago
- PhD Research Associate (Industry PhD Program) - Artificial Intelligence, SAP Labs Singapore
SAP · Singapore · 7mo ago
- 2027 Machine Learning Research Associate Program - PhD (New York)
Morgan Stanley · New York · 2mo ago
- Internet Measurement Research Intern - BGP Research Team
Cisco · London · 19h ago
Applying to this role
This Research Intern – Reinforcement Learning for Large Foundation Models role at Tencent runs through the firm's own careers portal and expects a CV and cover letter written specifically for the posting, not a portable submission carried across firms. Jorb AI's application agent tailors a CV and cover letter from your background to this posting and tracks the role alongside the rest of your applications.
Jorb AI tracks details for Research Intern – Reinforcement Learning for Large Foundation Models at Tencent. Postings refresh hourly from primary careers pages. Job details mirror the firm's posting; the apply link goes directly to the source. Last refreshed 2026-08-24.
