PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling

Canva·London·United Kingdom·Research / Applied Science

Canva is hiring a PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling in London. Posted 2026-08-07; applications close 2026-10-06 (in 57 days).

Role details

AI Research Internship (PhD) — September Start

Canva is redefining how the world experiences design. This internship is based in London, with a hybrid way of working that offers flexibility to work from home and collaborate in person on campus.

Our global HQ is in Sydney, Australia, and our London campus is in Hoxton Square, right in the middle of Shoreditch. The UK team comes together to connect, create, and collaborate.

Fun fact: the London team is one of the places where the AI powering Canva gets built.

About the Internship

We’re looking for current PhD students ready to bring their research into the real world and help shape the culture of AI at Canva.

This full-time, 16-week AI Research Internship starts in September. During your internship, you’ll work directly with Canva’s AI team on a live, industry-scale project, turning part of your PhD journey into real-world impact.

You’ll gain hands-on experience with real data, production infrastructure, and real deadlines, while learning from and working alongside the researchers and engineers creating Canva’s next generation of AI-powered experiences.

What You’ll Be Doing

At the moment, this role is focused on:

  • Designing and validating a rubric-guided, per-layer VLM judge for RGBA layer decomposition, calibrated against human evaluations.
  • Building VLM-based methods for automatic, human-aligned evaluation of multi-layer designs.
  • Turning VLM-based evaluators into reward functions to train generative models in a reinforcement learning setting.
  • Distilling judges into lightweight reward models that score layered images from learned representations, at a fraction of the inference cost.
  • Collaborating with research, engineering, and product teams to move findings toward production and Canva’s layered-generation roadmap.
  • Contributing to the broader research community through publication where appropriate.

The team builds the groundwork before you arrive—baselines reproduced, harnesses running, and data prepared—so you can focus on the novel parts of the work from week one.

Who You Are

You’re probably a match if:

  • You’re currently completing a PhD, ideally third year or later.
  • You have a strong diffusion or flow-matching background, with hands-on policy-gradient RL for generative models (e.g., GRPO, PPO, DPO, or similar).
  • You have experience fine-tuning VLMs (e.g., with LoRA) and designing prompts or rubrics for evaluation tasks.
  • You have reward modelling experience, including preference optimisation, pseudo-labelling, and distillation.
  • You can read a recent paper and reproduce results quickly.
  • You communicate technical work clearly in writing and presentations.
  • You enjoy working closely with researchers and engineers on hard problems.
  • You can juggle multiple threads at once and switch focus without losing context.
  • You can set your own priorities on a daily basis and between checkpoints.

Nice to Have

  • PyTorch at scale, and the ability to write research code for data processing, training, and evaluation.
  • Multi-GPU training (e.g., FSDP, DeepSpeed) and evaluation-harness engineering.
  • Experience with layered or RGBA generation, matting, or inpainting.
  • Familiarity with reward-hacking and score-compression diagnostics, or human-evaluation design.
  • Publications or open-source contributions in generative modelling, RLHF, or multimodal models.

What You Should Aim to Take Away

  • Publishable and potentially patentable contributions based on the work you complete.
  • A paper draft covering the work, with support on publication strategy.
  • Compute resources, base checkpoints, preference data, and an annotation budget (provided).
  • Four mentors: a coach for weekly 1:1s, plus a specialist lead on each workstream.
  • Work that feeds directly into a product used by hundreds of millions of people.

Additional Information

Canva makes hiring decisions based on experience, skills, passion, and how you can enhance Canva and our culture. When you apply, please include your pronouns and any reasonable adjustments you may need during the interview process.

We celebrate all types of skills and backgrounds. Even if you don’t feel your skills fully match everything listed, we still want to hear from you.

Note: interviews are conducted virtually.

More open roles at Canva

Other open Research / Applied Science roles

Applying to this role

This PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling role at Canva runs through the firm's own careers portal and expects a CV and cover letter written specifically for the posting, not a portable submission carried across firms. Jorb AI's application agent tailors a CV and cover letter from your background to this posting and tracks the role alongside the rest of your applications.

Jorb AI tracks details for PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling at Canva. Postings refresh hourly from primary careers pages. Job details mirror the firm's posting; the apply link goes directly to the source. Last refreshed 2026-08-09.

Canva careers

Save this role and tailor your cover letter with Jorb AI.