Anthropic Fellows Program, AI Safety & Security
Anthropic·London·United Kingdom·Research / Applied Science
Anthropic is hiring a Anthropic Fellows Program, AI Safety & Security in London. Posted 2026-08-11; applications close 2026-10-10 (in 52 days).
Role details
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team includes committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
Apply using this link: Apply using this link
We are accepting applications on a rolling basis for the next cohort expected to start in January 2027. In some circumstances, we can accommodate fellows starting outside the usual cohort timelines—please note in your application if the January 2027 start date does not work for you.
This page is specific to one of the Anthropic Fellows Workstreams. See also the main Anthropic Fellows posting.
Anthropic Fellows Program Overview
The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent regardless of previous experience.
Fellows primarily use external infrastructure (e.g., open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g., a paper submission). In one earlier cohort, over 80% of fellows produced papers. We run multiple cohorts each year and review applications on a rolling basis.
What to Expect
- 4 months of full-time research
- Direct mentorship from Anthropic researchers
- Access to a shared workspace (Berkeley, California, or London, UK)
- Connection to the broader AI safety and security research community
- Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (vary by country)
- Funding for compute (~$15k/month) and other research expenses
Interview Process
The interview process will include an initial application and reference check, technical assessments and interviews, and a research discussion.
We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every listed qualification. We also aim to include a range of diverse perspectives on our team.
Compensation
The expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension).
Fellows Workstreams
Due to the success of the Anthropic Fellows for AI Safety Research program, we are expanding it across teams at Anthropic. We expect significant overlap in the types of skills and responsibilities across roles and will by default consider candidates for all workstreams.
Some workstreams may include unique assessment steps. Please share your workstream preferences in the application. Current workstreams include:
- AI Safety Fellows
- AI Security Fellows
- ML Systems & Performance Fellows
- Reinforcement Learning Fellows
- Economics & Societal Impacts Fellows
Across the Workstreams: You May Be a Good Fit If You
- Are motivated by making sure AI is safe and beneficial for society as a whole
- Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
- Have a strong technical background in computer science, mathematics, or physics
- Thrive in fast-paced, collaborative environments
- Can implement ideas quickly and communicate clearly
Strong Candidates May Also Have
- A strong background in a discipline relevant to a specific Fellows workstream (e.g., economics, social sciences, or cybersecurity)
- Experience in areas of research or engineering related to their workstream
Candidates Must Be
- Fluent in Python programming
- Available to work full-time on the Fellows program
AI Safety Fellows
Mentors, Research Areas, & Past Projects
Fellows will undergo a project selection and mentor matching process. Potential mentors include:
- Sam Bowman
- Sara Price
- Alex Tamkin
- Nina Panickssery
- Trenton Bricken
- Logan Graham
- Jascha Sohl-Dickstein
- Joe Benton
- Collin Burns
- Fabien Roger
- Samuel Marks
- Kyle Fish
- Ethan Perez
Mentors will lead projects in AI safety research areas such as:
- Scalable Oversight: Developing techniques to keep highly capable models helpful and honest, even as they surpass human-level intelligence in various domains.
- Adversarial Robustness and AI Control: Creating methods to ensure advanced AI systems remain safe and harmless in unfamiliar or adversarial scenarios.
- Model Organisms: Creating model organisms of misalignment to improve empirical understanding of how alignment failures might arise.
- Model Internals / Mechanistic Interpretability: Advancing understanding of the internal workings of large language models to enable more targeted interventions and safety measures.
- AI Welfare: Improving understanding of potential AI welfare and developing related evaluations and mitigations.
On the Alignment Science and Frontier Red Team blogs, you can read about past projects, including:
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data — Alex Cloud and Minh Le, et al.; mentors include Samuel Marks and Owain Evans
- Open-source circuits — Michael Hanna and Mateusz Piotrowski; mentorship from Emmanuel Ameisen and Jack Lindsey
For a full list of representative projects for each area, please see:
- Introducing the Anthropic Fellows Program for AI Safety Research
- Recommendations for Technical AI Safety Research Directions
Unique Candidate Criteria
You might be a particularly great fit for this workstream if you:
- Are motivated by reducing catastrophic risks from advanced AI systems
- Have experience with empirical ML research projects
- Have experience working with large language models
- Have experience in one of the research areas mentioned above
- Have a track record of open-source contributions
Logistics
- Logistics Requirements: To participate, you must have work authorization in the US, UK, or Canada and be located in that country during the program.
- Workspace Locations: Shared workspaces are in London and Berkeley, where fellows will work from and mentors will visit. Remote fellows are also welcome in the UK, US, or Canada. The program will ask about your availability to work from Berkeley or London (full- or part-time) during the program.
- Visa Sponsorship: Anthropic is not currently able to sponsor visas for fellows. To participate, you need to have or independently obtain full-time work authorization in the UK, the US, or Canada.
- Program Duration: The program runs for 4 months full-time. If you cannot commit to the full duration, apply anyway and note your constraints in the application; requests are reviewed case-by-case.
- Offers: The program does not guarantee full-time offers. Strong performance may indicate fit for full-time roles at Anthropic. In previous cohorts, 25–50% of fellows received a full-time offer.
Applications and interviews are managed by Constellation, our recruiting partner. Clicking “Apply here” will take you to their portal, and updates will come from a Constellation address. Constellation also runs the Berkeley workspace and provides program support for fellows working on AI safety and security; fellows on capabilities-focused projects are supported directly by Anthropic.
Apply here: Apply here
More open roles at Anthropic
- Anthropic Fellows Program, The Anthropic Institute (Economics & Policy)
London · 2mo ago
- Anthropic Fellows Program
London · 4mo ago
Other open Research / Applied Science roles
- PhD Research Associate (Industry PhD Program) - Artificial Intelligence, SAP Labs Singapore
SAP · Singapore · 7mo ago
- PhD Research Scientist Intern - Reinforcement Learning for Diffusion Modelling
Canva · London · 12d ago
- 2027 Machine Learning Research Associate Program - PhD (New York)
Morgan Stanley · New York · 2mo ago
- PhD Student Intern - AI Research
SAP · Singapore · 4mo ago
- Data Science - Internship
DBS · Hong Kong · 16h ago
Applying to this role
This Anthropic Fellows Program, AI Safety & Security role at Anthropic runs through the firm's own careers portal and expects a CV and cover letter written specifically for the posting, not a portable submission carried across firms. Jorb AI's application agent tailors a CV and cover letter from your background to this posting and tracks the role alongside the rest of your applications.
Jorb AI tracks details for Anthropic Fellows Program, AI Safety & Security at Anthropic. Postings refresh hourly from primary careers pages. Job details mirror the firm's posting; the apply link goes directly to the source. Last refreshed 2026-08-19.
