S

Founding Member of Technical Staff – AI Safety Research (Evaluation, Methods & Infrastructure)

sampura research United Kingdom
Visa Sponsorship
Apply
AI Summary

Join Sampura Research as a Founding Member of Technical Staff to shape AI safety research focused on scalable oversight. You’ll design novel judge protocols, build high-quality evaluations, and develop infrastructure for hybrid human/AI judge systems. Own end-to-end research and engineering problems in evaluation, methods, and infrastructure with autonomy and impact.

Key Highlights
Founding role shaping AI safety research with autonomy and ownership over key projects
Focus on hybrid human/AI judge systems for scalable oversight of frontier models
Budget of $1M/quarter for compute and human data, with rapid hiring and growth
Key Responsibilities
Design and implement diverse, high-quality evaluations for judge protocols using existing datasets and novel data creation
Research and develop strategies to improve judge performance through hybridization (e.g., learning-to-defer, human assistance)
Build scalable infrastructure for large-scale inference, evaluation pipelines, and human rating platforms
Synthesize findings into papers, blog posts, and benchmark releases for public dissemination
Lead research and mentorship programs through paid fellowships or volunteer initiatives
Technical Skills Required
Python Machine Learning Large Language Models (LLMs)
Benefits & Perks
Visa sponsorship
Competitive salary (£100,000–£290,000 annually)
On-site location in London
Nice to Have
Experience with reinforcement learning, adversarial fine-tuning, or distillation applied to judges
Leadership in mentoring or team leadership for related research programs
Expertise in model confidence/calibration, interpretability, or human-computer interaction

Job Description


About the Org

Sampura Research is an AI safety research non-profit focused on scalable oversight for frontier models through improved human/AI “judges”.


Judges (LLM, human, or a hybrid) are essential for aligning highly capable models through their use in evaluation and training. Limitations in today's judges can lead to undesirable behavior such as reward hacking in training, or failure to spot misalignment in evaluation. As models become more capable and their data more complex, often beyond what the systems overseeing them can understand, we expect these limitations to grow in severity and to become harder to notice.


We are designing novel judge protocols which leverage the strengths of both humans and models to outperform either alone, and building high-quality evaluations to explore judge performance across a diverse set of safety-related tasks. We believe that our focused approach will rapidly advance the science of judges by providing a much-needed systematic measurement of which protocols actually work, and producing state-of-the-art judges which can be used in real world training and evaluation pipelines.


Backed by an $11M grant from Coefficient Giving ($7M for the first year, and another $4M pledged), Sampura Research was founded in 2026 by members of DeepMind’s alignment and human data teams.


Please see our announcement (https://sampura.org/news/announcing-sampura-research/) for more information.


Role Overview

As a founding Member of Technical Staff, you will help shape the research and engineering direction alongside the founding team, and own end-to-end research and engineering problems within our agenda:


Evaluation: Build diverse and high-quality evaluations for judge protocols by pulling from existing datasets/literature and creating novel datasets from scratch.

Methods: Research and implement strategies for improving judge performance through hybridization, such as learning-to-defer, human assistance, or other mechanisms.

Infrastructure: Build the tools required to enable and scale our research agenda, including large-scale inference/evaluation pipelines, a human rating platform, and a robust public leaderboard.


We've budgeted roughly $1M per quarter for compute and human data, across a team of fewer than ten, to ensure that we can commit to ambitious goals in all of these areas. You'll scope ambiguous problems independently, communicate progress clearly, and synthesize findings into papers, blog posts, and benchmark releases.


Staff hires may also lead research and mentorship programs through paid fellowships or volunteer research focused on evaluation or methods work.


If our work and mission resonate with you, please consider applying even if you don’t tick every box. We strongly believe in investing in excellent people regardless of prior or formal experience.


You may be a good fit if you have

  • A Bachelor's degree or higher in Computer Science, Mathematics, or a related field or equivalent experience.
  • Proficiency coding in Python or similar languages, as well as machine learning and analysis tooling (e.g., JAX, PyTorch, Pandas).
  • Experience working with large language models, human-in-the-loop data annotation systems, or AI evaluation frameworks.
  • A willingness to embrace AI tools intelligently, applying scrutiny and human review where necessary.
  • A track record of exploring and resolving open-ended research questions with rigor, in academic or industry labs, fellowships, or volunteer research programs.


Outstanding candidates will have

  • Demonstrated ownership across multiple parts of the AI research lifecycle (experiment design and execution, infrastructure, data collection and processing, inference).
  • Expertise in a highly relevant domain such as model confidence/calibration, interpretability, safety/alignment evaluation, human-computer interaction, or other fields related to scalable oversight.
  • Hands-on experience with post-training methods such as reinforcement learning, adversarial fine-tuning, or distillation, especially applied to judges or reward models.
  • Leadership as a mentor or team lead for related research, either inside a research lab or via programs like SPAR, MARS, or MATS.


Salary and Logistics

Junior hires: £100,000-£150,000 (~$136,000-$205,000)

Senior hires: £150,000-£200,000 (~$205,000-$273,000)

Staff hires: £200,000-£290,000 (~$273,000-$395,000)

*Approximate USD conversion as of 24 August 2026.


Exceptional students/new-grads and early-career individuals will often come in at the junior level. Researchers or engineers with several relevant years of experience will often come in at the senior level. Staff hires typically bring deep expertise in a directly relevant area at notable research labs, and have led research programs or teams. We'll agree a target level with you early in the process, and may reassess as you progress.


Location: This role is based on-site in London.

Visa Sponsorship: We offer visa sponsorship for this role.

Start Dates: We are hiring ~2 FTEs targeting a late-September 2026 start, and ~2 more starting January 2027. We will continue hiring more FTEs later in 2027. We will review applications on a rolling basis, so please apply ASAP.


Equal Opportunity

Sampura Research is an equal opportunity employer. We are committed to building a team with diverse backgrounds and perspectives and believe this makes our research and workplace better. We do not discriminate on the basis of race, religion or belief, sex, sexual orientation, gender reassignment, age, disability, marriage or civil partnership, pregnancy or maternity, or any other protected characteristic. If you need adjustments at any stage of the application process, let us know and we'll make them.


Similar Jobs

Explore other opportunities that match your interests

Machine Learning Runtime Development Engineer

Machine Learning
19h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Arm

United Kingdom

System Test Engineer - AI Hardware Accelerator Verification

Machine Learning
19h ago

Premium Job

Sign up is free! Login or Sign up to view full details.

•••••• •••••• ••••••
Job Type ••••••
Experience Level ••••••

Arm

United Kingdom

Lead Machine Learning Scientist

Machine Learning
1w ago
Visa Sponsorship Relocation Remote
Job Type Full-time
Experience Level Not Applicable

monzo

United Kingdom

Subscribe our newsletter

New Things Will Always Update Regularly