waypointjobs

Trajectory Labs, PBC

Head of Evals, AI Red Teaming

berkeley, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Trajectory Labs, PBC lists this Head of Evals, AI Red Teaming opportunity in berkeley, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

You'll own the evaluation pipeline for our prompt injection red-teaming line: what we test, how we test it, and what ships to frontier lab customers.

This is a senior individual-contributor role. You'll mostly be directing agents rather than managing people, and you'll split your time between reviewing evals, designing new ones, and building the agent tooling that scales your own judgment.

We're an early-stage startup, so you should expect and enjoy that your responsibilities will grow and priorities change quickly.

About us

Our mission is to automate AI safety , to pave the way for a future where the vast majority of AI safety work is done by AI models.

Frontier models already solve coding problems that take humans days, but a model that can be hijacked by a malicious email or web page can't be trusted to work on its own. Before AI can do the work that matters, including AI safety research itself, models have to be robust to attack. So we build the safety and alignment evals, red-teaming programs, and RL environments that find these failures and train them out.

Frontier labs use our evaluations to make their models robust to prompt injection. That only works if the evals are right: a subtly broken task or a wrong grade teaches the model the wrong lesson. You own that bar.

You’ll:

Hold the quality bar. Review tasks, transcripts, and red-teamer submissions, and decide what ships to frontier lab customers.

Design what we test next. Study where models struggle and why, then design the task types, methodologies, and environments that target the gaps. The goal is training the failure out, not just finding it.

Automate your own judgment. Turn your review patterns into agent skills, checkers, and pipeline automation, so the quality bar scales faster than headcount.

Improve how we work. Our internal pipelines are agentic too. Apply the same eval eye to them and keep raising how much one person can do.

You’ll work closely with our founders and the teams developing frontier models, with unusual autonomy to make consequential decisions. The work you review shapes system cards, deployment safeguards, and how much the world can trust the most capable AI systems.

Your primary focus will be in our AI Red Teaming workstream. However, we build evals across several areas, and your responsibilities may expand over time.

What we’re looking for

Calibrated eval judgment : you can tell when a task, transcript, grade, or environment is subtly wrong, and explain the evidence behind your call.

Agent-native engineering : you use LLMs and coding agents as core tools, and you can independently script and automate your own workflows.

Sustained attention to detail : you hold the same bar on the hundredth review as on the first, and you look for ways to automate the repeatable parts.

Early-stage startup drive : you enjoy a fast pace, shifting priorities, limited structure, and taking on whatever will have the biggest impact.

You’ve evaluated something rigorously : an eval, a benchmark, a grading pipeline, or quality assurance you owned for a technical product. Agentic evals, RL environments, and model-training data are the strongest version

Experience building LLM judges and rubric-based grading, and iterating on them as models change

Experience designing and building evaluation environments yourself; professional software engineering experience is a strong plus

Prompt injection, red teaming, or security experience, especially when paired with eval design or grading judgment

Familiarity with AI safety and the alignment research community

Experience collaborating with frontier labs or other demanding technical customers

These criteria are a guide, not a checklist. If you want to do your life's work making frontier models safer and this role excites you, we encourage you to apply.

Location and workspace : Our team works out of Constellation in Berkeley, CA, and we prefer someone who can work alongside us there. We are open to remote for the right candidate.

Compensation : $200,000–$400,000 plus equity. More for exceptional candidates.

Health coverage : You'll receive a generous monthly pre-tax allowance to choose the medical, dental, and vision coverage that best fits your needs.

401(k) : We offer a 401(k) retirement plan.

Visa sponsorship : We sponsor visas, although we cannot successfully sponsor every visa for every role or candidate. If we make you an offer, we will make every reasonable effort to secure the visa you need, with support from an immigration lawyer we retain to guide and coordinate the process.

#J-18808-Ljbffr

Worksite address

berkeley, CA, 94709, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Jackson Therapy Partners

WhatJobs

Certified Occupational Therapy Assistant

gallup, NM

Salary not specified

Jackson Therapy Partners is seeking a Certified Occupational Therapy Assistant (COTA) to deliver patient-centered rehab services in diverse clini…

Last received from source 2026-10-10View job

Jackson Therapy Partners

WhatJobs

Certified Occupational Therapy Assistant

marlborough, MA

Salary not specified

Jackson Therapy Partners is seeking a Certified Occupational Therapy Assistant (COTA) to deliver patient-centered rehab services in diverse clini…

Last received from source 2026-10-10View job

Jackson HealthPros

WhatJobs

Computed Tomography (CT) Technologist

mcminnville, OR

Salary not specified

Jackson Health Professionals is seeking a dedicated Computed Tomography (CT) Technologist to join our traveling healthcare team. In this role, yo…

Last received from source 2026-10-10View job

Performance Food Group

WhatJobs

Diesel Mechanic

godfrey, IL

Salary not specified

Performance Food Group is seeking a skilled Diesel Mechanic to keep our distribution fleet safe, reliable, and road-ready. In this on-site role, …

Last received from source 2026-10-10View job

Rula Health

WhatJobs

Senior Digital Health Clinical Psychologist

clovis, CA

Salary not specified

At Rula Technologies, a pioneer in digital healthcare solutions, we are seeking a Licensed Clinical Psychologist to join our dynamic team. Our mi…

Last received from source 2026-10-10View job

Jackson HealthPros

WhatJobs

Nuclear Medicine Technologist

annandale, VA

Salary not specified

Jackson HealthPros is seeking a full-time Nuclear Medicine Tech to provide high-quality diagnostic imaging services for partner healthcare facili…

Last received from source 2026-10-10View job