AI Red Team & Safety Testing Jobs

AI companies pay people to attack their own models: probing for jailbreaks, hunting failure modes and edge cases, stress-testing safety guardrails before the public finds the gaps. The work rewards a security mindset and creative persistence in equal measure. It is done from home, typically paid hourly, and flexible enough to run as a side gig alongside other work.

Applied Clinical Judgement tracks red-team and AI safety testing roles daily across Mercor, micro1 and Turing. We are a referral service, not the employer. Each card shows pay, hours and eligibility, and links straight to the platform’s application page.

20 live AI Red Team & Safety Testing roles · updated daily

Mercor$48.0-$62.0 / hourly

AI Safety Experts — English & Danish

3d ago
Global · remote

$48.0–$62.0 hourly. Mercor seeks native English and Danish speakers to red team conversational AI models, probing for vulnerabilities through adversarial inputs, jailbreaks, and prompt injection techniques. You'll generate structured datasets documenting failure modes, classify risks, and produce reproducible reports for customers. The role suits those with prior red teaming, cybersecurity, or socio-technical testing experience who can work methodically within taxonomies whilst remaining creatively adversarial. Text-based work; sensitive content exposure is optional and supported.

red teamingadversarial testingprompt injectionjailbreak testingAI safety+4
View role & apply
Mercor$48.0-$62.0 / hourly

AI Safety Experts — English & Dutch

3d ago
Global · remote

$48.0–$62.0 per hour. Mercor is recruiting native English and Dutch speakers to red team conversational AI models and identify safety vulnerabilities. The role involves adversarial testing, prompt injection, jailbreak techniques, and bias exploitation across text-based projects. You'll generate structured datasets, document reproducible attack cases, and produce reports that help customers strengthen their AI systems. Prior red teaming, cybersecurity, or socio-technical risk experience is valued.

red teamingadversarial testingprompt injectionjailbreak techniquesbias detection+3
View role & apply
Mercor$48.0-$62.0 / hourly

AI Safety Experts — English & Finnish

3d ago
Global · remote

$48.0–$62.0 per hour. Mercor is hiring native English and Finnish speakers to red team conversational AI models and identify vulnerabilities through adversarial testing. You will probe for jailbreaks, prompt injections, bias exploitation, and misuse cases, then document findings and generate datasets for customers. Prior red teaming or cybersecurity experience is expected; the role suits those comfortable working on sensitive content with clear guidance and wellness support.

red teamingadversarial testingprompt injectionjailbreak testingAI safety+3
View role & apply
Mercor$17.0-$25.0 / hourly

AI Safety Experts — English & Indonesian

3d ago
Global · remote

$17.0–$25.0/hourly. Mercor seeks fluent English and Indonesian speakers to red team conversational AI models through adversarial testing, jailbreaks, and prompt injection attacks. You'll generate high-quality vulnerability data, classify failures, and document reproducible attack cases. Prior red teaming or cybersecurity experience is essential. This role suits those with structured adversarial thinking who can uncover AI safety vulnerabilities and communicate risks clearly.

red teamingadversarial testingprompt injectionjailbreak testingAI safety+4
View role & apply
Mercor$48.0-$62.0 / hourly

AI Safety Experts — English & Norwegian

3d ago
Global · remote

$48.0–$62.0 per hour. Mercor seeks red team experts fluent in English and Norwegian to probe conversational AI models for vulnerabilities. You'll conduct adversarial testing—jailbreaks, prompt injections, bias exploitation—generate high-quality safety data, and document reproducible attack cases. Prior red teaming, cybersecurity, or socio-technical risk experience is valued. Work is text-based, remote, and optional for higher-sensitivity projects with support provided.

red teamingadversarial testingprompt injectionjailbreak testingbias exploitation+3
View role & apply
Mercor$29.0-$45.0 / hourly

AI Safety Experts — English & Portuguese (global)

3d ago
Global · remote

$29.0–$45.0 hourly. Mercor seeks fluent English and Portuguese speakers (excluding Brazilian Portuguese) to red-team conversational AI models. You'll conduct adversarial testing, generate attack datasets, classify vulnerabilities, and document reproducible findings. Prior red-teaming, cybersecurity, or socio-technical risk experience is expected. The role suits security-minded practitioners comfortable probing AI systems for jailbreaks, prompt injections, bias exploitation, and misuse cases.

red teamingadversarial testingprompt injectionjailbreakAI safety+3
View role & apply
Mercor$48.0-$62.0 / hourly

AI Safety Experts — English & Swedish

3d ago
Global · remote

$48.0–$62.0 hourly. Mercor seeks fluent English and Swedish speakers to red team conversational AI models as part of their adversarial safety programme. You'll probe systems for jailbreaks, prompt injections, bias exploitation and misuse cases, then document vulnerabilities and generate datasets. The role suits those with prior red teaming, cybersecurity or socio-technical testing experience who work methodically within established taxonomies. All work is text-based; exposure to sensitive content is optional and supported.

red teamingadversarial testingprompt injectionjailbreakingAI safety+3
View role & apply
Mercor$24.0-$35.0 / hourly

AI Safety Experts — English & Thai

3d ago
Global · remote

$24.0–$35.0 per hour. Mercor seeks AI safety experts fluent in English and Thai to red team conversational AI models. The role involves adversarial testing, jailbreak attempts, prompt injection, and bias exploitation to surface vulnerabilities before deployment. You'll generate high-quality annotated datasets, classify failure modes, and document reproducible attack cases. Prior red teaming, cybersecurity, or socio-technical risk experience is essential. Work is text-based; exposure to sensitive content is optional and pre-disclosed.

red teamingadversarial testingprompt injectionjailbreak testingAI safety+4
View role & apply
Mercor$17.0-$25.0 / hourly

AI Safety Experts — English & Vietnamese

3d ago
Global · remote

$17.0–$25.0 per hour. Mercor seeks native English and Vietnamese speakers to red team conversational AI models through adversarial testing. You'll probe for jailbreaks, prompt injections, bias, and misuse cases; annotate failures; and generate reproducible attack datasets. Prior red teaming, cybersecurity, or socio-technical risk experience preferred. Text-based work with optional higher-sensitivity projects and wellness support.

red teamingadversarial testingprompt injectionjailbreakAI safety+3
View role & apply
MercorFrom 35 hrs/week$60.0-$90.0 / hourly

LLM Red Team Specialist — Failure Modes & Edge Cases

1 week agoMaster's
USA

$60.0–$90.0 per hour. This full-time W-2 role with Cincinnatus LLC places you within a leading AI lab's GenAI team to identify failure modes and edge cases in frontier models. You will design multi-step probing tasks, document vulnerabilities in coding and analysis domains, and collaborate with researchers to strengthen evaluation benchmarks. Requires MSc/PhD in STEM, 1+ years in research or AI evaluation, Python proficiency, and demonstrated red-teaming or adversarial testing experience. Based in the United States, remote, ~35 hours weekly.

PythonGitLLM evaluationred teamingadversarial testing+2
View role & apply
Micro1$60-$100 / hr

InfoSec Expert

1 week ago
Global · remote

$60–$100/hr. micro1 seeks information security specialists for red-teaming and adversarial testing of enterprise SaaS platforms. You will design attack scenarios, conduct penetration tests on Microsoft 365, Google Workspace, Slack, GitHub and Snowflake, and identify AI agent vulnerabilities including prompt injection and privilege escalation. The role suits experienced red teamers, penetration testers and cloud security professionals who can think adversarially about SaaS exploitation. Prior AI/LLM security knowledge is valued but not required.

red teamingpenetration testingcloud securityadversarial threat modelingMicrosoft 365+10
View role & apply
Mercor$60.0-$70.0 / hourly

AI Safety Practitioner

2 weeks agoBachelor's
Global · remote

$60.0–$70.0 per hour. Mercor seeks AI safety specialists to evaluate frontier model outputs for safety, accuracy and alignment across sensitive policy domains. You'll assess responses, apply safety rubrics, identify failures and hallucinations, and provide feedback to improve model behaviour. Requires a bachelor's degree in relevant disciplines and five years' experience in AI safety, trust and safety, journalism, policy, science or security roles.

AI SafetyRLHFSFTTrust & SafetyContent moderation+3
View role & apply
Mercor$70.0-$84.0 / hourly

AI Safety Red Teamer

2 weeks agoBachelor's
Global · remote

$70.0–$84.0 per hour. Mercor seeks experienced AI safety red teamers to probe frontier models for vulnerabilities through adversarial testing. You will design challenging prompts, uncover weaknesses in model behaviour, and evaluate robustness across sensitive domains including biosecurity, misinformation, and cyber threats. The role involves documenting findings and collaborating with AI researchers on safety improvements. Requires a relevant bachelor's degree and five years' prior experience in AI safety, red teaming, Trust & Safety, or related disciplines.

AI safety red teamingadversarial prompt designjailbreak testingprompt engineeringAI alignment+5
View role & apply
Micro1$40-$65 / hr

LLM Red-Teamer

2 weeks ago
Global · remote

$40–$65 / hr. micro1 seeks LLM Red-Teamers to develop adversarial multi-turn conversations and evaluation rubrics for frontier language models. You'll craft complex scenarios, assess model responses, and document failure modes. No AI background required, but deep familiarity with LLMs, exceptional written English, and experience in RLHF, annotation, or technical writing strengthen applications. Output-based compensation; minimum weekly submissions required.

Adversarial prompt constructionRubric designLLM evaluationRLHFWritten English+1
View role & apply
Micro1$50-$90 / hr

AI Jailbreak & Prompt-Injection Security Expert

1mo agoMaster's
Global · remote

$50–$90 / hr. micro1 seeks experts in AI security to design and execute jailbreak and prompt-injection testing methodologies. You'll develop adversarial evaluation frameworks, create regression test suites for vulnerability detection, and collaborate with technical teams to strengthen model robustness. The role suits those with 2+ years in adversarial ML, LLM red teaming, or AI safety, ideally holding a master's or PhD, and demonstrated community credibility through research, tools, or conference work.

Ethical JailbreaksLLM Red TeamingPrompt InjectionTool-Use AbuseAdversarial Machine Learning+4
View role & apply
Micro1$50-$90 / hr

Red Team Lead (Offensive Cybersecurity)

1mo agoPro. cert
Global · remote

$50–$90 / hr. micro1 seeks Red Team Leads to develop taxonomies, evaluation frameworks, and scoring rubrics for offensive cybersecurity training. You'll design proxy tasks simulating attack vectors, review red-team benchmarks, and produce technical documentation. The role suits experienced practitioners in exploit development, vulnerability research, cloud security, or malware analysis who can communicate complex concepts clearly. No AI background required.

Offensive cybersecurityRed teamingExploit developmentVulnerability researchExploit chains+6
View role & apply
Mercor$20.0-$22.0 / hourly

AI Safety Experts — English & Assamese

1mo ago
Global · remote

$20.0–$22.0 per hour. Mercor seeks English and Assamese-fluent specialists to red team conversational AI models, uncovering vulnerabilities through adversarial inputs, prompt injections, and bias exploitation. You'll generate annotated datasets, document attack cases, and classify systemic risks using structured taxonomies. Prior red teaming, cybersecurity, or adversarial ML experience is expected. Work is text-based; exposure to sensitive content is optional and supported by wellness resources.

red teamingadversarial testingprompt injectionjailbreak testingAI safety+3
View role & apply
Mercor$20.0-$22.0 / hourly

AI Safety Experts — English & Bengali

1mo ago
Global · remote

$20.0–$22.0 per hour. Mercor seeks native English and Bengali speakers to red team conversational AI systems. You'll conduct adversarial testing, identify vulnerabilities, generate attack datasets, and document findings using structured frameworks. The role suits candidates with red teaming, cybersecurity, or socio-technical expertise who can probe systems methodically and communicate risks clearly. All work is text-based; sensitive content exposure is optional and supported.

red teamingadversarial testingprompt injectionjailbreakingAI safety+3
View role & apply
Mercor$20.0-$22.0 / hourly

AI Safety Experts — English & Odia

1mo ago
Global · remote

$20.0–$22.0 per hour. Mercor seeks experienced red teamers fluent in English and Odia to probe conversational AI models for vulnerabilities. You will conduct adversarial testing, identify biases and misuse cases, generate annotated datasets, and document reproducible attack scenarios. The role suits those with prior red teaming, cybersecurity, or socio-technical risk experience who can work systematically within established frameworks and communicate findings clearly.

red teamingadversarial testingAI safetyprompt injectionjailbreak+3
View role & apply
Mercor$20.0-$22.0 / hourly

AI Safety Experts — English & Punjabi

1mo ago
Global · remote

$20.0–$22.0 / hourly. Mercor seeks fluent English and Punjabi speakers to red team conversational AI systems through adversarial testing. You will probe models for vulnerabilities including jailbreaks, prompt injections, and bias exploitation, then document findings in structured reports. The role suits those with prior experience in adversarial work, cybersecurity, or socio-technical testing who can identify weaknesses automated systems miss.

red teamingadversarial AIprompt injectionjailbreakingbias exploitation+3
View role & apply

Frequently asked questions

Do I need a cybersecurity background?

It helps, but plenty of listings weigh adversarial creativity just as heavily: the knack for finding the prompt that breaks the rule. Conventional pen-testing skills transfer well. Each listing states its own requirements.

Is this legitimate work?

Yes. This is sanctioned testing commissioned by the AI companies on their own systems, under agreement, so weaknesses get fixed before deployment.

What does the day-to-day involve?

Writing adversarial prompts, documenting model failures systematically, and often rating responses from a safety perspective. Clear written English and methodical documentation matter as much as ingenuity.

Last Reviewed: