AI Content Red Team Analyst - Trust and Safety

San Jose, California, US Until 8/24/2026 3+ years exp H-1B sponsor history First posted June 25, 2026 Last posted June 25, 2026
Job description

Description

The Trust & Safety (T&S) GenAI & Emerging Product team's mission is to empower the development of GenAI models and applications. We do this by building a world-class safety, testing, and risk management system that ensures GenAI innovations are launched responsibly.

The AI Content Red Team sits within the T&S GenAI and Emerging Products pillar. The team is responsible for conducting unstructured adversarial testing of Bytedance's generative AI products and models to uncover emerging risks, alongside our structured evaluations.

This team combines attacker-minded testing, risk discovery, and clear operational feedback loops to inform product decisions, policy development, mitigations, and longer-term evaluation strategy. We probe models and product experiences across modalities, use cases, and abuse patterns to identify failure modes, stress-test safeguards, and help teams improve safety before and after launch.

We work closely with Trust & Safety teams (policy, product, engineering, data science, operations), and business teams across global markets. Success in this team requires strong judgment, creativity, analytical rigor, and the ability to translate ambiguous findings into actionable recommendations.

Responsibilities:
- Conduct structured adversarial testing on AI models, features, and policies to identify vulnerabilities and emerging risks.
- Explore product behavior across contexts and user journeys, to identify model failure modes that may not be captured in standard evaluations.
- Investigate jailbreaks, evasions, prompt-based attacks, and other adversarial techniques relevant to content safety.
- Document findings clearly and consistently, including risk descriptions, reproduction steps, severity assessments, and mitigation recommendations.
- Partner with cross functional stakeholders (policy, product, business teams) to ensure mitigation validation and root cause closure.
- Support development of testing playbooks, taxonomies, and internal knowledge bases.
- Stay updated on emerging adversarial trends (e.g., deepfakes, multimodal manipulation, coordinated abuse), and shifts in the external risk landscape.

Requirements

Minimum Qualification(s):
- Minimum 3 years of experience in Trust & Safety, cybersecurity, risk/adversarial testing, or related fields.
- Experience with prompt testing, jailbreak analysis, LLM evaluation, or adversarial QA.
- Familiarity with AI safety risks (jailbreaks, hallucinations, bias, misuse patterns).
- Strong interest in GenAI safety, and the ways AI systems can be compromised under adversarial conditions.
- Demonstrated ability to independently investigate ambiguous problems, identify non-obvious failure modes and abuse patterns, and produce clear, evidence-based conclusions.
- Ability to manage multiple priorities, and collaborate effectively with cross-functional teams.

Preferred Qualification(s):
- Experience working with agentic AI tools to scale your impact, including building/operating AI tools to make processes efficient and effective.

About this role

Summary

Conduct adversarial testing on AI models to identify vulnerabilities and improve safety.

Job title

AI Content Red Team Analyst - Trust and Safety

Experience level

3+ years

Minimum experience

3+ years exp

Industry

technology

Location requirements

San Jose, California; remote work not specified.

Salary

Not specified

Visa sponsorship

H-1B sponsor history

Management role

No

Skills & keywords

Required skills

adversarial testingprompt testingjailbreak analysisLLM evaluationAI safety

Preferred skills

agentic AI toolsAI tool development

Specializations

adversarial testingLLM evaluationAI safetycontent safetyrisk management
Locations

Structured locations inferred from the posting.

San Jose, CA, USA

Work arrangement unknown City