SafetyTech Client #1 | Adversarial Task Writer for AI Security RL Gyms

SD Solutions

Apply to this job
Serbia / Armenia / Bulgaria Until 9/21/2026 First posted July 23, 2026 Last posted July 23, 2026
Job description

On behalf of SafetyTech Client #1, SD Solutions is looking for a talented Adversarial Task Writer for AI Security RL Gyms.

SD Solutions is a staffing company operating globally. Contact us to get more details about the benefits we offer.

Responsibilities:

You design prompt injection scenarios in YAML, run them against frontier models, validate success rates, and submit passing tasks. 5 high-quality tasks per week (full-time equivalent). Per-task compensation, paid on acceptance.

    Requirements:

    • Adversarial mindset: you think like an attacker and understand how to exploit an AI agent’s helpfulness, authority assumptions, or trust in its environment
    • Prompt injection expertise: direct (role-play, encoding, context flooding) and indirect/environment-embedded (poisoned tool responses, malicious content in documents, cross-context leakage)
    • Technical writing in YAML
    • Comfortable with Docker, CLI tools, and running systematic tests against multiple models
    • Domain realism in at least one vertical: e-commerce, finance, HR, enterprise SaaS, healthcare, travel
    • Background in pentesting, appsec, LLM security research, or red teaming strongly preferred

    The Task

    You build adversarial prompt injection tasks for Alice’s RL Gym platform. Each task is a self-contained YAML scenario simulating a realistic AI agent deployment, testing whether the agent can be manipulated into violating its safety policies.

    About the company:

    A company building specialized evaluation infrastructure for AI safety and robustness testing. Their platform simulates adversarial conditions used by AI development teams to validate agent behavior before deployment. Currently expanding a freelance contributor pool for scenario and environment development.

    By applying for this position, you agree to the terms outlined in our Privacy Policy. Please take a moment to review our Privacy Policy https://sd-solutions.breezy.hr/privacy-notice, and make sure you understand its contents. If you have any questions or concerns regarding our Privacy Policy, please feel free to contact us.

    About this role

    Summary

    Design prompt injection scenarios for AI safety testing using YAML and security expertise

    Job title

    Adversarial Task Writer for AI Security RL Gyms

    Experience level

    preferably with pentesting, appsec, or red teaming background

    Industry

    software

    Location requirements

    Remote work allowed from Serbia, Armenia, Bulgaria

    Salary

    Not specified

    Management role

    No

    Skills & keywords

    Required skills

    yamldockercli toolsadversarial mindsetprompt injection

    Preferred skills

    pentestingappsecred teamingllm security research

    Specializations

    adversarial mindsetprompt injectionyamldockersecurity
    Locations

    Structured locations inferred from the posting.

    Serbia

    Remote Country

    Armenia

    Remote Country

    Bulgaria

    Remote Country