Design and execute adversarial evaluations of LLM-powered web agents and chatbots to uncover failure modes (prompt injection, tool abuse, data exfiltration)
Analyze model behavior (including internal thoughts) to forecast security risks in frontier models
Develop and prototype security-focused training/fine-tuning techniques
Augment web agents with security principles (information flow control, privilege control, contextual security)
Partner with engineers and privacy/ML/cryptographic researchers to translate findings into product changes, tests, and measurable security improvements
Requirements
Demonstrated experience in Secure AI (prompt injection defenses and security guardrails)
Strong publication record in AI/ML security or related areas
Track record of shipping products/features in fast-moving environments
Strong Python skills and familiarity with ML frameworks (PyTorch or JAX)
Experience training, fine-tuning, and evaluating large language models (LLMs)
Compensation & benefits
Remote or hybrid work; London preferred but North America, EU, UK accepted
Competitive compensation and room for personal growth
Exciting, well-funded international startup with international exposure