AI Safety Researcher

Obiguard
Obiguard

Software Engineering, Data Science

United States · Kuala Lumpur, Malaysia · Remote

Posted on Sep 5, 2026
Department
RESEARCH
Location
Kuala Lumpur · Remote-friendly
Type
Full-time

Open to part-time arrangements, and we welcome internship applications for this role.

How to apply

Email your CV to hr@obiguard.com with the role title in the subject line.

Apply now →
About the role

Attacks against LLMs and autonomous agents evolve weekly — new jailbreaks, new injection techniques, new exfiltration patterns. As AI Safety Researcher, you’ll track this landscape ahead of our customers, reproduce attacks against real model deployments, and work with engineering to turn findings into detectors that ship into the inspection pipeline within days, not quarters. This is a forward-deployed research role — you may be deployed on-site with customers during pilots and security reviews, for stretches of up to three weeks at a time, advising on emerging risks in person.

What you'll do
  • Research and reproduce emerging attack techniques against LLMs and autonomous agents — prompt injection, jailbreaks, data exfiltration, tool-call abuse.
  • Design evaluation suites and red-team harnesses to continuously test Obiguard’s own detectors against novel attacks.
  • Translate research findings into production detector specs and policy primitives, working closely with the engineering team.
  • Publish internal and external write-ups on notable findings, contributing to Obiguard’s standing as a research-driven vendor.
  • Map new attack classes and detectors to relevant compliance frameworks (NIST AI RMF, ISO 42001, EU AI Act).
  • Advise customers on emerging risks during pilots and security reviews.
What we're looking for
  • Strong background in ML/NLP, applied security research, or adversarial machine learning — academic or industry.
  • Hands-on experience prompting, fine-tuning, or red-teaming LLMs.
  • Comfortable reading and reproducing findings from security research papers and disclosures.
  • Able to communicate technical findings clearly to both engineers and non-technical stakeholders.
  • Programming proficiency in Python; comfortable building quick evaluation tooling from scratch.
  • Comfortable with extended client-site deployments — this is a forward-deployed role, with on-site stints of up to three weeks at a time during active engagements.
Nice to have
  • Publications or public write-ups on LLM/agent security, adversarial ML, or red-teaming.
  • Experience with CTFs, bug bounties, or formal security research.
  • Familiarity with agent frameworks (LangGraph, AutoGen, MCP-based tooling).