* Own technical execution end-to-end: implementation, code review, and release.
* Translate research workflows and feature requests into well-scoped tasks with realistic, risk-aware estimates the team can plan against.
* Manage day-to-day execution: unblock people, sequence work, catch problems early.
* Set and defend the technical bar: review rigor, testing discipline, documentation, architectural consistency. Partner with the researchers — and amplify them
* Embed with malware researchers to understand their workflow and capture the tacit knowledge and edge cases no spec ever wrote down.
* Translate that knowledge into reliable agentic tooling — and know when an agent is confidently wrong before it ever reaches a researcher.
* Spend roughly 5–10% of your time doing actual malware research (with structured onboarding) to stay close to how the tool is used.
* Be willing to tell a researcher when a proposed workflow won't automate well — and explain why. Be the technical authority and mentor
* Make the hard architecture and design trade-off calls.
* Mentor through code review, pairing, and design discussions. Raise the level of everyone around you.
* Dive deep on the critical, difficult features and bug fixes yourself.
* Design agentic workflows into the architecture from the start, and build the evaluations and guardrails that keep them trustworthy.
About Alice:
Alice is a trust, safety, and security company built for the AI era. We safeguard the communicative technologies people use to create, collaborate, and interact—whether with each other or with machines. In a world where AI has fundamentally changed the nature of risk, Alice provides end-to-end coverage across the entire AI lifecycle. We support frontier model labs, enterprises, and UGC platforms with a comprehensive suite of solutions: from model hardening evaluations and pre-deployment red-teaming to runtime guardrails and ongoing drift detection.
Must-have 5+ years of software development experience, with a track record of delivering products to production — not just prototypes or POCs. • Strong Python, including async (asyncio), modern typing, and a disciplined testing approach (pytest). • Hands-on Playwright experience in production — not one-off scripts. • Production experience with agentic workflows: building, deploying, and operating LLM-powered systems that plan, call tools, and execute multi-step tasks — using a modern agent framework (e.g., LangGraph, the Anthropic Claude Agent SDK, the OpenAI Agents SDK, or DSPy). • Experience building evaluations and guardrails to measure agent quality and catch regressions before they reach a user (e.g., MLflow GenAI evaluation & tracing, LangSmith, or Braintrust). • Proven experience leading development efforts: estimation, task breakdown, code r













