We are looking for a proactive and skilled Software Infrastructure Developer with a strong Agentic AI focus to join our team. In this role, you will design and build the infrastructure, tooling, and AI-powered automation that keep our engineering systems reliable, scalable, and intelligent.
You will be responsible for embedding GenAI and agentic capabilities into our internal developer platforms and CI/CD workflows – creating frameworks, libraries, and multi-agent systems that accelerate automation, improve system stability, and unlock new levels of engineering productivity. This is a hands-on role for someone who thrives at the intersection of infrastructure engineering and applied AI.
The ideal candidate brings deep technical expertise in Python (FastAPI) and/or TypeScript, Kubernetes, CI/CD (Jenkins, GitHub Actions), containerization, and cloud infrastructure – combined with hands-on experience building LLM-powered applications, agents, MCP servers, and RAG pipelines.
What you'll be doing:
AI Platform Development: Design and build internal AI platforms, reusable SDKs, libraries, tools, and MCP servers for agentic applications.
Agentic Workflow Development: Develop single- and multi-agent workflows covering planning, tool selection, delegation, state and memory management, execution, retries, fallbacks, error recovery, and human-in-the-loop approval.
AI-Powered Infrastructure & CI/CD: Integrate AI capabilities into Kubernetes infrastructure and Jenkins/GitHub Actions pipelines to improve development, testing, deployment, and operational workflows.
AI Observability & Evaluation: Implement agent tracing, prompt and tool-call logging, latency and error monitoring, token and cost tracking, regression testing, and failure analysis.
AI Security & Governance: Apply guardrails, access controls, auditability, and governance practices to ensure agents operate safely and reliably.
End-to-End Platform Ownership: Own AI platform capabilities from architecture through production and build reusable solutions that accelerate adoption across engineering teams.
Cross-Functional Collaboration: Collaborate with engineering and business teams to identify high-impact AI opportunities and translate them into scalable technical solutions.
AI Research & Innovation: Evaluate emerging AI frameworks, models, and cloud technologies and promote their adoption where they provide measurable value.
You will be responsible for embedding GenAI and agentic capabilities into our internal developer platforms and CI/CD workflows – creating frameworks, libraries, and multi-agent systems that accelerate automation, improve system stability, and unlock new levels of engineering productivity. This is a hands-on role for someone who thrives at the intersection of infrastructure engineering and applied AI.
The ideal candidate brings deep technical expertise in Python (FastAPI) and/or TypeScript, Kubernetes, CI/CD (Jenkins, GitHub Actions), containerization, and cloud infrastructure – combined with hands-on experience building LLM-powered applications, agents, MCP servers, and RAG pipelines.
What you'll be doing:
AI Platform Development: Design and build internal AI platforms, reusable SDKs, libraries, tools, and MCP servers for agentic applications.
Agentic Workflow Development: Develop single- and multi-agent workflows covering planning, tool selection, delegation, state and memory management, execution, retries, fallbacks, error recovery, and human-in-the-loop approval.
AI-Powered Infrastructure & CI/CD: Integrate AI capabilities into Kubernetes infrastructure and Jenkins/GitHub Actions pipelines to improve development, testing, deployment, and operational workflows.
AI Observability & Evaluation: Implement agent tracing, prompt and tool-call logging, latency and error monitoring, token and cost tracking, regression testing, and failure analysis.
AI Security & Governance: Apply guardrails, access controls, auditability, and governance practices to ensure agents operate safely and reliably.
End-to-End Platform Ownership: Own AI platform capabilities from architecture through production and build reusable solutions that accelerate adoption across engineering teams.
Cross-Functional Collaboration: Collaborate with engineering and business teams to identify high-impact AI opportunities and translate them into scalable technical solutions.
AI Research & Innovation: Evaluate emerging AI frameworks, models, and cloud technologies and promote their adoption where they provide measurable value.
Requirements:
Core Engineering:
Minimum of 5 years of hands-on development experience with Python and/or TypeScript.
Experience building asynchronous backend services and APIs using FastAPI or an equivalent modern framework.
Strong software-engineering, analytical, troubleshooting, and cross-functional collaboration skills.
Ability to build reliable, reusable internal platforms that improve developer productivity and accelerate technology adoption.
Infrastructure & DevOps:
Strong experience with Kubernetes and Helm, including deployment, networking, resource management, autoscaling, troubleshooting, chart development, templating, versioning, and release management.
Strong experience with Docker, Jenkins, GitHub Actions, and AWS and/or Azure.
Familiarity with Infrastructure as Code, GitOps, and Git-based collaborative workflows.
AI & GenAI:
AI Experience – Mandatory: Minimum of 2 years of hands-on production experience building LLM-powered applications, AI agents, multi-agent systems, or internal AI platforms.
Hands-on experience with LangChain, LangGraph, Claude Agent SDK, and OpenAI and/or Anthropic APIs.
Strong understanding of agent architecture, including planning, tool selection, delegation, state and memory management, retries, fallbacks, error recovery, and human-in-the-loop workflows.
Experience with prompt and context engineering, tool/function calling, and structured outputs.
Experience building and integrating MCP servers and clients, including tool schemas, permissions, reliable execution, and error handling.
Core Engineering:
Minimum of 5 years of hands-on development experience with Python and/or TypeScript.
Experience building asynchronous backend services and APIs using FastAPI or an equivalent modern framework.
Strong software-engineering, analytical, troubleshooting, and cross-functional collaboration skills.
Ability to build reliable, reusable internal platforms that improve developer productivity and accelerate technology adoption.
Infrastructure & DevOps:
Strong experience with Kubernetes and Helm, including deployment, networking, resource management, autoscaling, troubleshooting, chart development, templating, versioning, and release management.
Strong experience with Docker, Jenkins, GitHub Actions, and AWS and/or Azure.
Familiarity with Infrastructure as Code, GitOps, and Git-based collaborative workflows.
AI & GenAI:
AI Experience – Mandatory: Minimum of 2 years of hands-on production experience building LLM-powered applications, AI agents, multi-agent systems, or internal AI platforms.
Hands-on experience with LangChain, LangGraph, Claude Agent SDK, and OpenAI and/or Anthropic APIs.
Strong understanding of agent architecture, including planning, tool selection, delegation, state and memory management, retries, fallbacks, error recovery, and human-in-the-loop workflows.
Experience with prompt and context engineering, tool/function calling, and structured outputs.
Experience building and integrating MCP servers and clients, including tool schemas, permissions, reliable execution, and error handling.
This position is open to all candidates.





