What You'll Do
Design, implement, and manage our cloud infrastructure using tools like Terraform, ensuring environment consistency and scalability.
CI/CD Automation: Take full ownership of our deployment pipelines, optimizing for speed, reliability, and developer productivity.
Cloud Orchestration: Manage and scale our Kubernetes (K8s) clusters on AWS, ensuring high availability and efficient resource utilization.
Observability & Monitoring: Implement and maintain robust monitoring, logging, and alerting systems to ensure platform health and rapid incident response.
Security & Compliance: Drive security best practices across the infrastructure, including IAM management, network security, and vulnerability scanning.
Collaborate: Work closely with software engineers to optimize application performance, containerization strategies, and database reliability.
We are seeking a hands-on builder who can grow into a technical leader for our infrastructure. Someone excited to own hard problems and pick up the specifics of our stack quickly.
Must-Haves
5+ years of experience in DevOps or Site Reliability Engineering (SRE), ideally within a high-growth SaaS or data-heavy environment.
Hands-on experience running production workloads on AWS (e.g., EKS, RDS, S3, IAM, VPC).
Strong experience with Kubernetes and Docker in production environments.
Proficiency with CI/CD tooling (e.g., GitHub Actions, GitLab CI, or Jenkins) and infrastructure-as-code (e.g., Terraform).
Working proficiency in Python, Bash, or Go for automation and internal tooling.
A team player with strong communication skills who can explain infrastructure concepts to cross-functional partners.








