If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale, then come join us! This truly is a role at the forefront of AI/ML-you'll be working on features for the largest clusters, with the largest customers, for the largest AI models.
Key job responsibilities
Be a senior engineer on a team that builds and maintains the infrastructure that monitors and reports on functionality and performance of massive testing workloads run at scale. Use our internal CI/CD tools, Linux, and public AWS products to automate the delivery of our software to customers, saving developer time. Write Python code that effortlessly spools up large clusters and runs benchmarks and applications for ML and HPC workloads. Use AWS Managed Grafana and Athena to digest the massive amount of performance data generated by these workloads and create dashboards for developers and stakeholders. Invent automatic mechanisms to alert developers to functional and performance regressions so they never reach reach customers. Manage the complexity of infrastructure that covers many instance types, software stacks, Linux operating systems, cutting-edge releases and make it easy to evolve.
Basic Qualifications
– 5+ years of non-internship professional software development experience.
– 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
– 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
– 3+ years as a mentor, tech lead or leading engineering teams.
– 3+years experience in SW/HW Co-Design.
Preferred Qualifications
– Bachelor's degree in computer science or equivalent.
– Experience creating automated dashboards and visualization (such as Grafana).





















