
AI Infrastructure Engineer
Primary stack
Nice to have's
Job description
About the Position
We are seeking an AI Infrastructure Engineer to join our global team responsible for managing and supporting large-scale AI infrastructure environments. You will help ensure the availability, performance, and operational stability of critical AI infrastructure platforms, working closely with a distributed team across regions to provide continuous coverage and support.
This is an opportunity to work hands-on with some of the most advanced Kubernetes-based AI infrastructure in production today, while contributing to the platforms and processes that keep it running reliably at scale.
Responsibilities
- Manage and operate production AI infrastructure environments.
- Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations.
- Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.
- Conduct root cause analysis and drive product and operational improvements.
- Contribute and improve operational documentation and knowledge base.
- Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.
Requirements
- Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).
- Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.
- Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
- Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
- Effective verbal and written communication skills in English.
- Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.
- Willingness to occasionally work weekends and holidays.
Nice to Have
- Previous experience building, scaling, and running High-Performance Computing (HPC) environments.
- Hands-on experiences with managing large scale Kubernetes platforms in production.
- A good understanding of NVIDIA GPU technologies and the associated software stack.
- Proficiency in scripting languages (e.g., Python, Bash, Go).
We Offer
- Work with an established Silicon Valley leader in the cloud infrastructure industry.
- Work with exceptionally passionate, talented, and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies.
- Be a part of cutting-edge, open-source innovation.
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued.
- Professional development and training.
- Attend conferences and working groups.
- Company outings, happy hours, hackathons, and tech talks.
- Receive a competitive compensation package with a strong benefits plan.
About the Company
About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment-on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
© Mirantis. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
Our MissionOur mission is to be the world’s most effective vehicle for converting open source innovation into customer value.By providing continuous delivery of open source software via a managed services model, Mirantis enables organizations to easily consume infrastructure, rather than build it themselves. And, with an option to transfer operations control, enterprises can reap the benefits of open source standards and avoid vendor lock-in.Company OverviewMirantis delivers open cloud infrastructure to top enterprises using OpenStack, Kubernetes and related open source technologies. The company is a major contributor of code to many open infrastructure projects and follows a build-operate-transfer model to deliver Mirantis Cloud Platform and cloud management services, empowering customers to take advantage of open source innovation with no vendor lock-in. To date Mirantis has helped over 200 enterprises build and operate some of the largest open clouds in the world. Its customers include iconic brands such as AT&T, Comcast, Shenzhen Stock Exchange, eBay, Wells Fargo Bank and Volkswagen.Mirantis Fast FactsLeader in OpenStack and Kubernetes code contributions500+ employeesBased in Sunnyvale, CAPrivately heldTop-tier investors include August Capital, Dell Ventures, Ericsson, Goldman Sachs, Intel, Insight Venture Partners, Sapphire Ventures, Siguler Guff & Co., and WestSummit Capital.
More at Mirantis
All 41 roles
HPC Network Engineer
Mirantis

Intern — Data Analyst
Mirantis

Software Engineer, Infrastructure (Go)
Mirantis · United States

eCommerce Business Analyst & Project Coordinator
Mirantis · Uzbekistan
