
Senior Site Reliability Engineer
Primary stack
Nice to have's
Job description
About the Position
Senior Site Reliability Engineer
We are seeking a Senior Site Reliability Engineer to join our team in Arizona City, United States. This is a hybrid position, combining remote and on-site work.
About the Role
As a Senior Site Reliability Engineer, you will lead the operational excellence, reliability, and continuous improvement of AI-powered digital platforms supporting healthcare payer operations. You will oversee hybrid-cloud services leveraging Large Language Models (LLMs), MLOps practices, cloud-native technologies, and modern engineering frameworks to deliver secure, scalable, and compliant solutions that enhance member and provider experiences.
You will be a valued member of the Technology & Engineering team, collaborating closely with business stakeholders, product teams, platform engineers, AI specialists, and operations teams to ensure service stability, innovation, and regulatory compliance.
Responsibilities
- Own end-to-end service accountability for payer-focused AI, LLM, and ML-enabled platforms, ensuring high availability, performance, and compliance across hybrid environments.
- Lead operational governance and continuous improvement initiatives aligned with enterprise service management best practices.
- Oversee MLOps processes supporting LLM-powered applications, including model deployment, monitoring, retraining, validation, and rollback strategies.
- Coordinate the deployment and lifecycle management of containerized services utilizing Kubernetes to deliver scalable, resilient, and highly available solutions.
- Implement and govern Infrastructure as Code (IaC) practices using Terraform to provision and manage cloud and on-premises resources consistently and securely.
- Standardize configuration management through Ansible automation to improve operational efficiency and reduce service disruptions.
- Support the development, deployment, and operational management of Node.js-based services and APIs that integrate AI and machine learning capabilities.
- Establish best practices for Git-based version control, release management, code reviews, and repository governance across application and infrastructure teams.
- Drive operational excellence across Linux environments, including security hardening, patch management, performance optimization, and system monitoring.
- Guide Python-based development supporting data pipelines, AI orchestration, automation frameworks, and analytics workloads.
- Collaborate with business stakeholders, product owners, and healthcare domain experts to translate complex payer requirements into reliable technology services.
- Lead incident management, root cause analysis, problem management, and service restoration activities to minimize business impact.
- Monitor platform health through metrics, logs, traces, and observability tools while continuously improving service reliability and resilience.
- Foster a culture of knowledge sharing, operational excellence, automation, and continuous improvement across distributed teams.
Requirements
- Strong experience with Kubernetes and GCP (GKE)
- Strong experience in IaC (Terraform), Helm, and GitHub Actions
- Proficiency in Python, Ansible, Node.js
- Strong experience with Prometheus and Grafana observability stack
- Solid understanding of Linux systems and networking fundamentals
- Experience in incident management, on-call support, and production triage
- Hands-on experience with automation and CI/CD pipelines
- Strong understanding of AI/ML concepts and AIOps practices (model lifecycle, monitoring, or AI-driven alerting)
Nice to Have
- Google Cloud Architect Certification
- Certified Kubernetes Administrator (CKA)
- Experience in Java/J2EE, Spring Boot
- Experience supporting or operating ML/AI platforms or pipelines (MLOps)
- Exposure to AIOps tools, anomaly detection, or predictive analytics systems
- Experience with large-scale distributed systems and microservices architecture
- Experience with GPU-based workloads or ML infrastructure on GCP
- Knowledge of Kubeflow, Vertex AI, or ML pipelines
- Experience integrating AI-driven automation into monitoring and incident response
We Offer
Cognizant offers the following benefits for this position, subject to applicable eligibility requirements:
- Medical/Dental/Vision/Life Insurance
- Paid holidays plus Paid Time Off
- 401(k) plan and contributions
- Long-term/Short-term Disability
- Paid Parental Leave
- Employee Stock Purchase Plan
About the Company
At Cognizant, we're engineering modern businesses through innovation, technology, and human-centered solutions. You'll work with talented professionals, cutting-edge technologies, and industry-leading clients while helping shape the future of AI-driven healthcare solutions.
We're excited to meet people who share our mission and can make an impact in a variety of ways. Don't hesitate to apply, even if you don't meet every requirement. We value diverse experiences, transferable skills, and a passion for innovation.
Cognizant (Nasdaq: CTSH) is an AI Builder and technology services provider, bridging the gap between AI investment and enterprise value by building full-stack AI solutions for our clients. Our deep industry, process, and engineering expertise enables us to build an organization’s unique context into technology systems that amplify human potential, drive tangible outcomes, and keep global enterprises ahead in a fast-changing world. See how at cognizant.ai or follow us on Twitter.
Additional employment information
Compensation information is accurate as of the date of this posting. Cognizant reserves the right to modify this information at any time, subject to applicable law.
Applicants may be required to attend interviews in person or by video conference. In addition, candidates may be required to present their current state or government-issued ID during each interview.
Cognizant is an equal opportunity employer. Your application and candidacy will not be considered based on race, color, sex, religion, creed, sexual orientation, gender identity, national origin, disability, genetic information, pregnancy, veteran status, or any other characteristic protected by federal, state, or local laws.
If you have a disability that requires reasonable accommodation to search for a job opening or submit an application, please email [email protected] with your request and contact information.
© Cognizant Technology Solutions. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
Cognizant (Nasdaq-100: CTSH) is one of the world’s leading professional services companies, transforming clients’ business, operating and technology models for the digital era. Our unique industry-based, consultative approach helps clients envision, build and run more innovative and efficient businesses. Headquartered in the US, Cognizant is ranked 185 on the Fortune 500 and is consistently listed among the most admired companies in the world. Learn how Cognizant helps clients lead with digital atwww.cognizant.comor follow us @Cognizant.
More at Cognizant Technology Solutions
All 1104 roles
Oracle EPM Planning Cloud (EPBCS) Architect
Cognizant Technology Solutions · United States

Subject Matter Expert (SME) - JR
Cognizant Technology Solutions · Brazil

SailPoint Specialist
Cognizant Technology Solutions · India

SailPoint Specialist
Cognizant Technology Solutions · India