
Senior Site Reliability Engineer
Primary stack
Job description
About the Position
We are seeking a skilled and results-driven Senior Site Reliability Engineer to join our engineering team and bridge the gap between software development and systems operations. In this role, you will apply engineering principles to automate operations, scale infrastructure, and keep systems highly available, resilient, and performant. Your mission is to build, run, and safeguard the production environments that power our applications, reducing downtime and enabling fast, safe software delivery.
Responsibilities
- Architect, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation
- Develop and refine CI/CD pipelines to streamline software deployments, configuration management, and routine operational tasks
- Implement comprehensive logging, monitoring, and alerting systems using tools like Prometheus, Grafana, and Datadog
- Define clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Lead incident response efforts and troubleshoot production issues to restore services promptly
- Facilitate blameless post-mortems to uncover root causes and prevent future occurrences
- Collaborate with software developers to enhance system performance and plan for capacity needs
- Guarantee services scale effectively to accommodate growth and traffic surges
Requirements
- 3+ years of experience in systems administration, DevOps, or systems-oriented software development
- Competency in at least one scripting or programming language such as Python, Bash, Go, or Rust
- Hands-on experience with public cloud platforms such as AWS, Azure, or GCP, along with containerization tools like Docker and Kubernetes
- Solid understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS
- Demonstrated reliability mindset with a strong drive for automation, reducing toil, and designing resilient systems that fail gracefully
- English proficiency at B2 level or higher
We Offer
- Opportunity to work in a Hybrid role based in Argentina with other locations
- Competitive compensation package
- Opportunity to work with cutting-edge technologies and practices
About the Company
[Company description if present]
© EPAM. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
EPAM helps organizations innovate their business processes and rethink the way they manage their businesses so they can remain competitive in this new digital age.
More at EPAM
All 2384 roles
Lead Full-Stack Developer
EPAM · Argentina +1

Python Engineering Manager
EPAM · Mexico

Product Manager
EPAM · Argentina +2

Automation Tester (JavaScript)
EPAM · Brazil
Similar jobs
Senior DevOps / Site Reliability Engineer
N-iX
Senior DevOps / Site Reliability Engineer
N-iX
Senior DevOps / Site Reliability Engineer
N-iX
Senior DevOps / Site Reliability Engineer
N-iX