EPAM logo

Senior SRE Engineer

EPAM·Salary not specified

Primary stack

KubernetesCloud LoggingAWSCloud Monitoring

Nice to have's

multi-cloudhybrid environmentsInfrastructure as CodeTerraformCloudFormationlarge-scale environmentsregulated environmentsvector databasesRAG architectureGenerative AILLM platformsClaudeAmazon Bedrock

Job description

About the Position

We are seeking a Senior SRE Engineer with strong technical authority to influence design and operational decisions across architecture, engineering, security, and operations teams. The ideal candidate is a pragmatic problem-solver who remains calm and methodical under pressure, balancing reliability, security, cost, and delivery speed while clearly communicating complex technical concepts to diverse audiences. This role requires onsite presence at the client office three times a week.

Responsibilities

  • Design and operate observability platforms, including monitoring, logging, and alerting systems
  • Manage metrics, logs, APM, and alerting through Datadog
  • Apply SRE principles such as SLOs, error budgets, incident management, and reliability engineering
  • Collaborate with security teams to uphold cloud security principles
  • Apply cloud security best practices to ensure system reliability
  • Manage security incidents and drive mitigation and remediation measures
  • Monitor, optimize, and control costs associated with ML models and API-based AI services
  • Operate and maintain Kubernetes and containerized platforms
  • Manage AWS infrastructure, including EKS, ECS, EC2, networking, IAM, and managed services

Requirements

  • 6-9 years of hands-on technical experience in SRE, Platform Engineering, Infrastructure, or related roles
  • Strong experience with AWS services such as EKS, ECS, EC2, networking, IAM, and managed services
  • Deep hands-on experience with Kubernetes and containerized platforms
  • Proven experience designing and operating observability platforms, including monitoring, logging, and alerting
  • Hands-on experience with Datadog for metrics, logs, APM, and alerting
  • Strong understanding of SRE principles, including SLOs, error budgets, incident management, and reliability engineering
  • Solid understanding of cloud security principles and experience collaborating with security teams
  • Extensive hands-on experience in cloud security engineering, applying best practices to ensure system reliability, managing security incidents, and driving mitigation and remediation measures
  • Experience or working knowledge of FinOps practices, including monitoring, optimizing, and controlling costs associated with ML models and API-based AI services

Nice to Have

  • Experience supporting multi-cloud or hybrid environments
  • Exposure to Infrastructure as Code, such as Terraform and CloudFormation
  • Experience in large-scale, complex, or regulated environments
  • Knowledge of vector databases and RAG architecture for building internal SRE knowledge assistants
  • Knowledge of Generative AI and LLM platforms, such as Claude and Amazon Bedrock

About the Company

[Company description if present]

© EPAM. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

EPAM helps organizations innovate their business processes and rethink the way they manage their businesses so they can remain competitive in this new digital age.

More at EPAM

All 2376 roles

Similar jobs

Popular searches