Lead Site Reliability Engineer

On-siteSalary not specified
United States

Job Description, Responsibilities & Requirements

About the Position

We are seeking a Lead Site Reliability Engineer with strong experience in observability, monitoring, and SRE practices to join our team in Wilmington, Delaware.

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
  • Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
  • Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
  • Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
  • Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
  • Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
  • Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
  • Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
  • Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
  • Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
  • Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
  • Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
  • Identify reliability, operational, and technology risks requiring escalation to management.
  • Promote an environment that supports a culture of belonging and reflects the Client brand.
  • Maintain Client internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
  • Complete other related duties as assigned.

Requirements

  • Must have

    • Strong experience in observability and monitoring, including hands-on expertise with:
      • Dynatrace
      • OpenTelemetry (OTel)
      • Distributed tracing
      • Metrics collection and analysis
      • Centralized logging and log aggregation
      • Alerting and dashboard development
    • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
    • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
    • Experience with CI/CD pipelines, deployment automation, and operational tooling.
    • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
    • Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
    • Cloud & Platform Expertise
      • Strong experience with Microsoft Azure, including:
        • Azure App Services
        • Resource Groups
        • Azure networking concepts
        • Scaling and performance optimization
        • Deployment and release management
        • Application lifecycle management
      • Experience leveraging Azure-native operational tooling such as:
        • Azure Monitor
        • Application Insights
        • Log Analytics
        • Azure dashboards and alerting
      • Experience supporting cloud-native and hybrid infrastructure environments.
    • Reliability & Engineering Practices
      • Demonstrated experience implementing and operating SRE practices, including:
        • Service Level Objectives (SLOs)
        • Service Level Indicators (SLIs)
        • Error budgets
        • Incident management
        • Problem management
        • Root Cause Analysis (RCA)
        • Reliability automation
      • Ability to improve system reliability through:
        • Performance tuning
        • Capacity planning
        • Observability-driven insights
        • Proactive issue detection
        • Reliability engineering initiatives
      • Experience developing automated recovery mechanisms and self-healing solutions.
      • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
  • Nice to have

    • Experience supporting large-scale enterprise applications in regulated environments.
    • Strong analytical and troubleshooting skills related to production systems and distributed architectures.
    • Experience working in Agile and DevOps operating models.
    • Ability to work autonomously and lead complex reliability initiatives.
    • Strong organizational and time management skills.
    • Advanced verbal and written communication skills.
    • Experience driving project milestones and delivery commitments.
    • Proven experience leading major incident response and post-incident improvement efforts.
    • Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.
    • Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
    • Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferred.

We Offer

  • Competitive salary
  • Opportunity to work in a dynamic and innovative environment
  • Professional growth and development opportunities
  • Collaborative and inclusive work culture

About the Company

Luxoft is a global leader in digital transformation and technology services, empowering businesses to thrive in the digital age. With a focus on innovation, quality, and customer satisfaction, Luxoft delivers cutting-edge solutions that drive success for our clients.

Application

Interested candidates are invited to apply for the Lead Site Reliability Engineer position by submitting their resume and a cover letter detailing their relevant experience and skills.

Job Details

Company name:
Luxoft
Location:
United States
Employment Type:
Full-time
Work Mode:
On-site
Posted on TheJob:
Aug 7, 2026
Last checked:
Aug 7, 2026
Apply Now