Tech Stack
Job Description, Responsibilities & Requirements
About the Position
We are seeking a Senior Site Reliability Engineer to build and scale critical infrastructure behind every product at DraftKings. As a Senior Site Reliability Engineer, you'll build and scale the critical infrastructure behind every product. In this role, you'll take on complex challenges across global data centers, multiple cloud platforms, and on-premise systems-designing automation-first solutions that elevate performance and eliminate operational friction. You'll be trusted to drive stability at scale, influence architectural decisions, and build tools that empower our teams to move fast and deliver reliably. This is where your impact won't just be felt, it'll be foundational.
Responsibilities
- Drive stability and scalability across our global compute platform spanning numerous data centers, multiple public clouds, and on-premise environments, serving as the foundation for every product.
- Operate and evolve our GitOps delivery model, using Rancher Fleet and Flux with Helm to deploy core cluster services and application workloads declaratively and repeatably.
- Build self-healing, fault-tolerant infrastructure and internal tooling that eliminates repetitive operational work and reduces toil for both platform and application teams.
- Own cluster autoscaling and capacity strategy, including Karpenter, HPA, and KEDA, and predictive scaling driven by event and calendar data.
- Define SLOs and reliability metrics for platform components, using Datadog and our logging pipeline to surface cluster and workload health.
- Support technical growth by sharing knowledge, participating in design discussions, and contributing to a collaborative team culture, including on-call rotation.
Requirements
- Bachelor's degree in Computer Science or relevant education, experience, and training.
- At least 4 years managing distributed cloud and on-premise environments at scale, with strong hands-on AWS experience. Exposure to GCP, vSphere, or Nutanix is a plus.
- Deep expertise in container orchestration with Kubernetes, including the ability to design, scale, and troubleshoot complex workloads.
- Strong experience developing software for automation and infrastructure tooling such as Go and Python.
- Working knowledge of networking and Linux-based systems, including container runtimes such as Docker and containerd, packet-level debugging, and kernel troubleshooting.
- Experience with Infrastructure as Code (IaC) and configuration management tools to ensure scalable and repeatable infrastructure provisioning.
Nice to Have
- Experience with Rancher Fleet, Flux, Helm, Karpenter, HPA, KEDA, and Datadog.
We Offer
- Competitive salary range: $128,000 - $160,000 USD per year, plus bonus, equity, and benefits.
- Comprehensive health benefits, including medical, dental, and vision plans.
- Free access to programs such as free therapy sessions with Lyra Mental Health Solution, Lyra Employee Assistance Program, the Calm App, Virtual Yoga Classes, and more.
- 14 weeks of 100% paid parental leave for all global team members.
- Workplace lactation support and partnership with Care.com for backup childcare.
- Flexible PTO, pet insurance, gym reimbursement, financial planning support, commuter benefits, and tuition reimbursement.
About the Company
At DraftKings, AI is becoming an integral part of both our present and future, powering how work gets done today, guiding smarter decisions, and sparking bold ideas. It’s transforming how we enhance customer experiences, streamline operations, and unlock new possibilities. Our teams are energized by innovation and readily embrace emerging technology. We’re not waiting for the future to arrive. We’re shaping it, one bold step at a time. To those who see AI as a driver of progress, come build the future together.
Location: Boston, MA
Employment Type: Full-time, On-site
Categories: Engineering
Language: English (Native)
JR14235