Senior DevOps Engineer

On-siteSalary not specified

Tech Stack

PrometheusDevOpsLokiHelmKubernetesPostgreSQLGrafanaCI/CDTempoTerraform

Job Description, Responsibilities & Requirements

About the Position

Senior DevOps Engineer

Remote - Japan - APAC


About the Company

Alpaca is a US-headquartered self-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. Our recent Series D funding round brought our total investment to over $320 million, fueling our ambitious vision.

Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totaling over 9 million brokerage accounts.

Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.

Alpaca is proudly backed by top-tier global investors, including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Unbound, SBI Group, Derayah Financial, Elefund, and Y Combinator.


Responsibilities

As a Senior DevOps Engineer, you will design, build, and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You will have the autonomy to design and implement solutions against clearly defined goals and a real voice in shaping those goals with the team.

  • Design and evolve our cloud architecture on GCP - networking, interconnects, IAM, and high-availability topology - and express it entirely as code with Terraform, following GitOps as a first principle.
  • Build and own the CI/CD pipelines that plan, review, test, and safely apply IaC changes - Policy-as-Code guardrails, drift detection, and progressive rollout so infrastructure changes ship as confidently as application code.
  • Advance Platform-as-a-Product: build self-serve capabilities and paved paths so engineers can provision what they need, through a golden path rather than a hand-off.
  • Strengthen our observability stack - metrics, logs, traces, and alerting across Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager - so the platform is easy to run and reason about.
  • Operate our GKE clusters and the infrastructure services that run on them - Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ), and data stores.
  • Participate in our Follow-The-Sun on-call model: watch and triage alerts, join and declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and the post-actions that actually close the loop.
  • Embed SRE practices - SLIs/SLOs and error budgets, capacity planning - into how Core Infrastructure builds and operates, working closely with our SRE function.

Requirements

  • 5+ years in a DevOps, Platform/Infrastructure, or SRE role, with a proven track record operating large-scale, high-availability, high-performance systems in production.
  • Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud - landing zones, networking, IAM, and high-availability topology.
  • Strong Infrastructure-as-Code skills with Terraform, structuring large codebases across multiple environments, with GitOps as a first principle and least-privilege as a default mindset.
  • Proven experience building CI/CD pipelines for IaC - automated plan/apply, code review, Policy-as-Code, drift detection, and safe rollout.
  • Significant production experience with Kubernetes (ideally GKE) and packaging/deploying workloads with Helm.
  • Solid cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and comfort debugging cross-service connectivity.
  • Hands-on experience with a modern observability stack - Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager - across metrics, logs, traces, and alerting.
  • Operator-level familiarity with data stores such as PostgreSQL and Message Brokers (e.g., RabbitMQ, RedPanda) - able to run and troubleshoot them in production.
  • A good understanding of SRE practices - SLOs/error budgets, capacity planning - and a Platform-as-a-Product mindset.
  • Strong grasp of incident management end to end: joining and declaring incidents, structured debugging under pressure, escalation, clear documentation, and post-mortems that drive real change.
  • Able and willing to take part in a Follow-The-Sun on-call rotation from APAC hours, and to work effectively in a distributed, async-first team with strong written communication.

Nice to Have

  • Policy-as-code and IaC quality tooling (OPA/Conftest, Checkov, tflint, Atlantis, or similar).
  • Experience managing Terraform state, module registries, and versioning at scale across many teams.
  • Experience building self-serve developer platforms and internal golden paths (e.g., with Backstage, Tilt, or similar).
  • Experience with the Alloy collector and with incident tooling such as Rootly.
  • Working proficiency in Go for automation and tooling.
  • Strong Linux (Debian/Ubuntu) and container (Docker/containerd) fundamentals.
  • Security and compliance experience in a regulated environment (SOC 2, secrets management, audit logging).
  • Familiarity with trading, brokerage, or other regulated fintech domains, and with low-latency systems.

We Offer

  • Competitive Salary & Stock Options
  • Health Benefits
  • New Hire Home-Office Setup: One-time USD $500
  • Monthly Stipend: USD $150 per month via a Brex Card

*Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.

*

Recruitment Privacy Policy


Apply for this job

indicates a required field

Job Details

Company name:
Alpaca Expo Group
Location:
Japan
Work Mode:
On-site
Posted on TheJob:
Jul 18, 2026
Last checked:
Jul 18, 2026
Apply Now
© 2026 TheJob, Inc. All rights reserved.