Tech Stack
Job Description, Responsibilities & Requirements
About the Position
Site Reliability Engineer
Remote - EMEA
Alpaca is a US-headquartered self-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more. Our recent Series D funding round brought our total investment to over $320 million, fueling our ambitious vision.
Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totaling over 9 million brokerage accounts.
Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.
Alpaca is proudly backed by top-tier global investors, including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Unbound, SBI Group, Derayah Financial, Elefund, and Y Combinator.
Responsibilities
As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform reliable, observable, and operable as we grow. Your tasks will include:
- Operate production day-to-day - on-call, incident response, postmortems, and follow-ups that actually close the loop.
- Own reliability practice - define and refine SLIs/SLOs and error budgets, and help product teams live within them.
- Strengthen our observability across metrics, logs, traces, and alerting.
- Ship infrastructure through code in a GitOps workflow - cloud resources and Kubernetes workloads alike.
- Look after PostgreSQL: performance tuning, schema and migration review, online migrations on large tables, HA/DR, and CDC pipelines.
- Mentor engineers on reliability and database fundamentals through code review, design review, and pairing.
Requirements
- 4+ years in SRE, DevOps, Platform/Infrastructure, or backend engineering with significant production operations ownership.
- Hands-on experience operating production services on Kubernetes, and shipping infrastructure as code in a GitOps workflow.
- Solid working knowledge of PostgreSQL in production - query plans, pg_stat_*, indexing and schema trade-offs, and what a safe online migration looks like on a non-trivial table.
- Cloud networking fundamentals (VPCs, routing, L4/L7 load balancing, DNS, TLS) and comfort debugging cross-service connectivity.
- Comfortable with a modern observability stack and proficient with Linux at the operator level.
- Practiced in incident response - calm under pressure, structured debugging, postmortems that drive change.
- At least working proficiency in Go or Python, plus strong written and verbal communication.
- Genuine interest in databases and in growing your PostgreSQL/DBA expertise.
Nice to Have
- Deeper PostgreSQL experience: large clusters at OLTP load, online migrations on big tables, HA/DR ownership, connection pooling at scale, or change-data-capture pipelines.
- Experience with typed SQL access layers in Go (e.g. pgx, gorm, sqlc).
- Production experience with messaging systems at scale (e.g. RabbitMQ, Kafka, Redpanda).
- Security & compliance experience in a regulated environment (SOC 2, secrets management, audit logging).
- Familiarity with trading, brokerage, or other regulated fintech domains.
We Offer
- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card
Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.
Recruitment Privacy Policy
Apply for this job
indicates a required field