AgileEngine logo

Lead Data Engineer

AgileEngine·Salary not specified

Primary stack

AthenaDockerKubernetesPythonPostgreSQLAirflowGraphQLSQL

Nice to have's

TypeScript/ReactAWS BedrockLangChainPydantic AINxRedis/caching layersSageMakermarketing data structurescampaign management APIsdigital advertising metrics

Job description

About the Position

We are looking for a Lead Data Engineer to own the data pipeline and analytical architecture layer for a large-volume marketing analytics platform. You will make architectural decisions around partitioning strategy, file formats, schema design, and near-real-time processing for OLAP-oriented workloads built on an S3-backed data lake. You will design and govern ETL pipelines, define DAG-based orchestration strategies using Airflow, drive the AWS data stack including Athena and EKS, and lead a team of senior developers while enforcing code quality standards. The role requires a high degree of autonomy: you will often work on ad hoc or underspecified problems, defining the problem, gathering context, identifying constraints, and shaping the right technical approach before implementation.

Responsibilities

  • Design and own ETL pipelines that extract, transform, and validate data from internal databases and external APIs at scale.
  • Make architectural calls on partitioning, file formats, schema/data-type strategy, and near-real-time processing for large-volume, OLAP-oriented data systems built on an object-storage data lake.
  • Own the design of scheduled batch workflows (DAGs) on the client's Airflow setup - defining pipeline structure, dependencies, and triggering strategy, and driving architectural conversations about them. Not responsible for administering Airflow itself.
  • Drive use of the client's AWS data stack (S3-backed data lake, Athena, EKS/Kubernetes), and partner directly with the client's DevOps team to clarify functional and non-functional requirements.
  • Review PRs and enforce code quality standards.
  • Guide senior developers and ensure alignment with the client's engineering practices.

Requirements

  • 7+ years of engineering experience, with a proven track record designing and implementing ETL pipelines and making architectural decisions for large-volume data systems.
  • Hands-on experience with OLAP-style analytical data architecture - comparable experience with Athena, Trino/Presto, BigQuery, Snowflake, Spark SQL, ClickHouse, or similar is acceptable; a specific stack isn't mandatory as long as the OLAP depth is real.
  • Object-storage-backed data lakes: hands-on experience designing against a data lake sitting on object storage (S3 or equivalent) queried via a serverless engine - including partitioning strategy, file formats (Parquet/ORC), and the cost/performance tradeoffs that come with them. Athena specifically is a plus, not a requirement.
  • Task orchestration: Deep familiarity with DAG-style workflow definition and triggering. The client orchestrates most batch processing through Airflow, so this role needs either substantial prior Airflow experience they can draw on to drive architectural conversations, or enough depth in a comparable orchestrator (Dagster, Prefect, Luigi, Step Functions) to ramp on Airflow quickly and lead those conversations from day one. Managing the Airflow deployment itself is out of scope.
  • AWS Ecosystem: practical comfort across the client's AWS data stack - S3-backed data lake, serverless query engines (Athena or equivalent), and EKS/Kubernetes - with the ability to drive infrastructure conversations with DevOps.
  • Backend proficiency in Python (FastAPI or Flask).
  • Comfortable with REST and GraphQL.
  • Docker and PostgreSQL for the transactional/application layer.
  • Highly comfortable in Mac/Linux terminal-centric environments.
  • Practical, hands-on use of AI-assisted development tools (e.g., Claude Code), paired with the critical judgment to challenge AI output when it compromises long-term maintainability - including the leadership presence to set the standard for how the team uses AI tooling responsibly (e.g., flagging risky AI-driven shortcuts during PR review).
  • Strong soft skills: the ability to hold and defend a technical opinion - challenging a stakeholder's or a tool's proposed "quick fix" with sound reasoning in pursuit of a solution that scales and is maintainable long-term, while still being pragmatic enough to ship.
  • Comfort with ambiguity (mandatory): work is frequently ad hoc and underspecified. This role requires defining the problem - gathering context, identifying constraints, and framing the work - before solving it, rather than waiting for a specification. Experience limited to well-specified work executed through agent workflows is not a fit.
  • Upper-Intermediate English level.

Nice to Have

  • Direct production experience with Athena specifically.
  • Working knowledge of TypeScript/React - enough to guide integration and review frontend-adjacent PRs, even if not the primary focus.
  • Production AI features using AWS Bedrock, LangChain, Pydantic AI, or similar.
  • Monorepo tooling (Nx) or modern package managers (Poetry, UV, Yarn).
  • Redis/caching layers, SageMaker.
  • Experience with marketing data structures, campaign management APIs, or digital advertising metrics.

We Offer

  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support: access local well-being programs and people-focused support tailored to your location

About the Company

We are a remote-first company, so all our positions are remote. Schedules are flexible, but the main requirement is to have an overlap with your client’s schedule to ensure smooth collaboration. Specific details depend on the project and are discussed during the interview process.

© AgileEngine. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

AgileEngine is a privately held company established in 2010 and HQed in the Washington DC area. We rank among the fastest-growing US companies on the Inc 5000 list and the top-3 software developers in DC on Clutch.

More at AgileEngine

All 87 roles

Popular searches