
Lead Data Quality Engineer
Primary stack
Nice to have's
Job description
About the Position
We are seeking a Lead Data Quality Engineer to drive rigorous data validation, SQL-based testing, and quality automation across cloud data environments. You will verify complex transformations, ensure consistency across multiple systems, and support migration and deployment work within modern data ecosystems. Join a distributed team focused on trustworthy datasets and apply today.
Responsibilities
- Execute QA validation for bulk data products within the data exchange ecosystem
- Validate match and append processes and ensure correct deployment into data pipelines
- Provide QA support for data platform migrations, ensuring data integrity and functional correctness
- Verify data products and data fulfillment processes for accuracy and completeness
- Perform data validation and comparisons across systems, including source input files to cloud data warehouse tables, table-to-table checks, and confirmation of data mappings and transformations against specifications
- Use and enhance quality check frameworks built with PySpark scripts to automate data validation
- Investigate defects through data analysis, identifying root causes in data pipelines or transformation logic
- Collaborate with engineering and data teams to triage issues, validate fixes, and confirm production readiness
- Contribute to test automation for data validation and testing to increase efficiency and coverage
- Communicate findings, risks, and test results clearly to stakeholders
Requirements
- 5+ years of experience in Data Quality Engineering
- Expertise in SQL, including complex joins across multiple tables and large datasets
- Hands-on experience with cloud data warehouses such as BigQuery, Redshift, or Synapse Analytics
- Working knowledge of cloud platform services across AWS, Azure, or GCP
- Understanding of data validation practices and automated data comparison techniques
- Background in data transformation, validation, and mapping verification using specifications
- Capability to understand and work with PySpark-based quality frameworks
- Strong data analysis and debugging skills, with an ability to identify defects in data processing pipelines
- Excellent written and verbal communication skills
- Upper-Intermediate English proficiency (B2)
Nice to Have
- Proficiency in Python or PySpark development
- Background in data engineering or data pipeline testing environments
Locations
- Remote in Mexico
- Remote in 4 other locations
About the Company
[Company description if present]
© EPAM. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
EPAM helps organizations innovate their business processes and rethink the way they manage their businesses so they can remain competitive in this new digital age.