Preply logo

Data Engineer

Preply·Salary not specified

Primary stack

DevOpsProblem Solving

Job description

About the Position

Data Engineer

At Preply, we are dedicated to creating life-changing learning experiences. We help people discover the magic of the perfect tutor, craft a personalized learning journey, and stay motivated to keep growing. Our approach is human-led, tech-enabled - and it’s creating real impact.

We’ve just reached unicorn status with a $150M Series D, accelerating our vision to transform education through human-led, AI-enhanced learning. Today, 100,000+ tutors teach 90+ languages to learners in 180 countries - and we’re only getting started. As a category-defining company, we’re shaping what the future of learning looks like at global scale.

Every Preply lesson sparks change, fuels ambition, and drives progress that matters. Joining Preply means helping define the future of education at global scale, and building something that truly matters for millions of people, every day.

Responsibilities

As a Data Engineer in the Data Ingestion and Enrichment team, you will build and contribute to the data layer that powers both Preply's analytics, machine learning, and product. You will work closely with ML Platform, Applied/Data Scientists, Analytics Engineering, and Product squads to ensure that features, datasets, and pipelines are production-ready, observable, and reusable within the team.

Key Responsibilities:

  • Contribute to trusted ingestion & enrichment foundations (Data Lake and Data as a Product):

    • Build and maintain components of Preply's data lake.
    • Ensure every dataset has clear ownership, purpose, schemas, and quality expectations from first ingestion through downstream consumption by analytics, product, and ML teams.
  • Develop end-to-end ingestion pipelines (batch & streaming):

    • Build and operate reliable batch and streaming ingestion pipelines that support both real-time and analytical use cases.
    • Contribute to defining clear raw → standardized → consumption layers with explicit responsibilities, lineage, and retention strategies.
  • Data quality, contracts & early validation:

    • Implement data contracts between producers and consumers, covering schema, freshness, volume, and quality guarantees.
    • Embed validation, anomaly detection, and quality checks early in the ingestion lifecycle to catch issues before they propagate.
  • Enrichment, modeling & lifecycle management:

    • Build enrichment logic that joins, standardizes, and contextualizes data across domains using shared definitions and reusable patterns.
    • Support historical tracking, point-in-time correctness, and dataset versioning so downstream users can confidently analyze changes and impacts over time.
  • Observability, reliability & operational excellence:

    • Instrument ingestion pipelines with strong observability: freshness, latency, data quality, and cost metrics.
    • Contribute to SLOs, alerting, and incident response playbooks so data failures are visible, diagnosable, and recoverable.
  • Governance & compliance by design:

    • Apply consistent access control, classification, and privacy protections at ingestion time.
    • Ensure sensitive data is properly masked, minimized, or anonymized by default, and that all data flows you own are auditable and traceable.
  • Enable self-service & standardization:

    • Contribute to standardized ingestion templates, shared libraries, and platform tooling that enable teams to onboard new data sources independently.
    • Improve discoverability, documentation, and metadata so datasets you own are easy to find and trust without relying on tribal knowledge.
  • Cross-team collaboration & ownership:

    • Work closely with Product, Backend, Analytics, and ML partners to align on ingestion requirements and trade-offs.
    • Build strong working relationships across teams. Mentor junior team members and actively contribute to a culture of shared data quality standards and data contracts.

Requirements

  • Hands-on experience building components of large, high-scale applications (e.g., data pipelines, well-structured APIs, efficient algorithms).
  • Solid experience working in platform or data engineering teams (or equivalent) with the ability to deliver within a multi-stakeholder environment.
  • Familiarity with cloud platforms (AWS/GCP or equivalent) and modern DevOps practices.
  • Hands-on experience designing and implementing real-time and batch data processing pipelines using modern frameworks like Spark, Flink, Spark Streaming, Kafka, Debezium, etc.
  • Experience with orchestration tools such as Airflow, dbt, or similar.
  • Exceptional problem-solving skills paired with a proactive, innovative mindset focused on continuous improvement.
  • Strong communication and cross-functional collaboration skills (English level B2+)

Nice to Have

  • Cross-team collaboration
  • Mentoring

We Offer

  • An open, collaborative, dynamic, and diverse culture;
  • A generous monthly allowance for lessons on Preply.com, Learning & Development budget, and time off for your self-development.
  • Not in Barcelona? We offer an attractive relocation package to join us in our Preply Barcelona Hub
  • A competitive financial package with equity, leave allowance, and health insurance;
  • Access to free mental health support platforms;
  • Access to Gympass-partnered wellness and gym centers throughout Spain to promote and support well-being and physical health;
  • The opportunity to unlock the potential of learners and tutors through language learning and teaching in 175 countries (and counting!).

About the Company

Preply.com is committed to creating an inclusive environment where people of diverse backgrounds can thrive. We believe that the presence of different opinions and viewpoints is a key ingredient for our success as a multicultural Ed-Tech company. That means that Preply will consider all applications for employment without regard to race, color, religion, gender identity or expression, sexual orientation, national origin, disability, age or veteran status.

© Preply. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

Preply is an online language learning marketplace, connecting tutors to hundreds of thousands of learners in 180 countries worldwide. More than 90,000 tutors teach over 50 languages, powered by a machine-learning algorithm that recommends the best tutors for each learner.Backed by some of the world’s leading investors, Preply is on a mission to unlock human potential through learning. Fueled by a belief that live engagement with a teacher is still the most effective way to learn a new skill, Preply is building a personalized learning space that will enable individual learners to reach their goals in the fastest way possible.Powered by a tenfold increase in revenues over the last three years, Preply is now leading the online language tutoring segment globally and has 700+ employees of over 60 nationalities based across Barcelona, Kyiv, New York, and London. Preply is driven by a culture of experimentation and data-driven learnings, focused on building best-in-class consumer and enterprise solutions.Get to know more about life at Preply! Check out the resources below:— Engineering Blog:medium.com/preply-engineering— Discover what it’s like to work at Preply through the stories of our employees:wearepreply.medium.com— Check the news from Preply’s offices in Instagram:www.instagram.com/wearepreply— More recent investment round:techcrunch.com/... m-and-doubles-down-on-ai— Co-founder’s interview for Big Moneywww.youtube.com/watch?v=IY4oaRqFafQRelated articles:dou.ua/... cles/how-i-work-voloshyndou.ua/... icles/dou-books-voloshyndou.ua/lenta/articles/dockerdou.ua/lenta/articles/vagrantwww.gilles-bertrand.com/... ithout-budget-preply.htmldzone.com/... pment-process-with-docker

More at Preply

All 118 roles

Popular searches