OnHires logo

Senior Data Collection Engineer (Python / Web Scraping)

OnHires·Salary not specified

Primary stack

NoSQLDockerSeleniumPythonLinuxHTTPGitHTMLSQLPlaywright

Nice to have's

anti-bot techniquesbrowser fingerprintingCAPTCHA solvingproxy managementblocking mitigationasyncioCelerydistributed task queuesmonitoring systemsdata quality validationAI-assisted development tools

Job description

About the Position

Senior Data Collection Engineer (Python / Web Scraping)

Our client is a fast-growing, remote-first B2B SaaS company where large-scale data collection is at the core of the product. As the platform continues to grow globally, they're looking for an experienced Data Collection Engineer to help build and scale the infrastructure behind one of their core product capabilities.

This is a hands-on engineering role focused on designing resilient scraping infrastructure, overcoming sophisticated anti-bot systems, and collecting high-quality data at scale. You'll own the entire lifecycle of large-scale data collection pipelines, ensuring they remain reliable, scalable, and resilient as the product continues to grow.

Responsibilities

  • Infrastructure Strategy & Architecture: Architect, build, and maintain the core infrastructure behind our large-scale asynchronous data collection platform.
  • Advanced Resilience Engineering: Design, implement, and continuously improve sophisticated anti-blocking strategies, including browser fingerprinting, proxy rotation, CAPTCHA handling, and other techniques required to maintain reliable data collection.
  • Core Development: Design, develop, test, and maintain robust scraping components using Python and modern scraping frameworks such as Playwright, Scrapy, Selenium, Requests, and related tools.
  • Operational Excellence: Build monitoring, alerting, and logging systems that help identify issues quickly and continuously improve scraper reliability and data quality.
  • Data Pipelines & Integrations: Develop and maintain scalable data ingestion pipelines and integrations with internal and external REST APIs.
  • DevOps & Automation: Contribute to infrastructure automation using Docker, CI/CD pipelines, Linux environments, and related DevOps practices.
  • Collaboration: Work closely with other engineers to improve our scraping platform, establish engineering standards, and mentor less experienced teammates.

Requirements

  • Strong commercial experience building high-volume web scraping and data collection systems using Python.
  • Deep practical knowledge of anti-bot techniques, including browser fingerprinting, CAPTCHA solving, proxy management, and blocking mitigation.
  • Strong understanding of asynchronous programming, browser automation, HTML parsing, HTTP protocols, and REST APIs.
  • Hands-on experience with Playwright, Scrapy, Selenium, or similar scraping frameworks.
  • Experience working with Docker, Linux, Git, and modern software development practices.
  • Familiarity with SQL and NoSQL databases.
  • Strong ownership mindset with the ability to independently drive complex technical projects.
  • Fluent English communication skills.

Nice to Have

  • Experience with advanced asynchronous frameworks (asyncio, Celery, distributed task queues).
  • Experience building monitoring and data quality validation systems.
  • Experience mentoring engineers or helping technical teams scale.
  • Experience using modern AI-assisted development tools (e.g., Claude Code, Cursor, Codex, Windsurf, or similar).

We Offer

  • High ownership and the opportunity to make a measurable impact on a rapidly growing product.
  • Remote-first culture with flexible working arrangements.
  • Competitive compensation package.
  • Personal and professional development through ongoing learning and coaching.
  • A collaborative international engineering team solving technically challenging problems.
  • Optional office near Berlin at the Wildau Tech University campus.

About the Company

Our client is a fast-growing, remote-first B2B SaaS company where large-scale data collection is at the core of the product. As the platform continues to grow globally, they're looking for an experienced Data Collection Engineer to help build and scale the infrastructure behind one of their core product capabilities.

© OnHires. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

OnHires was founded in 2021 by Vasyl Grygorovych and Viktoriia Ihnatieva to help companies find and hire the best tech talent the world offers.As business owners, both Victoria and Vasyl were frustrated by how difficult and costly it was to quickly scale development teams. Having lived in multiple countries and traveled to every continent in the world, they realized that great talent could be found anywhere. At the same time, they realized that it’s difficult for both companies and talent to find each other outside of their own geographical location. In an effort to solve this problem, Viktoriia and Vasyl founded OnHires.You can read more about their background here:🚀 Our missionTo connect great companies with even greater talent. We believe that talent can be found anywhere but this is not always the case with opportunities. We are on a mission to bridge that gap and create equal opportunities for everyone. To achieve this, we are building an agency that will help companies unlock global talent.

More at OnHires

All 119 roles

Similar jobs

Popular searches