The Workly logo

AI QA engineer

The Workly·Salary not specified

Primary stack

DevOpsPythonRAGCI/CD

Job description

About the Position

Your vision. Our solutions.

We are seeking an AI QA engineer to ensure the reliability, safety, and ethical behavior of AI systems across multiple pods.

Responsibilities

  1. Own the Test Strategy (App + Data + Model)

    • Define the unified QA strategy across UI, backend, retrieval, and AI components.
    • Design test plans that match the pod’s frontier experiment, NFRs, and acceptance criteria.
    • Ensure coverage across functional, non-functional, and behavioral dimensions.
  2. Build Evaluation Suites

    • Develop evaluation frameworks for LLMs, retrieval, and agent workflows.
    • Implement scenario-based evals, confusion tests, regression tests, and behavioral probes.
    • Create both automated and manual eval paths for different model behaviors.
  3. Test Automation

    • Build automated test harnesses for Python services, agents, APIs, pipelines, and integrations.
    • Integrate test automation into CI/CD pipelines.
    • Maintain automated testing scripts that catch regressions early.
  4. Data Quality & ML Evaluation

    • Validate data correctness, consistency, and completeness for all RAG and agent pipelines.
    • Test embedding quality, retrieval accuracy, and ranking performance.
    • Identify hallucination patterns, reasoning failures, and model drift.
  5. Red-Team & Edge-Case Scenario Design

    • Simulate high-risk or adversarial scenarios to uncover weaknesses.
    • Create structured red-team tests for safety, compliance, and robustness.
    • Validate handling of ambiguous inputs, missing data, or malformed requests.
  6. Observability, Monitoring & Drift Detection

    • Define metrics and logs to monitor agent behavior, latency, cost, and error modes.
    • Work with AI Ops to implement dashboards and alerts for reliability tracking.
    • Detect and escalate drift, bias, or degradation trends quickly.
  7. Defect Management & Triage

    • Runs the defect triage workflow, partnering with the Tech Lead and engineers.
    • Diagnose root causes and categorize failures across UI, API, data, or model layers.
    • Ensure clear, crisp documentation with reproduction steps.

Requirements

Technical Skills

  • Strong Python scripting for test automation and scenario evaluation.
  • Experience with ML evaluation tools, LLM/RAG testing, or model benchmarking suites.
  • Familiarity with vector DBs, retrieval systems, and agent workflows.
  • Understanding of CI/CD pipelines, DevOps tooling, and observability platforms.
  • Ability to query data, validate embeddings, and test ranking/precision metrics.

QA & Risk Expertise

  • 5–6+ years in QA, SDET, testing, or evaluation-focused ML engineering.
  • Strong instincts for edge cases, risk modes, and adversarial failures.
  • Experience designing tests for systems with nondeterministic or probabilistic behavior (preferred).

Mindset

  • Curious, skeptical, and systematic.
  • Thrives on breaking things to make them better.
  • Strong communicator - crisp defect reporting is non-negotiable.
  • High ownership and discipline; loves clear structure and tight loops.

We Offer

  • Fully remote position with possible occasional in-person team sessions/workshops/gatherings (i.e. 1x quarter) likely to take place in Prague.
  • Minimum 2-6pm CET overlap with preferred 2-7pm CET.

About the Company

The Workly s.r.o.

4D CENTER, Kodaňská 1441/46,

101 00, Praha 10.

Contact

[email protected]

  • 420 775 371 991

©2023, All right reserved.

© The Workly. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

More at The Workly

All 265 roles

Popular searches