The Workly logo

Data Scientist

The Workly·Salary not specified

Primary stack

Machine LearningNLPPythonRAG

Nice to have's

Retrieval SystemsEmbeddingsRanking/RerankingGPUsDistributed Compute

Job description

About the Position

Your vision. Our solutions.

We are seeking a Data Scientist to join our team and work on cutting-edge AI projects, developing advanced client solutions.

Responsibilities

Graph AI, Knowledge Graphs & Graph Data Science

  • Design and implement graph-based solutions, including graph features, graph analytics, and graph ML workflows aligned to business use cases.
  • Build and operate knowledge graphs (schema/ontology, entity resolution, ingestion/update pipelines) in graph databases (e.g., Neo4j) and apply graph representation learning (embeddings; GNNs such as GCN/GAT) to support prediction, ranking, and reasoning.

NLP, LLMs, Semantic Engineering & RAG

  • Build LLM-powered capabilities end-to-end: prompt/tool design, semantic engineering, grounding strategies, and production-grade RAG pipelines (chunking, embeddings, reranking, citations/attribution where applicable).
  • Deliver NLP features such as text-to-SQL/semantic parsing, classification/extraction, and domain adapters; incorporate user/context signals to enable search, retrieval, and personalization across structured and unstructured sources.
  • Apply fine-tuning and adaptation approaches (lightweight tuning, instruction tuning, preference optimization where appropriate) and establish evaluation harnesses for quality, safety, and robustness.

Core Data Science, Machine Learning & Deep Learning (Core Requirement)

  • Formulate problems, define success metrics, and design experiments; select and train ML/DL models appropriate for the task, data shape, and operational constraints.
  • Perform rigorous evaluation and error analysis (offline metrics, slice analysis, ablations), and iterate on features, architectures, and data to improve performance and robustness.
  • Ensure reproducibility and responsible delivery: version data/models, document assumptions, manage bias/quality risks, and contribute to model monitoring and drift detection.

Software Engineering, Architecture & Workflow Design

  • Design end-to-end solution architectures and workflows that connect data sources, graph layers, retrieval, models, and application surfaces into coherent, maintainable systems.
  • Implement production services and APIs (batch + real-time) that expose model, retrieval, and graph capabilities with clear contracts, security considerations, and performance targets.
  • Translate R&D prototypes into production-ready components; collaborate with architects, pod leads, and UX/FE to ensure scalable designs and high-quality delivery.

Agile Delivery, Quality, Observability & E2E Ownership

  • Own delivery in agile pods: break down problems, estimate, ship increments each sprint, and continuously improve based on feedback and measured outcomes.
  • Build quality into the workflow: automated tests, evaluation suites for ML/LLM behavior, CI checks, code reviews, and clear operational runbooks.
  • Instrument systems for reliability: monitoring for latency/cost, failure modes, and drift; logging for auditability and troubleshooting; and support safe rollout/rollback patterns.

Compute, Performance & Scaling (Preferred)

  • Optimize training and inference pipelines for performance and cost, including efficient batching, caching, quantization/acceleration approaches where appropriate, and practical profiling.
  • Leverage GPUs and accelerated compute for deep learning/LLM workloads; understand memory constraints, throughput bottlenecks, and practical optimization trade-offs.
  • Apply distributed compute patterns when needed (e.g., data processing, training, retrieval indexing) and design systems that scale reliably across cloud resources.

Requirements

Required Skills & Experience

  • Graph, Graph Data Science & Knowledge Graphs: graph theory fundamentals; graph analytics/graph data science; knowledge graph modeling (schemas/ontologies) and semantic layers; graph databases (e.g., Neo4j); graph embeddings; GNNs and graph ML tooling (e.g., GCN, GAT).
  • NLP, LLMs & Semantic Engineering: modern NLP and language systems; semantic engineering; search/retrieval (lexical + vector + hybrid) and reranking; RAG; personalization; text-to-SQL; prompt engineering; fine-tuning/adaptation patterns.
  • Core ML / Deep Learning / Data Science (Core Requirement): strong grounding in machine learning and deep learning; applied data science; ability to design experiments, evaluate models, and reason about quality, robustness, and bias.
  • Hands-on Software Engineering & Delivery: solid, hands-on development experience (production Python and/or related stack); agile ways of working; R&D prototyping; architecture and workflow understanding and design across services, APIs, and data pipelines.

Technical Skills

  • Strong Python engineering experience (3–6+ years).
  • Experience with retrieval/search systems (lexical + vector + hybrid), embeddings, and ranking/reranking.
  • Exposure to graph data science and/or knowledge graphs (modeling, ingestion, graph features/embeddings; experience with graph databases such as Neo4j); experience with GNNs (e.g., GCN/GAT) is a plus.
  • Hands-on experience building RAG, agentic workflows, or similar AI patterns.
  • Familiarity with model evaluation, testing, and observability tools.
  • Solid knowledge of APIs, microservices, and data-centric integrations.
  • Comfort with cloud platforms (Azure, AWS, GCP) and CI/CD workflows.
  • Preferred: exposure to GPUs, accelerated training/inference, and/or distributed compute (e.g., Ray, Spark, Dask, or similar patterns).

Foundational Engineering Skills

  • Excellent debugging, profiling, and performance optimization skills.
  • Strong understanding of source control (Git), branching, PR reviews.
  • Ability to work in structured agile environments with clear sprint increments.

Mindset

  • High ownership, curiosity, and willingness to experiment.
  • Loves hard problems and iterating on prototypes.
  • Comfortable working in fast-changing frontier environments.
  • Values clarity, structure, and quality in code.

We Offer

  • Fully remote position with occasional in-person team sessions/workshops/gatherings (likely to take place in Prague).
  • Minimum 2-6pm CET overlap with US hours, preferred 2-7pm CET.

About the Company

We are seeking a Data Scientist to join our team and work on cutting-edge AI projects, developing advanced client solutions.

Start date: ASAP

HackerRank Challenge: Yes

Remote vs Onsite: Fully remote, with possible occasional in-person team sessions/workshops/gatherings (likely to take place in Prague)

US Hours overlap needed: Minimum 2-6pm CET, preferred 2-7pm CET

We Expect You to Have:

Thank you! We will be in touch shortly

© The Workly. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

More at The Workly

All 265 roles

Similar jobs

Popular searches