Data Scientist
Primary stack
Nice to have's
Job description
About the Position
Your vision. Our solutions.
We are seeking a Data Scientist to join our team and work on cutting-edge AI projects, developing advanced client solutions.
Responsibilities
Graph AI, Knowledge Graphs & Graph Data Science
- Design and implement graph-based solutions, including graph features, graph analytics, and graph ML workflows aligned to business use cases.
- Build and operate knowledge graphs (schema/ontology, entity resolution, ingestion/update pipelines) in graph databases (e.g., Neo4j) and apply graph representation learning (embeddings; GNNs such as GCN/GAT) to support prediction, ranking, and reasoning.
NLP, LLMs, Semantic Engineering & RAG
- Build LLM-powered capabilities end-to-end: prompt/tool design, semantic engineering, grounding strategies, and production-grade RAG pipelines (chunking, embeddings, reranking, citations/attribution where applicable).
- Deliver NLP features such as text-to-SQL/semantic parsing, classification/extraction, and domain adapters; incorporate user/context signals to enable search, retrieval, and personalization across structured and unstructured sources.
- Apply fine-tuning and adaptation approaches (lightweight tuning, instruction tuning, preference optimization where appropriate) and establish evaluation harnesses for quality, safety, and robustness.
Core Data Science, Machine Learning & Deep Learning (Core Requirement)
- Formulate problems, define success metrics, and design experiments; select and train ML/DL models appropriate for the task, data shape, and operational constraints.
- Perform rigorous evaluation and error analysis (offline metrics, slice analysis, ablations), and iterate on features, architectures, and data to improve performance and robustness.
- Ensure reproducibility and responsible delivery: version data/models, document assumptions, manage bias/quality risks, and contribute to model monitoring and drift detection.
Software Engineering, Architecture & Workflow Design
- Design end-to-end solution architectures and workflows that connect data sources, graph layers, retrieval, models, and application surfaces into coherent, maintainable systems.
- Implement production services and APIs (batch + real-time) that expose model, retrieval, and graph capabilities with clear contracts, security considerations, and performance targets.
- Translate R&D prototypes into production-ready components; collaborate with architects, pod leads, and UX/FE to ensure scalable designs and high-quality delivery.
Agile Delivery, Quality, Observability & E2E Ownership
- Own delivery in agile pods: break down problems, estimate, ship increments each sprint, and continuously improve based on feedback and measured outcomes.
- Build quality into the workflow: automated tests, evaluation suites for ML/LLM behavior, CI checks, code reviews, and clear operational runbooks.
- Instrument systems for reliability: monitoring for latency/cost, failure modes, and drift; logging for auditability and troubleshooting; and support safe rollout/rollback patterns.
Compute, Performance & Scaling (Preferred)
- Optimize training and inference pipelines for performance and cost, including efficient batching, caching, quantization/acceleration approaches where appropriate, and practical profiling.
- Leverage GPUs and accelerated compute for deep learning/LLM workloads; understand memory constraints, throughput bottlenecks, and practical optimization trade-offs.
- Apply distributed compute patterns when needed (e.g., data processing, training, retrieval indexing) and design systems that scale reliably across cloud resources.
Requirements
Required Skills & Experience
- Graph, Graph Data Science & Knowledge Graphs: graph theory fundamentals; graph analytics/graph data science; knowledge graph modeling (schemas/ontologies) and semantic layers; graph databases (e.g., Neo4j); graph embeddings; GNNs and graph ML tooling (e.g., GCN, GAT).
- NLP, LLMs & Semantic Engineering: modern NLP and language systems; semantic engineering; search/retrieval (lexical + vector + hybrid) and reranking; RAG; personalization; text-to-SQL; prompt engineering; fine-tuning/adaptation patterns.
- Core ML / Deep Learning / Data Science (Core Requirement): strong grounding in machine learning and deep learning; applied data science; ability to design experiments, evaluate models, and reason about quality, robustness, and bias.
- Hands-on Software Engineering & Delivery: solid, hands-on development experience (production Python and/or related stack); agile ways of working; R&D prototyping; architecture and workflow understanding and design across services, APIs, and data pipelines.
Technical Skills
- Strong Python engineering experience (3–6+ years).
- Experience with retrieval/search systems (lexical + vector + hybrid), embeddings, and ranking/reranking.
- Exposure to graph data science and/or knowledge graphs (modeling, ingestion, graph features/embeddings; experience with graph databases such as Neo4j); experience with GNNs (e.g., GCN/GAT) is a plus.
- Hands-on experience building RAG, agentic workflows, or similar AI patterns.
- Familiarity with model evaluation, testing, and observability tools.
- Solid knowledge of APIs, microservices, and data-centric integrations.
- Comfort with cloud platforms (Azure, AWS, GCP) and CI/CD workflows.
- Preferred: exposure to GPUs, accelerated training/inference, and/or distributed compute (e.g., Ray, Spark, Dask, or similar patterns).
Foundational Engineering Skills
- Excellent debugging, profiling, and performance optimization skills.
- Strong understanding of source control (Git), branching, PR reviews.
- Ability to work in structured agile environments with clear sprint increments.
Mindset
- High ownership, curiosity, and willingness to experiment.
- Loves hard problems and iterating on prototypes.
- Comfortable working in fast-changing frontier environments.
- Values clarity, structure, and quality in code.
We Offer
- Fully remote position with occasional in-person team sessions/workshops/gatherings (likely to take place in Prague).
- Minimum 2-6pm CET overlap with US hours, preferred 2-7pm CET.
About the Company
We are seeking a Data Scientist to join our team and work on cutting-edge AI projects, developing advanced client solutions.
Start date: ASAP
HackerRank Challenge: Yes
Remote vs Onsite: Fully remote, with possible occasional in-person team sessions/workshops/gatherings (likely to take place in Prague)
US Hours overlap needed: Minimum 2-6pm CET, preferred 2-7pm CET
We Expect You to Have:
Thank you! We will be in touch shortly
© The Workly. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
More at The Workly
All 265 rolesData Engineer/ Energy Trading Developer
The Workly · Germany
Project Manager for Calypso / Treasury Data Warehouse
The Workly · Austria
AWS Data Movement Engineer
The Workly · Prague
Senior Product Analyst
The Workly · Prague

