EPAM logo

Senior AI Engineer with Kubernetes

EPAM·Salary not specified

Primary stack

DockerOpenAI APIHugging FaceKubernetesGoPythonC++CohereMilvusJava

Nice to have's

LangChainLlamaIndexGPU schedulingresource optimizationinference accelerationhybrid searchmetadata filteringindex tuningLLM evaluationgovernancetracingmonitoringCI/CD pipelinesInfrastructure-as-Codecloud-native deployment practicesOil and Gas industryDataiku DSSSRE practices

Job description

About the Position

We are seeking a highly skilled Senior AI Engineer with Kubernetes and vector database technologies to join our team. In this role, you will design, build, and scale production-grade AI systems, working with cutting-edge LLM frameworks, embeddings, and cloud-native infrastructure to deliver robust and high-performance solutions.

Responsibilities

  • Deploy and manage Milvus vector databases, including schema design and index tuning (HNSW, IVF-FLAT)
  • Build and maintain embedding and LLM pipelines using OpenAI API, Hugging Face, or Cohere
  • Manage Kubernetes clusters, Helm charts, and containerized microservices in production
  • Develop and maintain Docker containerization workflows, including multi-stage builds and registry management
  • Design and deliver production-grade Python applications, integrating with Go, Java, or C++ where required
  • Integrate object storage systems such as AWS S3, MinIO, or Google Cloud Storage
  • Evaluate and implement alternative vector database solutions, including Qdrant, Pinecone, and Weaviate
  • Collaborate cross-functionally with team members to deliver reliable, scalable AI services
  • Ensure operational excellence, observability, and performance of deployed AI workloads

Requirements

  • Bachelor's degree in Engineering with 5+ years of relevant experience
  • Expertise in Milvus deployment, schema design, and index tuning (HNSW, IVF-FLAT)
  • Familiarity with vector database alternatives such as Qdrant, Pinecone, Weaviate, PGVector, or Chroma
  • Proficiency in building embedding and LLM pipelines using OpenAI API, Hugging Face, or Cohere
  • Skills in Kubernetes cluster management, Helm charts, and containerized microservices
  • Background in Docker containerization, multi-stage builds, and registry management
  • Production-level Python development along with Go, Java, or C++
  • Knowledge of object storage integration, including AWS S3, MinIO, or Google Cloud Storage
  • Excellent verbal and written communication skills with strong team collaboration abilities
  • Proficiency in English at an Upper-Intermediate level (B2) or higher

Nice to Have

  • Experience supporting large-scale RAG applications and multi-agent platforms
  • Hands-on familiarity with LangChain, LlamaIndex, or custom pipelines
  • Understanding of GPU scheduling, resource optimization, and inference acceleration
  • Production experience with hybrid search, metadata filtering, and index tuning
  • Implementation of LLM evaluation, governance, tracing, and monitoring tools
  • Familiarity with CI/CD pipelines, Infrastructure-as-Code, and cloud-native deployment practices
  • Prior work experience in the Oil and Gas industry
  • Experience with Dataiku DSS
  • Knowledge of SRE practices

We Offer

  • Opportunity to work remotely in Georgia, USA, or from any location
  • Competitive contractor terms

About the Company

EPAM Systems is a global software engineering and product development company, delivering digital platforms and large-scale systems for the world’s leading organizations.

© EPAM. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.

EPAM helps organizations innovate their business processes and rethink the way they manage their businesses so they can remain competitive in this new digital age.

More at EPAM

All 2376 roles

Similar jobs

Popular searches