Tech Stack
Job Description, Responsibilities & Requirements
About the Position
We are looking for a DevOps/MLOps Engineer to join our client’s distributed engineering team and help build and maintain the cloud infrastructure that powers AI-driven applications.
In this role, you will focus on infrastructure, deployment pipelines, automation, and observability, ensuring reliability, scalability, and efficient operations across production environments that support AI workloads.
Please note that this role is focused on infrastructure and platform engineering. You will not work on proprietary ML models or training assets.
Our client is a fast-growing US-based AI-first startup building next-generation technology for the fashion and e-commerce industry. Their platform leverages artificial intelligence to deliver photorealistic virtual try-on, AI-powered sizing prediction, and personalized outfit recommendations, helping leading fashion brands create more engaging shopping experiences.
Responsibilities
- Build, manage, and optimize AWS EKS/Kubernetes infrastructure supporting production services
- Containerize and deploy applications using Docker
- Provision and maintain infrastructure using Terraform
- Support GPU-enabled infrastructure and workload orchestration with Slurm
- Build and maintain CI/CD pipelines using GitHub Actions
- Implement and maintain monitoring and observability using Grafana, Prometheus, and AWS CloudWatch
- Develop Python scripts and automation tools to improve operational efficiency
- Monitor cloud infrastructure utilization and contribute to cloud cost optimization initiatives
- Collaborate closely with software engineers to ensure reliable and scalable platform operations
- Participate in code reviews and follow engineering, security, and infrastructure best practices
Requirements
- 4+ years of professional experience as a DevOps, or MLOps Engineer
- Strong production experience with AWS, Kubernetes (EKS), Docker, and Terraform
- Experience with Slurm for workload orchestration in production environments
- Experience building and maintaining CI/CD pipelines using GitHub Actions
- Experience with monitoring and observability tools, including Grafana, Prometheus, and AWS CloudWatch
- Python scripting experience for automation and infrastructure tooling
- Experience supporting production cloud infrastructure
- Strong communication skills and ability to collaborate with distributed teams
- Upper-Intermediate or higher English
Nice to Have
- Experience supporting GPU-based infrastructure or AI/ML workloads
- Experience with MLflow
- Experience with FinOps or cloud cost optimization
We Offer
- Competitive salary
- Vacation (up to 20 working days)
- Paid sick leaves (10 working days)
- National Holidays as time off
- Flexible working schedule, remote format
- Direct cooperation with the customer
- Dynamic environment with low level of bureaucracy and great team spirit
- Challenging projects in diverse business domains and a variety of tech stacks
- Communication with Top/Senior level specialists to strengthen your hard skills
- Online teambuildings
Locations
- Kyiv, Ukraine
- Warsaw, Poland
- Sao Paolo, Brasil
- London, United Kingdom
- Remote
Contact
Phone: +44 772 611 66 87
Email: [email protected]
Development offices:
- Kyiv, Ukraine
- Warsaw, Poland
- Sao Paolo, Brasil
- London, United Kingdom