MLOps / Cloud Engineer
Primary stack
PrometheusDockerKubernetesPythonAWSGrafanaCI/CDKubeflowAzureTerraform
Job description
About the Position
Your vision. Our solutions.
We are seeking an experienced MLOps / Cloud Engineer to join our team in a remote role based in Europe.
Project Description
- Design, development, and maintenance of reliable cloud and AI/ML platforms for massive enterprise utilization
- Collaboration within a dedicated DevOps team on the development of end-to-end ML workflows and solving complex business challenges through the integration of GenAI and LLM models into a production environment
- Architecture and management of infrastructure as code with intensive use of AWS, Azure, Kubernetes, Terraform technologies, and advanced development in Python
- Building and optimizing CI/CD pipelines through GitHub Actions and ensuring comprehensive secure deployment
- Setting up and maintaining advanced monitoring and logging using the Prometheus, Grafana, and Loki stack (without the need for on-call support)
- Containerizing applications via Docker and managing the machine learning lifecycle using tools like Kubeflow and SageMaker
- Collaboration takes place in a 100% remote mode within the entire EU
Responsibilities
- Design, development, and maintenance of reliable cloud and AI/ML platforms
- Collaborate on the development of end-to-end ML workflows
- Architecture and management of infrastructure as code
- Building and optimizing CI/CD pipelines
- Setting up and maintaining advanced monitoring and logging systems
- Containerizing applications and managing the machine learning lifecycle
Requirements
- Advanced experience with:
- Software architecture development and design in Python (minimum 5 years of experience)
- Cloud platform, primarily AWS or Azure (including services like SageMaker or Bedrock)
- Container orchestration and cluster management in Kubernetes (e.g., EKS)
- Infrastructure as code and its automation using Terraform
- MLOps tools, especially integration and maintenance of Kubeflow
- Experience with:
- Containerization and optimization of performance and security using Docker
- Creation and debugging of complex CI/CD pipelines, ideally via GitHub Actions
- Deployment and work with end-to-end ML workflows and integration of LLM and GenAI
- Monitoring and logging systems like Prometheus, Grafana, or Loki
- Advanced knowledge of:
- CI/CD principles, security practices, and DevOps culture
- Theoretical foundations and production deployment of a wide range of ML algorithms
- Model registry, model performance monitoring, and data quality monitoring
Nice to Have
- Experience with advanced Kubernetes features, such as operators
- Experience with the Dynatrace platform
We Offer
- Opportunity to work in a 100% remote role within the entire EU
- Collaborate with a dedicated DevOps team on cutting-edge projects
- Work with state-of-the-art technologies in cloud and AI/ML
About the Company
The Workly s.r.o.
4D CENTER, Kodaňská 1441/46,
101 00, Praha 10.
Contact:
[email protected]
- 420 775 371 991
©2023, All right reserved.
© The Workly. This job description was sourced from the employer's public career page. TheJob is not the employer — we index the posting and route candidates to the source. All content rights and hiring decisions belong to the employer.
More at The Workly
All 266 rolesTH
Data Engineer/ Energy Trading Developer
The Workly · Germany
Salary not specified checked 2w ago
TH
Project Manager for Calypso / Treasury Data Warehouse
The Workly · Austria
Salary not specified checked 2w ago
TH
AWS Data Movement Engineer
The Workly · Prague
Salary not specified checked 2w ago
TH
Senior Product Analyst
The Workly · Prague
Salary not specified checked 2w ago