Lead Platform Engineer - HPC, Kubernetes
0
4500 - 6300 €/mėn.
Prieš mokesčius
Job description:
We are looking for a Lead Platform Engineer to support a customer that develops and manages several HPC clusters across AWS, CoreWeave, GCP and other providers, operating several thousand GPUs today and scaling 10x. This role is Kubernetes-heavy and requires strong software engineering skills to operate multi-cloud platform infrastructure where misconfigurations or failed upgrades cost thousands of GPU-hours, and at this scale, new kinds of failure come up routinely.
Feel free to work remotely from anywhere across Latvia or connect with colleagues at our Riga office.
- Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle, node pool management, networking policy, and stability during rapid growth
- Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI, with more providers coming
- Manage job scheduling to allocate GPU compute across training and inference workloads
- Define and maintain SLIs/SLOs
- Build monitoring and alerting systems
- Take part in incident response and write post-incident reviews
- Build tooling and automation in production-quality code
- Coordinate daily with the Networking, Storage, Security and AI/ML platform teams
Requirements:
- 5+ years of experience in infrastructure engineering, cloud platforms or HPC
- Expertise in Kubernetes at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
- Proficiency in advanced Python with experience writing production-grade tools, not only scripts
- Proficiency in Terraform, writing and reviewing infrastructure as code daily
- Working knowledge of AWS, including EC2, S3, EFS and FSx for Lustre
- Background in Site Reliability Engineering
- English proficiency at B2 level or higher
Nice to have
- Familiarity with Amazon Elastic Kubernetes Service, Google Kubernetes Engine and Google Cloud Platform
- Knowledge of High-performance computing (HPC), Slurm, Lustre and Amazon FSx
- Familiarity with Go, Rust and C++
- Familiarity with CI/CD
Company offers:
- Engineering Heritage: Best-in-class experts sharing a culture of engineering excellence and tackling complex engineering challenges for over 30 years.
- Advanced Tech Stack: Innovative projects where you can apply or enhance your expertise in Cloud, Data, AI, and other emerging technologies.
- World-Class Clients: Work closely with 340+ of the Forbes Global 2000 on creating disruptive solutions that make a global impact.
- Professional Growth: Exceptional support for career development with comprehensive resources for upskilling or reskilling in pioneering practices.
- GenAI Community: Strong AI competencies with 600+ experts across 55+ locations driving GenAI-enabled transformation journeys.
- Entrepreneurial Culture: If you're passionate and dedicated to improving business transformation, we provide the support you need to bring your ideas to life.
- Hybrid Setup: The flexibility to work from any location in Latvia, whether it's your home or our office in Riga.
- Other Benefits: Additional vacation and trust days, private health insurance, Employee Stock Purchase Plan and more.
Miestas:
Latvija
Nuotolinis darbas:
Taip
Laikas:
Visa darbo diena
Galioja iki:
01/11/2026
Kandidatavimas vyks EPAM SISTEMOS, UAB įmonės puslapyje
Persiųsti