Grįžti į skelbimus
EPAM SISTEMOS, UAB

Lead Platform Engineer - HPC, Kubernetes

0
4500 - 6300 €/mėn.
Prieš mokesčius

Job description:

We are looking for a Lead Platform Engineer to support a customer that develops and manages several HPC clusters across AWS, CoreWeave, GCP and other providers, operating several thousand GPUs today and scaling 10x. This role is Kubernetes-heavy and requires strong software engineering skills to operate multi-cloud platform infrastructure where misconfigurations or failed upgrades cost thousands of GPU-hours, and at this scale, new kinds of failure come up routinely.

Feel free to work remotely from anywhere across Latvia or connect with colleagues at our Riga office.

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale, including cluster lifecycle, node pool management, networking policy, and stability during rapid growth
  • Provision HPC infrastructure through CI/CD across AWS, CoreWeave, GCP and OCI, with more providers coming
  • Manage job scheduling to allocate GPU compute across training and inference workloads
  • Define and maintain SLIs/SLOs
  • Build monitoring and alerting systems
  • Take part in incident response and write post-incident reviews
  • Build tooling and automation in production-quality code
  • Coordinate daily with the Networking, Storage, Security and AI/ML platform teams

Requirements:

  • 5+ years of experience in infrastructure engineering, cloud platforms or HPC
  • Expertise in Kubernetes at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets
  • Proficiency in advanced Python with experience writing production-grade tools, not only scripts
  • Proficiency in Terraform, writing and reviewing infrastructure as code daily
  • Working knowledge of AWS, including EC2, S3, EFS and FSx for Lustre
  • Background in Site Reliability Engineering
  • English proficiency at B2 level or higher

Nice to have

  • Familiarity with Amazon Elastic Kubernetes Service, Google Kubernetes Engine and Google Cloud Platform
  • Knowledge of High-performance computing (HPC), Slurm, Lustre and Amazon FSx
  • Familiarity with Go, Rust and C++
  • Familiarity with CI/CD

Company offers:

  • Engineering Heritage: Best-in-class experts sharing a culture of engineering excellence and tackling complex engineering challenges for over 30 years.
  • Advanced Tech Stack: Innovative projects where you can apply or enhance your expertise in Cloud, Data, AI, and other emerging technologies.
  • World-Class Clients: Work closely with 340+ of the Forbes Global 2000 on creating disruptive solutions that make a global impact.
  • Professional Growth: Exceptional support for career development with comprehensive resources for upskilling or reskilling in pioneering practices.
  • GenAI Community: Strong AI competencies with 600+ experts across 55+ locations driving GenAI-enabled transformation journeys.
  • Entrepreneurial Culture: If you're passionate and dedicated to improving business transformation, we provide the support you need to bring your ideas to life.
  • Hybrid Setup: The flexibility to work from any location in Latvia, whether it's your home or our office in Riga.
  • Other Benefits: Additional vacation and trust days, private health insurance, Employee Stock Purchase Plan and more.

Miestas:
Latvija
Nuotolinis darbas:
Taip
Laikas:
Visa darbo diena
Galioja iki:
01/11/2026

Kandidatavimas vyks EPAM SISTEMOS, UAB įmonės puslapyje