Grįžti į skelbimus
Vinted, UAB

Lead Site Reliability Engineer, ML Platform

13
6767 - 9150 €/mėn.
Prieš mokesčius

Job description:

As Lead Site Reliability Engineer, you’ll own the reliability strategy and architectural evolution of Vinted’s ML platform. Your mission is to ensure that every machine learning model—from classic algorithms to Large Language Models (LLMs)—is trained and served with industry-leading stability and performance.

You’ll join the ML Platform team, which owns the tooling for ML/LLM development, deployment and other platform capabilities that increase ML/AI delivery speed at Vinted. You won’t just be managing services; you’ll be shaping the architectural strategy for how Vinted serves AI at scale, and you'll act as a multiplier for the engineering function by mentoring senior engineers and driving domain-wide standards.

In this role, you’ll partner closely with Data Scientists and Platform teams to translate Vinted’s ML innovation into production reality. You will be responsible for orchestrating one of the region’s largest GPU fleets—over 200 high-performance GPUs dedicated to ML model training, production inference and LLM serving, including next-generation hardware designed for the most demanding AI workloads.

Our tech stack: Kubernetes, Terraform, Chef, Google Cloud Platform, Go, Kafka, Vespa, Redis, Vitess.

  • Lead the technical roadmap for the ML Platform, ensuring our training and inference infrastructure scales to meet Vinted’s growing ML/AI demands.
  • Architect high-performance inference solutions, focusing on low-latency delivery and efficient GPU utilisation across our clusters.
  • Define and drive reliability standards, including SLOs, error budgets, and sophisticated observability for model-serving systems.
  • Mentor and grow senior engineers within the team, fostering a culture of technical excellence, pragmatism, and continuous learning.
  • Identify and solve systemic bottlenecks in our deployment lifecycle by writing RFCs and sparring with colleagues on long-term architectural improvements.
  • Oversee the integration of deployment workflows with automated experimentation patterns, such as A/B testing and shadow deployments.
  • Communicate with stakeholders about architectural risks, planned changes, and strategic reliability trade-offs.

Requirements:

  • Have a track record of leading technical problem-solving for complex, domain-level systems.
  • Deeply experienced with Kubernetes, Linux internals, and high-scale networking.
  • Have working experience with programming or scripting (Go is a plus, but experience with other languages is also welcome).
  • Choose the right tool for the job based on long-term impact rather than personal preference.
  • Excellent written and spoken English, with the ability to articulate technical risks to both technical and non-technical stakeholders.
  • Natural coach and mentor who finds more value in enabling an entire team than in writing every line of code yourself.

Nice to have

  • Advantage: Experience with GPU workloads and inference frameworks like KServe, vLLM, or NVIDIA Triton.

Company offers:

  • The opportunity to benefit from our share options programe
  • 25 working days of holiday
  • Access to all the tools & tech needed for work
  • Home office support: we provide IT workstation equipment and a personal budget of up to 540 for home workplace furniture
  • Private health insurance
  • Mental and emotional health support through the Mindletic app
  • Frequent team-building events
  • A personal monthly budget for shopping on Vinted
  • The opportunity to spend up to 90 days per year - 21 of which can be spent working outside of the EU - on workation
  • A dog-friendly office
  • In Vilnius office: Gym & in-house meals at friendly prices
  • In Kaunas office: a monthly lunch allowance, and a once-a-week provided in-house lunch and breakfast

Miestas:
Vilnius
Nuotolinis darbas:
Ne
Laikas:
Visa darbo diena
Galioja iki:
23/10/2026

Kandidatavimas vyks Vinted, UAB įmonės puslapyje