Senior Site Reliability Engineer, ML Platform
Job description:
As Senior Site Reliability Engineer, you will be responsible for the reliability, scalability, observability, and operational excellence of Vinted’s ML Platform services across on-prem and cloud environments - including SLOs, incident response, capacity planning, automation, and participating in the on-call rotation for critical platform services. You’ll join the ML Platform team with a product mindset - delivering resilient, self-service platform capabilities for the Vinted Feature Store, embedding extraction pipelines, and dataset tooling so data science teams can move fast without operational friction. You will work closely with Data Infrastructure, Production Engineering, and ML teams across Vinted.
Here are some of the technologies we use: Kubernetes, Terraform, Chef, Google Cloud Platform, Go, Kafka, Vespa, Flink, Temporal, Vitess.
- Own the reliability and performance of the Vinted Feature Store, ensuring it meets strict latency SLOs at 400k+ feature-read RPS.
- Evaluate, benchmark, and prototype next-generation distributed storage technologies as potential replacements for Vespa to support long-term online store scalability.
- Build and automate reliable, self-service embedding extraction pipelines, abstracting compute and serving complexity for internal users.
- Scale and operate the Dataset & Drift Platform, enabling self-service dataset discovery, versioning, and drift monitoring across hybrid training workflows.
- Lead reliability work for ML Platform services (SLOs, alerting, runbooks, postmortems, toil reduction).
- Build automated data governance into the platform, ensuring feature and dataset lifecycles are GDPR-compliant by design.
- Treat internal ML teams as platform customers—driving ML data standards, improving developer experience, consulting on adoption, and navigating reliability trade-offs.
Requirements:
- Have a solid understanding of Linux operating systems and distributed systems architecture.
- Experience building and running internal platforms or high-throughput, low-latency distributed data systems (NoSQL, Key-Value stores, Vespa, Kafka, etc.).
- Proven track record with core SRE practices (SLOs, observability, incident management) paired with a platform mindset (self-service, automation, reducing developer cognitive load).
- Have experience in parts of our technology stack or have worked with similar tools and platforms.
- Strong programming or scripting skills with a desire to build platform services and tooling (Go is a plus, but experience with other languages is also welcome).
- Excellent written and spoken English, with great cross-team consulting and communication skills.
- You’re curious, motivated to grow, and eager to deepen your expertise in the concepts and technologies we use.
Nice to have
- Advantage: experience with feature stores, stream processing, workflow orchestration (Temporal), data drift tooling, or data compliance (GDPR).
Company offers:
- The opportunity to benefit from our share options programme
- 25 working days of holiday
- Access to all the tools & tech needed for work
- Home office support: we provide IT workstation equipment and a personal budget of up to 540 for home workplace furniture
- Private health insurance
- Mental and emotional health support through the Mindletic app
- Frequent team-building events
- A personal monthly budget for shopping on Vinted
- The opportunity to spend up to 90 days per year - 21 of which can be spent working outside of the EU - on workation
- A dog-friendly office
- In Vilnius office: Gym & in-house meals at friendly prices
- In Kaunas office: a monthly lunch allowance, and a once-a-week provided in-house lunch and breakfast