Senior Site Reliability Engineer II, Observability Infrastructure
Job description:
As a Senior SRE Engineer II in the Observability Infra team, you’ll build and operate the software and infrastructure behind Vinted’s observability platforms. These platforms process metrics, logs, and traces from more than 7,000 physical servers and more than 30 kubernetes clusters across multiple regions.
You’ll develop backend services,libraries, APIs, automation, and self-service capabilities, primarily using Go and Kubernetes. You’ll take ownership of solutions throughout their lifecycle: from technical design and implementation to deployment and production operation and help shape the team’s technical direction.
This role is suited to an experienced backend engineer with hands-on Kubernetes experience who wants to work on large-scale infrastructure. Previous infrastructure operations or site reliability experience is valuable, but you can develop this expertise in the role if you have strong software engineering fundamentals and are motivated to take operational ownership.
In this position, you’ll
- Design and develop backend services, APIs, and automation for Vinted’s observability platforms.
- Build and operate reliable, scalable services and stateful workloads on Kubernetes.
- Develop self-service capabilities for engineering teams using our metrics, logs, and tracing products.
- Evolve our Prometheus, Thanos, OpenTelemetry, and Elasticsearch-based platforms.
- Lead technical initiatives from design to production, solving complex software and infrastructure problems along the way.
- Participate in the on-call rotation for observability products, currently approximately once every five to seven weeks.
Our technologies include Go, Kubernetes, OpenTelemetry, Prometheus, Thanos, Elasticsearch, Opensearch, Doris, Vector, Jaeger, Grafana, ArgoCD, Istio, Kafka, Redis and Linux.
Requirements:
- Have strong backend software engineering and system design skills, ideally with production experience in Go; experience with other backend languages is also welcome.
- Have hands-on experience deploying, operating, and troubleshooting services on Kubernetes.
- Can design maintainable distributed systems and own them throughout their production lifecycle.
- Understand reliability, scalability, monitoring, safe deployments, and failure recovery.
- Have infrastructure operations or site reliability experience, or are motivated to develop these skills and participate in on-call.
- Communicate technical decisions clearly, collaborate effectively across teams, and have excellent written and spoken English.
Nice to have
- Experience with OpenTelemetry, Prometheus, Thanos, Elasticsearch, or other large-scale observability technologies is an advantage.
Company offers:
- The opportunity to benefit from our share options programme
- 25 working days of holiday
- Access to all the tools & tech needed for work
- Home office support: we provide IT workstation equipment and a personal budget of up to 540 for home workplace furniture
- Private health insurance
- Mental and emotional health support through the Mindletic app
- Frequent team-building events
- A personal monthly budget for shopping on Vinted
- The opportunity to spend up to 90 days per year - 21 of which can be spent working outside of the EU - on workation
- A dog-friendly office
- In Vilnius office: Gym & in-house meals at friendly prices
- In Kaunas office: a monthly lunch allowance, and a once-a-week provided in-house lunch and breakfast