Site Reliability Engineer, Hardware
Job description:
As a Site Reliability Engineer on the Hardware team, you’ll be responsible for maintaining 10,000+ servers across Vinted’s data centers.
You’ll help monitor and debug production servers to ensure we meet our established SLOs. This will involve diving deep into hardware configuration, BIOS, BMC, and Linux systems to identify bugs, troubleshoot issues, and test performance. You’ll also automate server provisioning, build Grafana dashboards, and set up monitoring alerts.
You’ll work closely with data center administrators, who handle physical server repairs, helping them identify issues and validate fixes after repairs are completed. You’ll also collaborate with other teams to support them with hardware - and software-related challenges.
- Lead data center infrastructure reliability initiatives across 10,000+ physical servers.
- Build up automation tooling for server provisioning, configuration management, testing, and large-scale deployments.
- Create scripts, tools, and workflows to reduce manual operational work and improve infrastructure efficiency.
- Assist with troubleshooting hardware, BIOS, BMC, Redfish, Linux OS, networking, and performance-related issues.
- Communicate with infrastructure, platform, and engineering teams to support secure, scalable, and performant services.
- Participate in production incident response, root cause analysis, and on-call rotation.
- Monitor data center infrastructure, server health, SLOs, Grafana dashboards, alerts, and production systems.
- Oversee testing and benchmarking of new server hardware, CPU, and GPU models.
- Keep track of data center assets, hardware configurations, and infrastructure documentation in DCIM/NetBox.
- Work with the Platform Foundations DC team to create reliable, automated, and scalable infrastructure solutions.
- Maintain configuration-as-code repositories and automated tests.
- Improve server provisioning and configurations.
- Support other engineering teams who are running their services on servers.
- Collaborate with engineers to identify infrastructure bottlenecks and implement long-term improvements.
Requirements:
- Have experience with large-scale physical servers infrastructure and linux operating systems.
- Have a strong understanding of hardware automation and monitoring.
- Care deeply about scalability and production stability.
- Excellent written and spoken English
- Comfortable with production troubleshooting, debugging, and participation in an on-call rotation.
- Good at analyzing performance issues and benchmarking new hardware.
- Well-versed in Grafana and monitoring tools.
Nice to have
- Advantage: Knowledge of Chef
- Advantage: Experience in/with Netbox.
Company offers:
- The opportunity to benefit from our share options program
- 25 working days of holiday
- Access to all the tools & tech needed for work
- Home office support: we provide IT workstation equipment and a personal budget of up to 540 for home workplace furniture
- Private health insurance
- Mental and emotional health support through the Mindletic app
- Frequent team-building events
- A personal monthly budget for shopping on Vinted
- The opportunity to spend up to 90 days per year - 21 of which can be spent working outside of the EU - on workation
- A dog-friendly office
- In Vilnius office: Gym & in-house meals at friendly prices
- In Kaunas office: a monthly lunch allowance, and a once-a-week provided in-house lunch and breakfast