bout The Opportunity
We're building a highly resilient hybrid multicloud platform spanning Nutanix and public cloud, while delivering a major migration from VMware. We're looking for a Platform Reliability Engineer who is passionate about automation, resilience, recoverability, and operational excellence. This is not a traditional infrastructure operations role. We want someone who can demonstrate how they have improved platform reliability through engineering, automation, monitoring, disaster recovery testing, and continuous improvement.
What You'll Be Doing
- Ensuring the reliability, performance, and recoverability of critical platform services
- Managing and validating backup and recovery capabilities using enterprise tools such as Rubrik
- Planning and executing disaster recovery testing and recovery assurance activities
- Supporting platform and workload migrations from VMware to modern cloud platforms
- Developing automation to reduce operational effort and improve service reliability
- Improving observability, monitoring, alerting, and operational readiness across the platform
What We're Looking For
- 8 to 10 plus years of strong experience with enterprise infrastructure, cloud, or platform engineering
- Experience with Nutanix, VMware, public cloud platforms, and backup/recovery technologies
- Strong PowerShell, Python, or similar automation skills
- Experience improving platform resilience, recoverability, and operational efficiency through automation
- Proven ability to reduce operational toil and improve reliability at scale
For this role, you must be able to demonstrate:
- Reliability improvements you've delivered and how success was measured
- Disaster recovery programmes or recovery testing you've led
- Operational processes you've automated and the impact achieved
- How you validate recoverability, not just successful backups
- Major migration or transformation programmes you've supported
Incase you would like to apply to this job directly from the source, please click here