Staff Platform Engineer
As a Staff Software Engineer / Staff Platform Engineer, you'll join our platform area, where we build the foundations every product team relies on to ship quickly and run safely in production. You'll work alongside a team of strong senior engineers, setting technical direction, raising the reliability bar, and helping the whole team sharpen its skills — when they hit a wall, they'll escalate to you.
What the role involves
Reliability & Operational Excellence
- Make operational processes — deployments, upgrades, migrations — routine, safe, and reversible
- Lead incident response end-to-end, owning postmortems and turning findings into lasting systemic improvements
- Automate repeated manual work, treating recurring tasks as issues to be resolved rather than accepted
Observability Depth
- Build monitoring and alerting that reflects customer impact, so any engineer can move from "something broke" to "here's why" in minutes
- Own the reliability and efficiency roadmap, identifying the biggest gaps and coordinating fixes across teams
Technical Leadership & Culture
- Set technical direction through design reviews, RFCs, and pairing, raising the bar for strong senior engineers
- Champion rigour and consistency across the platform, ensuring teams understand trade-offs, not just rules
- Run blameless postmortems that surface causes rather than blame, helping the team learn and improve each cycle
What we are looking for
- Extensive experience in SRE, software, and platform engineering, including operating, debugging, and reasoning about large-scale distributed systems in production
- Proven track record in incident leadership and reliability work, including running complex incidents, writing postmortems, and delivering systemic changes that stick
- Hands-on experience with observability stacks (metrics, logs, traces), including tools such as Prometheus/VictoriaMetrics, Loki, Grafana, and OpenTelemetry, alongside distributed tracing
- Sound judgement to identify when an architecture is unsuitable for a given workload and to define what should replace it
- Strong software engineering skills, ideally in Go, with deep experience in AWS and Kubernetes
- A mindset that treats repeated manual work as a problem to solve, with automation as second nature
- Ability to act as a technical reference for senior engineers and drive change across teams without formal authority, through design rigour and credibility
- Comfortable operating in ambiguous, high-impact environments, with the ability to plan a year or two ahead for platform needs
Nice to have
- Experience building or operating a large-scale time-series, logging, or continuous-profiling platform (e.g. VictoriaMetrics, Loki, Pyroscope, or similar)
- Strong understanding of the cost and performance trade-offs of data stores, and how these shape architectural decisions
What we offer
- Medical / Health Insurance
- Open Annual Leave
- Employee Assistance Programme
- Training & Learning Development
Incase you would like to apply to this job directly from the source, please click here