What you'll be doing
AI & High-Performance Network Architecture
- Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale.
- Design networks that integrate closely with bare-metal provisioning and cluster management systems.
- Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS.
- Establish reference architectures and engineering standards that can be implemented consistently across Nscale's data centre estate.
- Identify systemic risks and architectural gaps and drive durable solutions that improve scalability, reliability, and operational simplicity.
Network Automation & Infrastructure as Code
- Lead Nscale's network automation strategy using a GitOps operating model.
- Build and guide Python and Ansible tooling for provisioning, configuration validation, compliance, and operational workflows.
- Drive version-controlled configuration and CI/CD-based network change across multi-vendor environments.
- Apply Infrastructure-as-Code and Network-as-Code principles to reduce manual intervention and improve operational consistency.
- Continuously identify opportunities to automate repetitive operational tasks and reduce reactive toil.
Network Security & Edge Infrastructure
- Design and engineer perimeter and network security infrastructure across WAN and data centre edge environments.
- Own architecture across firewalls, NAT, VPN, security policies, and multi-tenant segmentation.
- Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure.
Reliability, Observability & Operations
- Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues.
- Establish measurable SLOs and operational standards for network services.
- Set technical direction for network observability, telemetry, monitoring, and alerting.
- Ensure clear visibility into fabric health, traffic patterns, performance, and capacity.
- Develop runbooks, automation, and engineering improvements that systematically reduce operational toil.
- Act as a senior 3rd/4th line escalation point for complex networking issues.