About the job
Company Overview
Open Innovation AI is a global technology company that specializes in developing advanced solutions for managing AI workloads. Its flagship product, the Open Innovation Cluster Manager (OICM), orchestrates complex AI tasks efficiently across diverse infrastructures. The platform is hardware-agnostic, optimized for various GPUs and accelerators hardware, and facilitates seamless integration and scalability for enterprise AI applications. Open Innovation AI focuses on optimizing and simplifying AI workload management and making AI technologies accessible to organizations of all sizes. With its innovative solutions, companies can reduce operational costs, accelerate time to value, and maximize their return on investment, ensuring that their AI strategies contribute directly to enhanced business outcomes.
Role Overview:
The SRE Manager leads Open Innovation AI's L2 Support - Site Reliability Engineering team, a multidisciplinary function responsible for the reliable operation of GPU-dense, on-premises AI platforms across compute, storage, networking, virtualization, and Kubernetes. The role owns the people, processes, technical coordination, service performance, and continuous improvement required to deliver reliable L2 support across customer environments. This is a management and operational-leadership role. The holder must be technically credible enough to understand complex incidents, assess priority and risk, guide recovery, and make sound assignment and escalation decisions, without needing to be the deepest hands-on specialist in every domain. The role is accountable for SLA performance, 24/7 coverage, ITIL-aligned practices, operational readiness, reliability improvement, and the development of the engineering team, working closely with the Service Desk, Service Delivery Managers, L3, product teams, delivery teams, customers, and technology partners
Role Responsibilities:
- Manage the L2 Support - Site Reliability Engineering team day to day, including workload allocation across customers, environments, services, and issues, and rebalance assignments as priorities and operational risks change.
- Own L2 service performance and maturity, including SLA compliance, operational KPIs, backlog health, ticket ageing, repeat incidents, escalation quality, and customer-specific support readiness.
- Coordinate the team's incident response so that every incident has the right technical owner, is tracked through recovery and resolution, and is escalated promptly when deeper expertise or additional authority is required.
- Act as, or appoint, the technical incident lead for P1 and P2 incidents, ensuring clear technical ownership, coordinated recovery, timely stakeholder updates, evidence preservation, and post-incident follow-up.
- Plan and own on-call rotations and shift coverage to provide sustainable 24/7 continuity for key accounts
- Provide technical oversight and judgement on complex incidents by understanding the problem, assessing risk and options, and deciding on priority, assignment, recovery approach, and escalation, while relying on senior engineers for deep hands-on troubleshooting.
- Ensure consistent, ITIL-aligned Incident, Problem, and Change Management across the L2 function, including change risk assessment, execution readiness, rollback planning, and follow-up of corrective actions.
- Support Service Delivery Managers in customer operational reviews, escalations, SLA analysis, service-improvement plans, and communication of technical risks, while maintaining clear boundaries between technical operations and commercial account ownership.
- Coordinate closely with the Service Desk for smooth ticket flow and accurate triage, and with L3, product, and delivery teams on root-cause analysis, fix validation, knowledge transfer, and operational readiness.
- Own the L2 operational-readiness assessment for new releases, platforms, and customer environments, including monitoring, access, documentation, runbooks, training, escalation paths, recovery procedures, rollback plans, and known limitations.
- Oversee operational health through regular review of system status, cluster integrity, capacity, utilization, availability, performance, platform behavior, and emerging risks across supported environments.
- Coordinate technical escalations with vendors, delivery partners, customer infrastructure teams, and other third parties when incidents or risks cross organizational boundaries.
- Ensure that the team maintains accurate and usable SOPs, runbooks, troubleshooting guides, knowledge-base articles, architecture references, support matrices, and shift-handover records.
- Lead all people-management activities for the team, including hiring, onboarding, objective setting, performance management, coaching, succession planning, skills development, and building sufficient technical depth and redundancy.
- Prepare operational reports and P1/P2 post-incident reports with clear root-cause analysis, customer impact, timeline, corrective actions, owners, and due dates.
- Keep the Head of Technical Operations and relevant stakeholders informed about service performance, team capacity, operational events, material risks, dependencies, and continuous-improvement initiatives.
- Exercise the authority required to meet operational commitments, including reprioritizing work, reassigning engineers, escalating resource conflicts, requiring incident reviews, recommending emergency changes, and rejecting operationally unready service handovers
Required experience & Qualification
- Bachelor’s degree in computer science, Information Technology, Engineering, or a related field.
- 8 or more years of experience in L2/L3 support, SRE, systems engineering, or infrastructure operations, including large-scale on-premises production environments, with at least 3 years of direct people-management responsibility.
- Proven people-management experience covering hiring, onboarding, objective setting, performance management, coaching, career development, and management of underperformance.
- Experience managing customer-facing production services in a multi-customer environment and balancing operational priorities, service commitments, risk, and available engineering capacity.
- Demonstrated major-incident leadership, including technical coordination, recovery governance, stakeholder communication, escalation, and post-incident corrective-action management.
- Experience owning SLAs, operational KPIs, on-call coverage, backlog governance, service reviews, and continuous-improvement plans.
- Proven ability to build, scale, or mature an SRE, infrastructure-operations, or technical-support function, including establishing operating practices, technical ownership, skills coverage, and knowledge management.
- Technically credible across key SRE domains including GPU-dense compute, Linux, Kubernetes, high-performance networking over Ethernet, InfiniBand and RoCE, distributed storage, and virtualization. The candidate must be able to understand incidents, assess priority and risk, and make informed assignment and escalation decisions; deep hands-on expertise in every layer is not required.
- Working knowledge of distributed middleware and data-layer technologies such as Kafka, Redis, and PostgreSQL.
- Solid knowledge of ITIL-aligned Incident, Problem, and Change Management, including operational risk assessment and change-readiness governance.
- Strong coordination, decision-making, and communication skills, with the ability to maintain a clear view of ownership, priorities, dependencies, and risk and to communicate effectively with engineers, customers, partners, and senior leadership.
- Experience working in secure, isolated, air-gapped, or compliance-driven on-premises environments.
- Ability to obtain and maintain the clearance required for regular access to security-controlled customer sites.
- Fluent written and spoken English
Preferred Skills:
- Hands-on experience with Kubernetes and HPC or AI platform-management tooling.
- Experience coordinating technical escalations with hardware, storage, network, platform, or software vendors.
- Relevant certifications such as ITIL, CKA or CKAD, RHCE or RHCA, CCNP, or VMware VCP.
- Reporting To: Head of Technical Operations
- Manages: The L2 Support - Site Reliability Engineering team, spanning compute, storage, networking, virtualization, Kubernetes, platform operations, and reliability engineering.
Incase you would like to apply to this job directly from the source, please click here