Senior Systems Administrator (Contract) at Digital Research Alliance of Canada
, , -
Full Time


Start Date

Immediate

Expiry Date

16 Sep, 26

Salary

117990.0

Posted On

18 Jun, 26

Experience

5 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Kubernetes, Ceph, Linux Systems Administration, Infrastructure as Code, Ansible, Terraform, Git, Distributed Systems Architecture, Network Security, Virtualization, Object Storage, High Performance Computing, Incident Management, Capacity Planning, Technical Leadership, Scripting Languages

Industry

Research Services

Description
  SENIOR SYSTEMS ADMINISTRATOR (CONTRACT) ABOUT THE ALLIANCE   The Digital Research Alliance of Canada (the Alliance) serves Canadian researchers, with the objective of advancing Canada’s position as a leader in the knowledge economy on the international stage. By integrating, championing and funding the infrastructure and activities required for advanced research computing (ARC), research data management (RDM), and research software (RS), we provide the platform for the research community to access tools and services faster than ever before.    We have an ambitious mandate – to transform how research across all academic disciplines is organized, managed, stored and used. We work with other ecosystem partners and stakeholders across the country to help provide Canadian researchers with the support they need for leading-edge research excellence, research, innovation and advancement across all disciplines.   POSITION SUMMARY The Senior Systems Administrator provides senior technical leadership for mission-critical infrastructure platforms that support the Distributed Storage and Compute Grid (DSCG) and the Alliance’s broader digital research infrastructure ecosystem. The role is responsible for the implementation, evolution, operational stewardship, and long-term sustainability of a core infrastructure platform, including container-based compute services and distributed object and file storage systems. Acting as a technical authority within the Grid Operations team, the position is accountable for designing implementing, and sustaining scalable, secure, and reliable infrastructure platforms that support nationally distributed services operating continuously across multiple sites. The role develops technical standards, drives automation and operational excellence, and collaborates with infrastructure operations, security, architecture, research data management, user support teams, vendors, host sites, and national partners to resolve complex technical challenges, coordinate platform changes, and support service delivery. Operating in a highly collaborative and multi-stakeholder environment, the position requires advanced technical expertise, strong systems thinking, sound judgement, and the ability to influence technical decisions across organizational boundaries. Reporting to the Director, Infrastructure Operations and Support, the role works closely with the Grid Operations Lead and other technical leaders across the Alliance and national digital research infrastructure ecosystem. This is a contract position with a term until March 31, 2028.   RESPONSIBILITIES Platform Leadership and Technical Stewardship * Define, implement, test, and support interoperability between infrastructure platforms and Alliance services. * Serve as a senior technical authority for a core infrastructure platform by contributing to platform architecture, technical standards, operational practices, and long-term roadmaps. * Optimize platform performance, reliability, scalability, and cost efficiency by managing containerized workloads, distributed computing environments, data services, and highly available production infrastructure. * Design, deploy, maintain, and continuously improve infrastructure supporting advanced research computing workloads, including compute clusters, storage systems, networking, monitoring, security, and platform services. * Evaluate emerging technologies and recommend improvements to enhance platform capabilities and service delivery. * Serve as a senior technical authority for a core infrastructure platform by helping design and evolve platform architecture, technical standards, operational practices, documentation, and long-term roadmaps. Automation and Platform Engineering * Develop and maintain automation, infrastructure-as-code, and platform lifecycle management solutions. * Implement tooling to support deployment, configuration management, monitoring, observability, and operational efficiency. * Improve consistency, repeatability, and reliability through automation and operational standardization. * Support ongoing modernization and optimization of platform management practices. Operational Excellence, Security, and Incident Management * Lead complex operational activities including platform upgrades, infrastructure changes, migrations, and service transitions. * Provide senior technical leadership during major incidents and service recovery efforts. Collaboration and Service Integration * Collaborate with infrastructure operations, user support, security, architecture, and research data management teams to support service delivery and ensure alignment with organizational standards and objectives. * Act as a senior escalation point for complex platform and infrastructure issues. * Lead cross-functional technical initiatives and manage dependencies impacting platform delivery, operational effectiveness, and service integration.  * Mentor junior technical staff and contribute to knowledge sharing and operational maturity. Partner and Stakeholder Engagement * Collaborate with host site teams, research institutions, infrastructure providers, vendors, and national partners to support distributed service delivery. * Foster effective partnerships across multiple organizations to support platform operations, service integration, and infrastructure initiatives. * Represent the Alliance in technical discussions with external stakeholders and contribute to the development of shared operational practices and service objectives. Strategic Planning and Service Development * Contribute technical expertise to infrastructure planning, service roadmaps, funding proposals, and investment decisions. * Support capacity planning, technology selection, and long-term platform sustainability initiatives. * Provide recommendations regarding infrastructure growth, modernization, risk mitigation, and operational improvements. * Contribute to the evolution of national digital research infrastructure capabilities and services.   QUALIFICATIONS  * Post-secondary degree in Computer Science, Computational Science, Information Technology, Engineering, or a related discipline. * 7 to 10 years of progressive experience designing, implementing, operating, and evolving complex infrastructure platforms in large-scale or distributed environments. * Demonstrated expertise with Kubernetes, Ceph, or comparable enterprise-scale infrastructure platforms. * Strong understanding of Linux systems administration, networking, storage technologies, virtualization, and distributed systems architecture. * Experience supporting research computing, digital research infrastructure, higher education, public sector, or large-scale scientific computing environments is considered an asset. * Experience with distributed storage systems, object storage technologies, high-performance computing environments, or national-scale infrastructure platforms is considered an asset. * Experience developing and maintaining automation, infrastructure-as-code, and configuration management solutions using tools such as Ansible, Terraform, Git, scripting languages, or equivalent technologies. * Experience supporting highly available, mission-critical services operating in 24x7 production environments. * Strong knowledge of infrastructure security principles, vulnerability management, incident response, and operational resilience. * Experience diagnosing and resolving complex technical issues involving multiple systems, platforms, and stakeholders. * Demonstrated ability to lead technical initiatives and coordinate work across teams without formal supervisory authority. * Strong analytical, troubleshooting, and problem-solving skills with the ability to balance operational, technical, security, and stakeholder considerations. * Demonstrated ability to influence, negotiate, and facilitate technical discussions, build consensus, and work effectively across teams with differing priorities and perspectives. * Experience working in distributed, multi-stakeholder environments involving multiple organizations, institutions, or service providers. * Strong verbal and written communication skills with the ability to communicate effectively with technical and non-technical audiences. The Alliance is strongly committed to equity and inclusion within the community and encourages applications from all qualified candidates, including women, members of racialized groups, people of colour, persons with disabilities, and Indigenous and 2SLGBTQIA+ identified people.   The expected salary range for this position for candidates residing in Canada is between $78,660 CAD - $117,990 CAD. Placement within this range will be determined based on several factors, including a candidate’s qualifications, skills, experience, demonstrated performance, overall alignment with the requirements of the role, and internal equity considerations, and other relevant organizational factors. The range reflects the Alliance’s commitment to equitable pay practices and to ensuring fair compensation for all employees. Please apply here: Careers at the Alliance! [https://workforcenow.adp.com/mascsr/default/mdf/recruitment/recruitment.html?cid=61008c6c-69b0-4ac1-8727-d86f6a6d54f5&ccId=19000101_000001&type=JS⟨=en_CA]
Responsibilities
Provide senior technical leadership for mission-critical infrastructure platforms supporting the Distributed Storage and Compute Grid. Responsible for designing, implementing, and sustaining scalable, secure, and reliable container-based compute and distributed storage systems.
Loading...