Reliability & Observability Analyst I at IREN

Sydney, New South Wales, Australia -

Full Time

Start Date

Immediate

Expiry Date

07 Aug, 26

Salary

0.0

Posted On

09 May, 26

Experience

2 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

Skills

Reliability Analysis, Observability, Automation, Linux, Networking, SRE Concepts, Splunk, Datadog, Prometheus, Python, Bash, AIOps, Incident Management, Data Analysis, KPI Reporting, Technical Support

Industry

technology;Information and Internet

Description

Job Type: Full-time | Location: Sydney |Work Location Type: #onsite IREN is a leading AI Cloud Service Provider, delivering large-scale GPU clusters for AI training and inference. IREN’s vertically integrated platform is underpinned by its expansive portfolio of grid-connected land and data centers in renewable-rich regions across the U.S. and Canada. With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever-evolving applications of high-performance compute. We believe that human progress is invaluable, but it should be done in the right way – responsibly, sustainably and having a positive impact on the communities we operate in. We are seeking an IOC Reliability & Observability Analyst I with a strong reliability, observability, and automation mindset to support our 24/7 HPC Data Center Operations. The role focuses on analyzing operational signals, improving incident quality, and supporting AIOps enabled automation and tooling and is designed for candidates early in their careers who want to grow into Site Reliability, Infrastructure Operations, or Platform Engineering paths. This is an entry‑level (Level 1) IOC role focused on operational analysis, data quality, and reliability signal validation rather than system design or engineering ownership. You will support IOC, engineering, and operations teams by analyzing incidents, validating operational signals, and identifying opportunities to improve detection quality and operational reliability under established processes and guidance. * 1-3 years of experience in IOC, NOC, SRE‑adjacent operations, systems analysis, or technical support roles * Bachelor's degree in Computer Science, Data Science, Statistics, or equivalent hands-on experience * Exposure to 24/7 production environments supporting infrastructure, cloud, or data center operations * Foundational awareness of SRE concepts such as service health, MTTR/MTTD, and the incident lifecycle, with the ability to apply these concepts in operational analysis. * Working knowledge of Linux-based systems, basic networking concepts, and infrastructure dependencies * Experience working with metrics, logs, and alerting systems across infrastructure or application environments * Familiarity with observability platforms (e.g., Splunk, Datadog, Prometheus-style metrics) * Ability to assess alert quality, identify noise, and recognize monitoring gaps * Awareness of AIOps concepts such as anomaly detection, event correlation, and alert noise reduction, primarily for the purpose of reviewing and validating automated insights * Experience validating automated insights and supporting alerting or observability automation * Ability to read automation artifacts (Python, Bash, or configuration-based workflows) and assist with minor updates under documented procedures and guidance * Ability to analyze incident trends and system behaviors with strong attention to data accuracy, signal integrity, and identify recurring issues or improvement opportunities * Clear communication skills and comfort working cross-functionally with operations and engineering teams Other important requirements * This role operates in a 24×7 IOC/NOC environment and works 12‑hour rotating shifts on a 4‑days‑on / 3‑days‑off, alternating with 3‑days‑on / 4‑days‑off schedule * Pre-employment screening, including background check and substance testing may be required according to company policies * Analyze incident data, system behaviors, and operational signals across GPU clusters, networks, and facilities to identify risks and trends * Identify detection gaps, alert delays, false positives, and under-monitored systems, and document findings for review by IOC leadership or engineering teams * Validate ticketing and incident data for accuracy, completeness, and reporting integrity * Support continuous improvement of observability by evaluating metrics, logs, alerts, and dashboards * Assist in refining operational views focused on service health, reliability, and signal quality * Generate post-incident insights highlighting trends, risks, and improvement opportunities * Support AIOps-enabled capabilities by reviewing outputs from anomaly detection, alert correlation, and event clustering, and flagging accuracy or data-quality issues * Validate automated insights and escalate tuning or accuracy issues to IOC and engineering teams * Assist with testing automation related to alert routing, enrichment, and suppression, and submit recommended changes through established change and review processes * Produce and maintain SLA/KPI dashboards and reliability reports using established templates, definitions, and data sources * Provide data-driven insights and recommendations to inform preventive measures, workflow improvements, and monitoring enhancements * Contribute to runbook updates, operational documentation, and reliability initiatives in partnership with IOC and engineering teams * Develop foundational SRE skills in preparation for expanded operational responsibility * This role operates under defined IOC processes and supervision, with increasing responsibility as skills and experience develop At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader total rewards package, thoughtfully designed to support your health, well-being, and long-term success. * Compensation & Rewards * Competitive salary range finalized based on experience and impact * Short and long-term incentive programs designed to reward both results and long term company success * Wellbeing & Benefits * Paid vacation to recharge, travel, or simply enjoy more life outside of work We value diverse perspectives and believe that skills can be developed. If you’re passionate about this role, we want to hear from you — whether you meet every criteria or not. Your unique experiences might be exactly what we need! IREN Limited is an equal opportunity employer that is committed to creating an inclusive workplace. We evaluate qualified applicants without regard to race, colour, religion, age, sex, sexual orientation, gender identity, genetic information, national origin, disability, veteran status, and other legally protected characteristics.  This job will remain posted until filled. While we appreciate all applications we receive, we are only able to contact candidates under consideration.  By applying for this position and submitting your resume and application materials, you consent to the processing of your personal information in accordance with our Job Applicant Privacy Statement available on our website at www.iren.com [http://www.iren.com/].

Responsibilities

Analyze operational signals and incident data across GPU clusters and networks to identify risks and improve detection quality. Support AIOps automation and maintain reliability reports and SLA/KPI dashboards to enhance system health.