Design, implement, and maintain observability, monitoring, and synthetic testing to detect issues early and validate service behavior end-to-end.
Participate in on-call rotation and provide timely incident response through PagerDuty and related operational processes.
Lead or support disaster recovery planning, testing, and execution to ensure business continuity and recovery readiness.
Define, track, and continuously improve SLOs/SLIs to align platform reliability with business expectations.
Produce clear and actionable monthly operations reports summarizing service health, trends, incidents, and improvement opportunities.
Leverage Datadog to analyze logs, create and maintain monitoring metrics, and configure dashboards and alerts that support CloudOps activities and platform reliability.
Analyze operational data and platform metrics to identify performance bottlenecks, reliability gaps, and opportunities for automation or optimization.
Design, build, and maintain automation and internal tooling to improve efficiency, consistency, and scalability across the CloudOps environment.
Deploy, configure, and maintain application and platform layers of the service, ensuring changes are implemented safely and in accordance with standards.
Collaborate with architects, developers, and other technical stakeholders to review designs, provide feedback, and help shape scalable and resilient solutions.
Review peer code and infrastructure changes to ensure quality, maintainability, and compliance with team standards and best practices.
Develop, test, and document automation and platform components, including runbooks and operational procedures, to support long-term maintainability and operational readiness.