SRE Engineer- Sunnyvale, CA, the US at Kody
Sunnyvale, California, United States -
Full Time


Start Date

Immediate

Expiry Date

15 Sep, 26

Salary

0.0

Posted On

17 Jun, 26

Experience

5 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

AWS, GitHub Actions, Git Workflows, CI/CD, Incident Management, Observability, Monitoring, Log Management, Scripting, Mandarin Fluency, English Fluency, Production Reliability, SLA Management, Troubleshooting, Post-mortems, Platform Engineering

Industry

Financial Services

Description
About the Role We are seeking a high-caliber Senior Site Reliability Engineer (SRE) based in California to ensure the scalability, reliability, and runtime efficiency of our next-generation platform. In this role, you will bridge the gap between development and operations, working closely with our global engineering teams. We are looking for a unique engineering mindset: someone who brings a positive, collaborative energy to the daily grind, but can instantly pivot into a hyper-focused, high-ownership responder when an incident strikes. Key Responsibilities Production Reliability & Guardrails: Partner with the Platform Engineering team to implement reliability guardrails, ensuring applications running on AWS meet strict uptime and SLA requirements. CI/CD & Repository Management: Own the deployment pipelines and code management practices extensively via GitHub. Incident Management: Lead rapid-response troubleshooting during production incidents; conduct thorough blameless post-mortems to continuously harden our systems. Observability & Performance: Implement advanced monitoring, logging, and alerting systems to proactively detect and mitigate system anomalies. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment. 1. Technical Focus Ecosystem Expertise (Must-Haves): Deep, practical experience managing application deployment and runtime environments on AWS, alongside master-level knowledge of advanced Git workflows and actions on GitHub. Core Toolkit: Strong proficiency in monitoring tools, log management, and scripting for quick triaging and troubleshooting. 2. Soft Skills & Characteristics Ownership & Transparency: You are radically open, highly responsive, and communicative. You don't just clear tickets; you own the production environment's health end-to-end. Pressure-Resistance: High psychological resilience. You maintain a happy, positive attitude during smooth operations, yet feel a healthy, driving sense of urgency and laser-focus during high-stakes incidents. Bilingual Capability: Absolute fluency in Mandarin and English (verbal and written) is mandatory for effective technical alignment across our global teams. - Competitive base salary + equity packages aligned with California market standards.
Responsibilities
Ensure the scalability and reliability of a next-generation platform by implementing reliability guardrails and managing CI/CD pipelines. Lead incident response and troubleshooting while acting as a technical bridge between US and international engineering teams.
Loading...