Senior Principal Software Engineer, AI Infra Compute at Oracle Risk Management Services
Austin, Texas, United States -
Full Time


Start Date

Immediate

Expiry Date

11 Jun, 26

Salary

251600.0

Posted On

13 Mar, 26

Experience

10 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

Distributed Systems Engineering, Architectural Changes, GPU Delivery, Health Monitoring, Triage Automation, Diagnostic Services, AI/ML/HPC Workloads, RoCE, Infiniband, Monitoring Solutions, Repair Solutions, AI Infrastructure, GPU Control Plane, GPU Data Plane, Technical Leadership, Cross-functional Collaboration

Industry

IT Services and IT Consulting

Description
Our team is the GPU Availability and Monitoring team in the Compute Org. we are responsible for designing and developing architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services. These are essential for running distributed AI/ML/HPC workloads across thousands of GPUs, leveraging technologies like RoCE and Infiniband. We are looking for a highly skilled and motivated distributed systems engineer who can architect solutions to scale and optimize Monitoring and Repair solutions for AI infrastructure components like GPU control plane and GPU data plane that provide computing resources to customer AI workloads. You will provide technical leadership to the team and bring clarity to ambiguous problems and come up with innovative solutions. You will collaborate with cross-functional teams to enhance our AI infrastructure to deliver exceptional customer experience and peak performance.  Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com [accommodation-request_mb@oracle.com] or by calling 1-888-404-2494 in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

How To Apply:

Incase you would like to apply to this job directly from the source, please click here

Responsibilities
The role involves designing and developing architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services essential for running distributed AI/ML/HPC workloads. The engineer will architect solutions to scale and optimize Monitoring and Repair solutions for AI infrastructure components like the GPU control plane and data plane.
Loading...