Systems Integration and Debug Lead at Advanced Micro Devices, Inc
Bengaluru, karnataka, India -
Full Time


Start Date

Immediate

Expiry Date

17 Sep, 26

Salary

0.0

Posted On

19 Jun, 26

Experience

10 year(s) or above

Remote Job

Yes

Telecommute

Yes

Sponsor Visa

No

Skills

System Level Debug, RAS Knowledge, Hardware Debugging, Firmware Debugging, GPU Architecture, PCIe, HBM, Direct Liquid Cooling, Cluster Networks, Electrical Probing, Root Cause Analysis, Technical Leadership, System Initialization, SoC Debug, Oscilloscopes, Board Level Power Analysis

Industry

Semiconductor Manufacturing

Description
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. SMTS SILICON DESIGN ENGINEER THE ROLE: The Datacenter Platform Engineering Group (DPEG) organization is looking for an experienced system level debug engineer. Individual will be part of a team that will be responsible for Data Center deployments as well as ensuring availability/uptime of a large number of systems. Individual should be familiar with System level debug at all levels as well have RAS knowledge to quickly root cause issues. Person should have not only strong technical skills but be able to direct junior level engineers to solving problems and guide them to be successful in handling service level tickets. THE PERSON: Experience in debugging of complex HW/FW issues is a must, understand the flow of a GPU through the different layers of a system and be able to validate the items connecting to the GPU SOC (pcie, vr’s, RMs, retimers, HBM, internal networking). Handling of Data Center related issues that come with electrical complexities, direct liquid cooling and complex cluster networks. Communication Is essential in working with different owners of the functional code stack as well as the ability to drive issues via phone calls, chat messages, e-mails. Hands on experience with Hardware in a DataCenter environment will be required. KEY RESPONSIBLITIES: Debug / triage engineer and understanding of industry tools for root causing complex issues Understanding of GPU/System level HW and SW flow Ability to probe parts of a board; check electrical and power currents and validate a system RAS knowledge and flow of issues seen at a GPU level Provide leadership for driving to root cause issues and guide junior level engineers Document flows and methods of bring-up, boot-up, system initialization and debug PREFERRED EXPERIENCE: Minimum 10 yrs experience in Systems Integration, Debug Proven ability to drive resolution of critical problems within a lab, Datacenter Relationship with external customers/partners and able to help resolve problems in a Data Center Relationship with external customers/partners on ability to work manufacturing issues/failures Relationship with external customers/partners on ability to define rqmts for deployments/debug IDEAL CANDIDATE: Significant experience in SoC and/or System debug of complex issues Develop / Document debug capabilities on a given SOC and System Go-to-person at a Data Center for resolving issues, looking at complex issues Collaborate with internal teams on root causing issues, finding optimum resolutions Hands-on experience in using industry debug tools, scopes as well examine board level power ACADEMIC CREDENTIALS: Bachelors or Masters degree in Computer Engineering/Electrical Engineering Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Responsibilities
Lead the debug and triage of complex hardware and firmware issues within Data Center deployments to ensure system availability and uptime. Provide technical leadership and guidance to junior engineers while documenting bring-up and initialization flows.
Loading...