Solving complexity. Accelerating results.

At Penguin Solutions, we understand the boundless potential of technology and support our customers in turning cutting-edge ideas into outcomes—faster, and at any scale.

With over two decades of experience as trusted advisors, Penguin Solutions is an end-to-end technology company solving the industry’s most complex challenges in computing, memory, and LED solutions. Penguin designs, builds, deploys, and manages high-performance, high-availability enterprise solutions, allowing customers to achieve their breakthrough innovations.

Solving complexity. Accelerating results.

At Penguin Solutions, we understand the boundless potential of technology and support our customers in turning cutting-edge ideas into outcomes—faster, and at any scale.

With over two decades of experience as trusted advisors, Penguin Solutions is an end-to-end technology company solving the industry’s most complex challenges in computing, memory, and LED solutions. Penguin designs, builds, deploys, and manages high-performance, high-availability enterprise solutions, allowing customers to achieve their breakthrough innovations.

Sr. Failure Analysis Engineer

Date Posted:  Jul 24, 2026
Requisition ID:  2051
Location: 

Newark, CA, US, 94560

Brand:  Penguin Solutions

Senior Failure Analysis Engineer

 

At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.

 

Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.

 

Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.

 

At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.

 

Overview

We are seeking a seasoned and highly skilled Senior DRAM Failure Analysis Engineer with extensive expertise in high-speed, high-capacity DRAM modules. The ideal candidate will have 5+ years of experience in DRAM module reliability testing, high-speed signal integrity analysis, and failure analysis to ensure optimal performance in mission-critical environments. This leadership role involves owning and advancing our burn-in methodologies, spearheading high-speed testing strategies, and driving system-level failure analysis to guarantee the reliability of DDR4/DDR5 based memory solutions. You will work closely with cross-functional teams to define and enhance product performance, quality, and reliability, while also mentoring junior engineers.

 

Responsibilities

  • Lead and conduct complex failure analysis (FA) on DRAM modules and memory subsystems, utilizing high-speed signal integrity tools and oscilloscopes.
  • Drive the determination of failure root causes at the chip, module, and system level, presenting findings to technical and leadership teams.
  • Architect and execute debug strategies for system-level failures by analyzing memory controller interactions, BIOS tuning, and DIMM register settings in server and cloud environments.
  • Investigate and resolve critical performance bottlenecks, intermittent failures, and memory errors caused by power integrity (PI), signal integrity (SI), and thermal stress.
  • Lead the development and implementation of advanced stress test strategies for high-performance DRAM modules.
  • Mentor junior engineers in best practices for high-frequency waveform analysis, signal integrity debugging, and root cause analysis.
  • Generate and present detailed reports (8D) summarizing failure analysis findings, root causes, and strategic recommendations for product and process improvements.
  • Act as a technical lead in cross-functional teams to enhance product design, quality, and manufacturability.

 

Qualifications

  • Bachelor’s degree in Electrical Engineering or a related field; Master's degree is a plus.
  • 5+ years of experience in DRAM module burn-in, stress testing, and failure analysis.
  • Expert-level understanding of high-speed memory interfaces (DDR4, DDR5, HBM) and advanced SI/PI concepts.
  • Proven track record of complex problem-solving and root cause failure analysis in a high-performance computing environment.
  • Demonstrated mastery of server memory modules and system-level debugging in Linux and Windows environments.
  • Expertise with high-frequency test equipment such as oscilloscopes.
  • AI & Automation Fluency: Strong proficiency in applying modern generative AI tools and Agentic frameworks to real-world business problems.
  • Advanced proficiency in data analysis and scripting (e.g., Python, Perl, or similar).
  • Deep knowledge of JEDEC reliability standards, ECC error handling, and memory RAS (Reliability, Availability, and Serviceability) features.
  • Extensive experience with BIOS tuning, memory controller optimizations, and DRAM memory training algorithms.
  • Excellent communication skills with the ability to explain complex technical concepts effectively to both technical and non-technical audiences.
  • Strong interpersonal and leadership skills for mentoring and driving cross-functional team interaction.

 

Location

This opportunity is in Newark, California.

 

Travel

Willingness to travel occasionally to customer sites and industry events as needed.

 

Compensation & Benefits

The base pay range that the Company reasonably expects to pay for this position in California is $145,000 - $165,000; the pay ultimately offered may vary based on business considerations, including job-related knowledge, skills, experience, and education. The position is bonus-eligible, and there are medical, dental, and vision benefits available. There is a 401k saving plan and other benefits, such as Paid Time Off, Life Insurance, and an Employee Assistance Plan.   

 

Inclusion & Belonging Statement

We are committed to creating an inclusive environment that embraces differences and fosters belonging for all.

 

Equal Opportunity Statement

We are an Affirmative Action/Equal Opportunity Employer and strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, national origin, race, ethnicity, creed, gender, disability, veteran status, or any other characteristic protected by law.

 


Nearest Major Market: San Francisco
Nearest Secondary Market: Oakland