Sr. Data Center Technician
GA, US
At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.
Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.
Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.
At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.
Job Overview
We are looking for an experienced Senior Onsite Data Center Professional to apply specialized technical expertise and advanced judgment in a fast-paced, complex AI/HPC environment. Rather than merely executing established procedures, this role demands independent analysis, complex problem-solving, and decisive action to address large-scale infrastructure challenges. You will take full ownership of technical solutions and strategic recommendations, serving as a primary authority for issue resolution rather than simply escalating problems.
In this senior capacity, you will drive process optimizations, provide critical technical guidance to cross-functional teams, and actively shape operational strategies to support cloud-scale compute and storage environments. Adaptability, advanced technical judgment, and the ability to influence outcomes are central to your success in this role.
This position will be onsite in Columbus, Georgia at the customer’s data center.
Responsibilities
- Advanced Diagnostics & Problem Solving: Exercise independent technical judgment to lead advanced hardware diagnostics, isolate root causes, and own the resolution of complex AI/HPC issues (e.g., GPU kernel hangs, interconnect anomalies) ensuring optimal cluster stability.
- Technical Authority & Ownership: Serve as the definitive technical authority and primary escalation point within the client ticketing system, taking full ownership of complex hardware and network challenges to develop solutions rather than merely escalating them.
- Continuous Process Improvement: Drive operational efficiencies by evaluating current workflows and implementing systemic process improvements. Lead the documentation and continuous refinement of operational strategies and Standard Operating Procedures (SOPs).
- Technical Guidance & Mentorship: Provide expert technical guidance, mentorship, and training to junior staff on complex ticket resolution, physical interventions, and safety protocols, heavily influencing team development and decisions.
- Risk Assessment & RCA: Identify risks, assess systemic impacts, and lead cross-functional Root Cause Analysis (RCA) investigations for recurrent hardware or facility failures, recommending and implementing corrective actions to prevent future outages.
- Cross-Functional Collaboration: Collaborate directly with infrastructure engineering teams to maintain overarching cluster health, apply specialized expertise to optimize node uptime, and guide the execution of strategic data center maintenance.
- Vendor Strategy & Management: Serve as the primary technical liaison with third-party vendors, using technical judgment to evaluate options, manage advanced RMA escalations, and ensure SLA compliance for hardware replacements.
- Advanced Network Remediation: Evaluate complex fabric topologies and remediate physical layer outages across HPC cluster networks, applying deep expertise in InfiniBand and high-bandwidth optical networks.
- Strategic Incident Response: Respond decisively to critical facility, network, and server events, evaluating impacts and ensuring the physical environment aligns with strict AI/HPC workload requirements (including after-hours support and on-call rotation).
- Security & Compliance Leadership: Enforce physical Security Best Practices, safety guidelines, and compliance standards, proactively identifying vulnerabilities and recommending operational enhancements.
Qualifications
- 5+ years of experience as a data center technician or in a similar complex IT infrastructure environment.
- Demonstrated capability in independent analysis, exercising technical judgment, and making strategic decisions to address complex infrastructure issues.
- Proven track record of taking full ownership of technical solutions, driving process improvements, and contributing to long-term operational strategy.
- Extensive specialized expertise in installing, monitoring, and maintaining high-density data center equipment, particularly in AI/HPC environments.
- Strong ability to provide technical guidance, evaluate complex options, and influence outcomes across engineering and support teams.
- Exceptional skills in identifying risks, assessing broad impacts, and successfully implementing corrective actions.
- Excellent English communication skills to clearly articulate complex technical guidance, risks, and strategies to stakeholders and clients.
- NCA-AIIO, CompTIA ServerPlus, CompTIA Network, or CCNP certification is a plus.
Location
Onsite in Columbus, Georgia
Travel
None
Compensation & Benefits
The base pay range that the Company reasonably expects to pay for this position in Columbus, Georgia is $76,000 - $94,000; the pay ultimately offered may vary based on business considerations, including job-related knowledge, skills, experience, and education. The position is bonus-eligible, and there are medical, dental, and vision benefits available. There is a 401k saving plan and other benefits, such as Paid Time Off, Life Insurance, and an Employee Assistance Plan.
Inclusion & Belonging Statement
We are committed to creating an inclusive environment that embraces differences and fosters belonging for all.
Equal Opportunity Statement
We are an Affirmative Action/Equal Opportunity Employer and strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, national origin, race, ethnicity, creed, gender, disability, veteran status, or any other characteristic protected by law.