Principal AI Agent Architect
Bangalore, IN, 560048
At Penguin Solutions (Nasdaq: PENG) – The AI Factory Platform Company – we’re building a team of innovators who thrive on collaboration, creativity, and the opportunity to help shape the future of AI. As part of the AI technology revolution, our teams design, build, deploy, and manage AI factories for enterprises, sovereign AI initiatives, and neocloud providers worldwide.
Headquartered in Silicon Valley, California, Penguin Solutions operates globally through a network of R&D, manufacturing, and sales locations. For nearly three decades, we have operated at the intersection of memory and AI/HPC infrastructure. That engineering expertise positions us to power the next generation of AI workloads, from training to inference and agentic AI at scale.
Penguin Solutions brings together differentiated infrastructure software, advanced memory, compute systems, end-to-end services, and industry-leading partner solutions in a full-stack AI factory platform designed to help customers deploy and scale AI workloads with speed and precision.
At Penguin Solutions, we value ideas over hierarchy and believe in servant leadership, where leaders enable teams to do their best work. We empower employees to take ownership, drive innovation, and grow through challenging work, continuous learning, and exposure to advanced AI tools and technologies. With flexibility where it matters and a strong focus on outcomes, Penguin Solutions is a place to do your best work, grow your career, and make a meaningful impact.
Job Overview
We are looking for a Principal AI Agent Architect to design and deliver key architectural components of ClusterWareAI's next-generation AI Operational Agent. You will help build a secure, reliable, and measurable agentic AI platform capable of reasoning over enterprise infrastructure, orchestrating tools and workflows, and executing actions with appropriate levels of autonomy.
This is a hands-on technical leadership role for a senior technical expert who will design complex production AI systems, partner closely with engineering and technical leadership, mentor engineers, and contribute to the evolution of ClusterWareAI's AI platform architecture.
Responsibilities
- Design and implement key architectural components of the ClusterWareAI AI Operational Agent, including planning, reasoning, memory, state management, tool orchestration, and autonomous or human-approved execution.
- Design scalable agentic workflows and enterprise AI architectures, including multi-agent approaches where appropriate.
- Design knowledge management, RAG, context engineering, agent memory, and Model Context Protocol (MCP) capabilities for enterprise infrastructure automation.
- Design reliable execution patterns for long-running and asynchronous agent workflows, including state management, retries, recovery, fallbacks, and idempotency.
- Contribute to AI evaluation and observability frameworks covering task success, quality, reliability, latency, cost, tool execution, and failure modes.
- Design and implement guardrails and secure execution patterns for autonomous operations, including appropriate controls for tool access, privileges, data protection, tenant isolation, and human approval.
- Partner with Product Management, AI, platform, and infrastructure engineering teams to translate customer and operational requirements into scalable AI capabilities.
- Evaluate model and architecture choices, including model selection, routing, inference performance, reliability, and cost tradeoffs.
- Mentor engineers, contribute to technical standards and engineering best practices, and provide technical guidance on complex AI system design.
- Evaluate emerging AI technologies, frameworks, and platform capabilities and provide recommendations for adoption.
Required Qualifications
- 10+ years of software engineering experience, including demonstrated experience designing and delivering complex production systems.
- Deep hands-on experience designing, building, and operating production-grade LLM or agentic AI systems involving planning, reasoning, tool use, orchestration, and autonomous or human-approved execution.
- Strong expertise in Python and production software engineering.
- Strong expertise in distributed systems and fault-tolerant architectures, including asynchronous workflows, state management, retries, recovery, failure handling, and idempotency.
- Experience with agent orchestration frameworks such as LangGraph, Semantic Kernel, AutoGen, or similar, and familiarity with Model Context Protocol (MCP) or equivalent tool-integration patterns.
- Experience with RAG architectures, vector databases, knowledge retrieval, context engineering, and agent memory.
- Strong understanding of enterprise software architecture, APIs, microservices, and cloud-native application development.
- Hands-on experience with Kubernetes and cloud-native infrastructure, including deployment, APIs, automation, observability, and production operations.
- Experience designing or implementing AI evaluation and observability approaches with measurable quality, reliability, latency, and cost metrics.
- Strong understanding of AI/agent security, including prompt injection, tool security, excessive agency, privilege management, data protection, tenant isolation, and secure execution.
- Experience evaluating model selection, routing, inference performance, and cost tradeoffs in production AI systems.
- Demonstrated technical leadership, architecture experience, and ability to mentor engineers and influence complex technical decisions.
Preferred Qualifications
- 15+ years of software engineering experience, with significant experience in architecture or senior technical leadership roles.
- Experience building AI agents for infrastructure, cloud operations, SRE, Kubernetes, or enterprise IT operations.
- Experience with infrastructure management, observability, virtualization, storage, networking, or cloud platforms.
- Experience designing human-in-the-loop or human-on-the-loop systems and risk-based levels of agent autonomy.
- Experience with long-running agent workflows, persistent memory, multi-agent systems, and advanced context engineering.
- Experience operating AI systems in multi-tenant enterprise environments with strong security, auditability, and reliability requirements.
- Experience with GPU-based AI infrastructure, inference platforms, or AI workload orchestration.
- Contributions to open-source AI agent, Kubernetes, AI infrastructure, or distributed-systems projects.
- Published technical content, conference presentations, patents, or other demonstrated thought leadership in AI or distributed systems are a plus.
Location
Hybrid – Bangalore, India