Notes - The Man Who Built OpenAI's First Chip

TechTechPotato | September 29, 2026

Overview of Richard Ho and OpenAI's Custom Silicon

  • TechTechPotato hosted an interview with Richard Ho, the Vice President of Hardware at OpenAI, to discuss OpenAI's custom hardware announcement.
  • At the Hot Chips conference, OpenAI unveiled its custom machine learning accelerator chip designed to advance AI infrastructure and optimize inference workloads.
  • The discussion covered Ho's background in supercomputing and verification, the genesis of OpenAI's hardware team, the architectural principles of their chip, and the integration of AI models into the chip design process.

Richard Ho's Engineering Career and Background

  • Richard Ho began his career at Stanford University under John Hennessy, working as a junior engineer on a multi-node multiprocessor project.
  • Because thesis topics were taken by PhD students, Ho specialized in hardware verification, earning his PhD in verification and co-founding an Electronic Design Automation (EDA) verification startup.
  • After his startup was acquired, Ho joined D. E. Shaw Research to build supercomputers for molecular dynamics and drug discovery, aiming to assist in cancer research.
  • At D. E. Shaw Research, Ho contributed to building Anton 1, Anton 2, and Anton 3, which served as structural precursors to dedicated AI ASICs and specialized networking.
  • Ho was subsequently recruited as one of the earliest hires for Google's TPU program under secret conditions.
  • The Google TPU initiative was triggered by Jeff Dean's realization that Google Translate was consuming an unsustainable portion of data center compute power.
  • Over his tenure at Google, Ho helped design seven to eight generations of TPUs with a core team of approximately 300 engineers, relying on an ASIC partner for back-end physical design.
  • Ho briefly left Google to join a silicon photonics startup focusing on optical communication, believing optical interconnects would become critical for scaling compute infrastructure.
  • Founding an EDA startup taught Ho to engineer specifically for customer needs, challenge mainstream technical assumptions, and feel comfortable making large architectural bets.

Founding OpenAI's Hardware Team and Philosophy

  • OpenAI originally approached Ho to explore hardware possibilities, leading to one-on-one discussions with CEO Sam Altman, who was deeply committed to hardware, fabs, and long-term compute capacity.
  • Ho assembled the hardware team using an "Ocean's 11" approach, recruiting top talent from Google, Meta, Amazon, and other industry leaders.
  • The hardware group operates as an internal startup within OpenAI, benefiting from a "blank slate" without legacy hardware constraints or backward compatibility requirements.
  • To maximize alignment, the hardware team sits directly alongside AI researchers, compiler developers, kernel optimization experts, and inference teams in the same space.
  • Ho prioritized hardware programmability and fast execution speed over rigid fixed-function logic to prevent over-fitting to rapidly evolving AI models.
  • While OpenAI is diligent with its resources, Altman provided strong executive backing and substantial freedom to build the necessary hardware rapidly.

Architectural Design and Features of the "Jalapeño" Chip

  • The internal code name "Jalapeño" originated from a spice rack in the cafeteria at OpenAI's Mission District building, leading the team to adopt "Hot Peppers" as their group identity.
  • The Jalapeño package features a compute die, High Bandwidth Memory (HBM) stacks, and a dedicated I/O chiplet.
  • To address memory bottlenecks, the chip implements a NUMA-like architecture (NUMA 64) that tightly couples HBM memory banks directly to individual compute cores, drastically reducing high-energy data movement.
  • Rather than adopting heterogeneous disaggregation across prefill and decode nodes, OpenAI designed Jalapeño as a single homogeneous device capable of running the entire spectrum of inference workloads (prefill, speculative decoding, and decode).
  • This homogeneous strategy gives data centers maximum optionality and resource fungibility, avoiding rigid capital expenditure commitments to fixed hardware ratios.
  • The development roadmap included an intentional "zero stepping" (A0) to bring up software and system layers early, followed by a B0 stepping for final performance tuning and qualification.
  • To evaluate performance objectively, the team benchmarked Jalapeño using the third-party open-source InferenceX benchmark suite.
  • During a late development sprint for Hot Chips, Jalapeño's Single Token Prediction (STP) throughput beat the published Multi-Token Prediction (MTP) numbers of competing hardware.

Utilizing AI Models and Co-Design in Chip Development

  • The hardware team integrated OpenAI's raw, un-fine-tuned internal foundation models directly into the hardware development pipeline.
  • When logic failed to fit within strict physical area constraints, AI models performed multidimensional optimization, reducing total logic area by over 13% without performance degradation.
  • The team utilized XLS, an open-source high-level synthesis language with Rust-like syntax developed by team member Chris Leary, because AI models understood software semantics far better than traditional Verilog.
  • Ho views AI models as tools that transform engineers into "super-engineers," while stressing that traditional EDA sign-off flows and verification remain essential guardrails to verify 100% functional correctness.

Industry Partnerships and Manufacturing Realities

  • OpenAI partnered with Broadcom for back-end physical design, key intellectual property (SerDes and controllers), packaging expertise, and wafer allocation access as one of TSMC's largest volume clients.
  • OpenAI engages directly with memory manufacturers like Samsung and ix to formulate long-term system architectures, custom HBM configurations, and 3D memory integration.
  • Ho notes that computer engineering has transitioned from the "Golden Age of Architecture" into the "Golden Age of Packaging," where system-on-wafer, panel integration, co-packaged optics, and 3D stacking dictate system performance.

Long-Term Vision and Industry Bottlenecks

  • Ho identifies human mindset and executive hesitation as the single biggest bottleneck in AI infrastructure, as industry leaders fail to grasp the exponential scale required and under-invest relative to actual future demand.
  • In OpenAI's compute strategy, older hardware generations are deprecated down to lower-cost or free user tiers until energy consumption necessitates decommissioning.
  • The ultimate objective of OpenAI's hardware program is full-stack optimization across data centers and edge infrastructure to dramatically reduce the cost of compute, making AI intelligence pervasive and accessible globally.