Unlock Full Resume Report

August 10, 2026

Senior Performance Co-Design Engineer, TPU

Senior • On-site

Sunnyvale, CA

Minimum qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, a related field, or equivalent practical experience.
  • 5 years of experience in performance modeling/engineering, computer architecture, co-design, or systems engineering.
  • Experience programming in C++ or Python.

Preferred qualifications

  • Master's degree or PhD in Electrical Engineering, Computer Engineering, or Computer Science, with an emphasis on computer architecture.
  • Experience with hardware/software co-design problems, especially performance analysis and identification at the pre-silicon stage.
  • Experience enabling and optimizing large-scale ML models, such as LLMs and large embedding models.
  • Experience with ML infrastructure, profiling tools, or deep learning inference/serving optimizations.
  • Familiarity with accelerator architectures.

Description

Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action.

The TPU Chip Architecture and Performance Co-design team is at the forefront of optimizing custom AI silicon for next-generation machine learning models. As a Senior Performance Co-Design Engineer, you will focus on conducting LLM serving studies. You will analyze and optimize the serving performance of emerging models and use cases on custom hardware, and work closely with hardware architects to influence the evolution of custom ML accelerators.

The AI and Infrastructure team delivers AI and Infrastructure at scale, efficiency, reliability, and velocity. Customers include Googlers, Google Cloud customers, and billions of Google users worldwide. Teams work across software and hardware, including TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and more.

Responsibilities

  • Conduct comprehensive serving performance studies on current and emerging LLMs, including 1P and 3P models.
  • Develop and maintain advanced simulation, profiling, and modeling tools to identify bottlenecks, understand key characteristics, and project serving workload performance.
  • Partner with model researchers and software and hardware teams to co-design architectural improvements tailored to LLM inference latency and throughput.
  • Drive data-backed decisions that influence the roadmap for future TPU and Cloud Silicon architectures.

Similar jobs you might like