Unlock Full Resume Report

New offer - be the first one to apply!

September 22, 2026

Principal AI Software Engineer

Senior

188,000 - 304,200 USD/yr

Redmond, WA

Quick Facts

  • Role: Principal AI Software Engineer

Description

Lead end-to-end system software prototyping to build proof-of-concepts for hardware/software co-designed capabilities that reduce memory TCO for Azure compute. Drive workload characterization, correlation, and performance modeling to identify and evaluate optimizations for LLM inference—especially KV cache capacity, placement, migration, and utilization across GPU HBM, host DRAM, CXL memory expansion/pooling, SSD, and emerging memory tiers. Collaborate with workload experts to influence hardware architecture and technical direction for multi-year product readiness.

Responsibilities

  • Prototype full system software for memory tiering/pooling and overcommit solutions for Azure usage and deployment scenarios
  • Characterize workloads to find systems optimization opportunities
  • Engineer TCO-optimized solutions for Azure general-purpose and specialized compute fleets with internal and partner experts
  • Influence hardware architecture and industry alignment with data-driven analysis and recommendations
  • Characterize and optimize LLM inference workloads (KV cache, capacity, placement, migration, utilization)
  • Build evaluation frameworks and proof-of-concepts for memory-tiering architectures for AI inference (CXL pooled memory, context-memory platforms, SSD-backed cache tiers)
  • Run characterization studies for agentic, multi-turn, coding, reasoning, and long-context AI workloads
  • Analyze end-to-end data movement across GPU/CPU/storage/network; optimize using GDS/GDR, peer-to-peer transfers, and distributed inference pipelines
  • Develop software prototypes, instrumentation, and extensions to evaluate KV cache offload, prefetching, migration, compression, deduplication, and memory overcommit
  • Build performance models and simulation frameworks to predict the impact of memory hierarchy innovations

Requirements

  • Bachelor’s degree in Computer Science or related technical field AND 6+ years of technical engineering experience with coding in C/C++/C#/Java/JavaScript/Python (or equivalent)
  • Ability to meet Microsoft, customer, and/or government security screening requirements; pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter

Preferred Qualifications

  • 12+ years systems software experience (OS kernel, memory management, I/O stacks, virtualization) with success guiding architecture/software enablement
  • 10+ years leading hardware/software co-design projects influencing technical direction
  • Deep expertise in Linux kernel internals, memory management, I/O subsystems, NUMA, DMA, and GPU/CPU/storage/network data paths
  • Hands-on NVIDIA GPU software stack experience (CUDA, NCCL, GPUDirect Storage (GDS), GPUDirect RDMA (GDR))
  • Understanding of AI inference infrastructure, large GPU clusters, inference serving architectures, and workload performance optimization
  • Experience optimizing KV cache intensive workloads (long-context, agentic, multi-turn, coding, reasoning)
  • Experience with inference frameworks (vLLM, SGLang, TensorRT-LLM)
  • Familiarity with KV cache technologies (LMCache, SGLang HiCache, cache offload, cache sharing, memory tiering)
  • Experience designing/extending inference runtimes, scheduling systems, memory management components, or KV cache subsystems
  • Understanding of disaggregated prefill/decode and distributed inference serving with multi-node cache-sharing topologies
  • Experience with CXL memory expansion, memory pooling, and memory tiering in large-scale deployments
  • Strong software development in C/C++, Python, CUDA, and distributed systems
  • Ability to partner with and influence architects, hardware engineers, and software leads
  • Collaboration and communication skills for technical and non-technical stakeholders

Benefits

  • Certain roles may be eligible for benefits and other compensation (details provided via Microsoft corporate pay page)

Similar jobs you might like