Unlock Full Resume Report

New offer - be the first one to apply!

September 29, 2026

Software Development Engineer, AI/ML (AWS Neuron Inference)

Mid

180,000 - 210,000 USD/yr

Cupertino, CA

Quick Facts

  • Role: Software Development Engineer for AI/ML inference acceleration using AWS Neuron

Description

Help lead distributed inference support for PyTorch in the Neuron SDK. Design, develop, and optimize ML models, frameworks, and inference infrastructure to maximize performance and efficiency on AWS Trainium and Inferentia accelerators. Work across the stack (compiler, runtime, frameworks, and hardware) to build kernels, profile bottlenecks, and enable large-scale LLM families for customers.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators
  • Support all stages of the ML system lifecycle: distributed architecture design, profiling, hardware-specific optimizations, testing, and production deployment
  • Build infrastructure to analyze and onboard multiple models with diverse architectures
  • Design and implement high-performance kernels and ML operation features leveraging the Neuron architecture
  • Analyze and optimize system-level performance across Neuron hardware generations
  • Use profiling tools to identify and resolve performance bottlenecks
  • Implement optimizations such as fusion, sharding, tiling, and scheduling
  • Conduct unit and end-to-end model testing with continuous deployment and releases through pipelines
  • Work with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate with cross-functional teams to develop innovative optimization techniques

Requirements

  • Bachelor’s degree in Computer Science or equivalent
  • 3+ years of non-internship professional software development experience
  • 3+ years of non-internship systems design/architecture experience
  • Fundamentals of machine learning and LLMs, including architecture and training/inference lifecycles, plus optimization experience for improving model execution
  • C++ and Python software development experience (at least one language required)
  • Strong understanding of system performance, memory management, and parallel computing
  • Deep understanding of computer architecture and OS-level software with working knowledge of parallel computing
  • Proficiency in debugging, profiling, and implementing best practices in large-scale systems

Preferred Qualifications

  • Master’s degree or Ph.D. in Computer Science or equivalent
  • Familiarity with PyTorch, JIT compilation, and AOT tracing
  • Familiarity with CUDA kernels or equivalent low-level ML kernels (e.g., CUTLASS, FlashInfer)
  • Familiar with Triton-like tile-level semantics
  • Experience with online/offline inference serving platforms in production (e.g., vLLM, SGLang, TensorRT)

Benefits

  • Health insurance (medical, dental, vision, prescription, Basic Life & AD&D) and optional Supplemental life
  • EAP and mental health support; Medical Advice Line
  • Flexible Spending Accounts; Adoption and Surrogacy Reimbursement
  • 401(k) matching
  • Paid time off and parental leave

Similar jobs you might like