Field Application Engineer Manager, Cloud AI Infrastructure
Software Engineering, Other Engineering, Data Science
Austin, TX, USA · Kirkland, WA, USA
Posted on Sep 25, 2026
info_outline
In accordance with Washington state law, we are highlighting our comprehensive benefits package, which is available to all eligible US based employees. Benefits for this role include:- Health, dental, vision, life, disability insurance
- Retirement Benefits: 401(k) with company match
- Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
- Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
- Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
- Baby Bonding Leave: 18 weeks
- Holidays: 13 paid days per year
Minimum qualifications:
- Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science, or IT-related field, or equivalent practical experience.
- 8 years of experience with Linux/Unix systems and experience in debugging issues across the hardware/software boundary on enterprise-grade server infrastructure.
- 5 years of experience in technical leadership.
- 3 years of experience with technical infrastructure (e.g., deployment, maintenance, and troubleshooting), and with quality and reliability of technical infrastructure.
- 3 years of debug or validation experience with CPU, dGPU, or TPU.
- Experience troubleshooting and triaging technical issues across the stack (e.g., hardware faults, low-level software, networking, virtualization, kernel drivers, firmware, or performance).
Preferred qualifications:
- Experience working directly with AI/ML computing hardware, including GPUs or other accelerators.
- Experience with systems automation, and with systems design and debug.
- Experience working with distributed systems, and familiarity with common solutions, design patterns, or best practices.
- Experience with ML frameworks (e.g., TensorFlow, PyTorch), and understanding of the AI/ML training and inference lifecycle.
- Advanced understanding of memory and high-speed IO technologies.
- Familiarity with containerization and orchestration technologies like Kubernetes or Slurm in an on-prem or cloud environment.