Software Engineer, ML Platform

Austin, Texas
IDj-4166
Job TypeDirect Hire
Remote TypeHybrid

Software Engineer, ML Platform

Company: Autonomous Mobility Company
Location: Austin, TX
Job Type: Full-Time
Work Setting: Hybrid
Base Compensation: $150K–$160K + annual bonus + equity

About the Team

The ML Platform team at Autonomous Mobility Company builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle.

About the Role

As a Software Engineer, ML Platform, you'll own critical pieces of the ML stack, including workflow orchestration, distributed execution, resource governance, and performance.

You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience.

What You Will Do

  • Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration.

  • Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and I/O.

  • Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, and removing bottlenecks in storage, networking, and resource contention.

  • Work directly with ML teams to debug complex workload issues, drive root-cause analysis, and turn recurring problems into platform-level fixes.

  • Evaluate, integrate, and extend open-source tooling such as Argo Workflows, Ray, and the Kubernetes ecosystem to meet evolving platform needs.

What You Will Need

  • Strong proficiency in Python or Go; C++ is a plus.

  • Track record of designing and building scalable, maintainable systems and services.

  • Experience operating production services end-to-end, including APIs, reliability practices, and observability.

  • Deep knowledge of Kubernetes, including scheduling, resource management, controllers, and pod lifecycle behavior.

  • Solid Linux and systems-debugging skills, including performance investigation, networking, and storage/I/O.

  • Ability to troubleshoot complex production issues across logs, metrics, and traces and drive them to resolution.

Nice to Have

  • Experience with Argo Workflows, Ray, MLflow, or comparable distributed ML tooling.

  • Hands-on experience building or operating large-scale ML training systems, including GPU scheduling, distributed training, and training data pipelines.

  • Track record of optimizing resource usage and performance in distributed environments.

Compensation & Benefits

  • $150K–$160K base salary + annual bonus + equity

  • Employer-subsidized medical, dental, and vision insurance

  • Employer-paid disability and life insurance

  • 401(k) retirement plan

  • Generous paid time off

  • Covered lunches and additional employee benefits

Additional Details

  • Hybrid work setting in Austin, Texas

  • U.S. work authorization required

  • Background check and onboarding verification required

  • Interview process includes technical and cross-functional interviews

Drag & Drop Resume

(PNG, JPEG, PDF, DOC, TXT)

Message & data rates may apply to all numbers allowed to receive messages

Message frequency varies. Text STOP to opt-out or HELP for assistance