Senior Software Engineer - Simulation Orchestration

Applied Intuition Applied Intuition · Robotics · Sunnyvale, CA · Data Engineering

This role focuses on developing and owning the compute platform for simulation orchestration, involving workload scheduling, resource allocation, and lifecycle management, with a strong emphasis on distributed systems challenges like GPU scheduling and autoscaling within a Kubernetes environment. The goal is to expand this into a company-wide compute platform.

What you'd actually do

  1. Develop a multi-environment compute platform for workload scheduling, resource allocation, and node lifecycle.
  2. Collaborate with customers and internal teams to translate compute needs into platform features.
  3. Solve distributed systems challenges including GPU scheduling, autoscaling, and resource efficiency.
  4. Ensure high reliability and scalability for production simulation systems.
  5. Drive architecture decisions to expand simulation orchestration into a company-wide compute platform.

Skills

Required

  • 5+ years of backend engineering experience
  • 4+ years of coding in Golang, Python, Java, or C++
  • 4+ years building complex backend systems on major cloud providers
  • Deep Kubernetes expertise, including pod scheduling, node lifecycle, and autoscaling under load
  • Experience building or operating workload scheduling systems (e.g., dispatch, bin-packing, quota enforcement)
  • Self-starter with experience driving production system development from 0 to 1

Nice to have

  • Experience scheduling GPU workloads or managing GPU capacity at scale
  • Multi-cloud (AWS, OCI, GCP, Azure) or on-prem Kubernetes deployments
  • Experience managing multi-cluster environments using Infrastructure as Code
  • Experience building custom Kubernetes controllers or operators
  • Experience with cloud cost optimization strategies for high-scale compute workloads

What the JD emphasized

  • track record of owning production distributed systems
  • Deep Kubernetes expertise
  • building or operating workload scheduling systems
  • driving production system development from 0 to 1