Staff Software Engineer
Uber is seeking a Staff Software Engineer to focus on building and scaling ML inference services. This role involves optimizing models for production environments, ensuring low latency and high throughput, and collaborating with research and product teams to deploy ML models.
What you'd actually do
- Develop and deploy scalable, high-performance machine learning inference services.
- Optimize ML models for low-latency and high-throughput serving.
- Collaborate with ML researchers and product teams to bring cutting-edge models to production.
- Design and implement robust monitoring and alerting systems for ML inference.
- Contribute to the development of internal ML platforms and tooling.
Skills
Required
- Experience with ML model serving frameworks (e.g., TensorFlow Serving, TorchServe, Triton).
- Proficiency in Python and C++.
- Strong understanding of distributed systems and cloud computing (AWS, GCP, or Azure).
- Experience with containerization technologies (Docker, Kubernetes).
- Familiarity with ML model optimization techniques (quantization, pruning, distillation).
Nice to have
- Experience with ML performance profiling and debugging.
- Knowledge of GPU acceleration for ML inference.
- Experience with CI/CD pipelines for ML systems.
Other signals
- Develop and deploy scalable, high-performance machine learning inference services.
- Optimize ML models for low-latency and high-throughput serving.
- Collaborate with ML researchers and product teams to bring cutting-edge models to production.