Platform Software Engineer

Oracle Oracle · Enterprise · United States

Platform Software Engineer to build the execution layer for Oracle's AI and GPU growth, focusing on data platforms, GPU scheduling, orchestration, and MLOps tooling to create scalable, production-grade systems.

What you'd actually do

  1. Integrate with Oracle data platforms (OCI Streaming, 23ai, Object Storage, Data Flow) to move training and inference data reliably across customer and internal pipelines.
  2. Build and extend GPU schedulers and capacity-aware placement logic for A100, H100, H200, and Blackwell fleets.
  3. Develop orchestration pipelines for training, fine-tuning, and inference workloads using Kubernetes, Argo, and Slurm where appropriate.
  4. Ship MLOps agentic tooling — observability, automated triage, cost and SLO agents — that reduces operator load on large GPU deployments.
  5. Partner with PMs, SAs, and customer-facing teams to convert field requirements into reusable components, not one-off scripts

Skills

Required

  • Python
  • Kubernetes
  • container runtimes
  • workflow engine (Argo, Airflow, Prefect, or Temporal)
  • distributed systems
  • data platforms
  • ML infrastructure
  • GPU workloads
  • inference serving frameworks (vLLM, TensorRT-LLM, Triton)

Nice to have

  • Go
  • Rust
  • multi-cloud or hybrid deployments
  • QEC
  • quantum-classical hybrid workloads
  • scientific computing
  • open-source contributions to AI infra ecosystem
  • LLM agent frameworks (LangGraph, CrewAI, or similar)

What the JD emphasized

  • MLOps agentic tooling
  • GPU scheduling
  • orchestration pipelines
  • inference data
  • training data
  • GPU workloads
  • LLM agent frameworks

Other signals

  • building execution layer for AI/GPU growth
  • GPU scheduling and orchestration
  • MLOps agentic tooling
  • shipping production-grade systems at scale