Senior Network Engineer

Together AI Together AI · Data AI · San Francisco, CA · Engineering

Senior Network Engineer responsible for designing, implementing, and maintaining network infrastructure for AI company's user-facing services and production systems. Focus on routing, switching, network security, and protocols, with an emphasis on automation and HPC data center networking. Experience with large-scale hybrid data center networks, TCP/IP, BGP, OSPF, VXLAN, EVPN, QoS, Python/Ansible automation, network troubleshooting tools, multi-tenant networks, and major network device vendors (Cisco, Arista, Juniper, Mellanox). Cloud network experience (AWS, GCP, Azure) and Linux environment proficiency are also required. Preferred knowledge includes RoCE, Infiniband, Docker, Kubernetes, Slurm, and understanding AI training workloads.

What you'd actually do

  1. Design, deploy, manage and maintain global multi-vendor, multi-protocol high performance compute networks.
  2. Analyze data to diagnose and identify root causes to network issues to minimize downtime
  3. Evaluate and recommend network technologies, hardware, and software solutions.

Skills

Required

  • TCP/IP networking architecture and technologies such as BGP, OSPF, VXLAN, EVPN, and QoS
  • Python, Ansible, or other languages/tools utilized in infrastructure automation
  • Wireshark, tcpdump, nmap, MTR, and curl
  • multi-tenant networks
  • Cisco, Arista, Juniper, and Mellanox
  • AWS, GCP, and Azure
  • Linux environment

Nice to have

  • RoCE and Infiniband protocols
  • Docker, Kubernetes, or Slurm
  • AI training workloads

What the JD emphasized

  • 8+ years of professional experience building, managing, and supporting large-scale hybrid data center networks (excluding enterprise networks).