Production Engineer, Network

Meta Meta · Big Tech · Menlo Park, CA

Production Engineer role focused on building and managing large-scale network infrastructure critical for AI training and inference. Responsibilities include designing and implementing automation, ensuring network reliability, and supporting the deployment of new network products for AI workloads. Requires strong networking, systems, and coding experience, with a demonstrated ability to integrate AI tools into workflows.

What you'd actually do

  1. Conceptualize, build, and maintain automation and tools to support the next generation of network products, network deployment, release engineering and operations.
  2. Develop operational process improvements and implement them in scalable, automated workflows to enhance operational efficiency.
  3. Design and develop solutions that scale across a variety of network platforms.
  4. Lead enhancements of automation for continuous integration, validations, testing infrastructure, release, and configuration management across our global data center network fleet.
  5. Conduct thorough investigations into complex technical issues across networks, ranging from automated tooling to hardware failures and network issues.

Skills

Required

  • networking
  • systems
  • automation
  • tooling
  • software development
  • network device configurations
  • Python
  • Go
  • C++
  • TCP
  • IPv4/6
  • Routing Protocols (BGP, MPLS, ISIS)
  • DHCP
  • DNS
  • software solutions for managing network infrastructure
  • scalability
  • reliability
  • software and network debugging
  • profiling
  • instrumentation
  • distributed systems
  • automated testing infrastructure
  • IB/RDMA/RoCE Networks
  • RDMA congestion control mechanisms
  • AI training workloads
  • integrate AI tools to optimize/redesign workflows
  • responsible, ethical AI practices
  • prompt/context engineering
  • agent orchestration

Nice to have

  • Master's degree or graduate work experience in Computer Science, Computer Engineering, or a related technical field.

What the JD emphasized

  • critical role in supporting it
  • automation plays a critical role
  • ensure the scalability and reliability
  • power one of the largest networks in the world
  • Make the global Datacenter network fleet reliable and available
  • operating and bringing to production all the new network products that enable networking for AI training and Inference
  • operate at a unique intersection of low level systems engineering
  • operating a massively distributed fleet that is uniquely available at Meta
  • exposed to bleeding edge technology
  • automation and tools to support the next generation of network products
  • scalable, automated workflows
  • solutions that scale across a variety of network platforms
  • global data center network fleet
  • complex technical issues
  • weekly on-call rotation
  • Proactively find operational gaps
  • execution plan
  • drive the project directly and through influence
  • peer mentorship
  • 5+ years of experience developing and understanding network device configurations
  • 5+ years of coding experience
  • Demonstrated knowledge of TCP, IPv4/6, Routing Protocols
  • 6+ years of experience building software solutions for managing network infrastructure, with a focus on scalability and reliability
  • In-depth knowledge of software and network debugging, profiling, and instrumentation techniques
  • Proven experience designing, developing, and operating distributed systems at scale
  • Experience designing and maintaining automated testing infrastructure
  • Knowledge of IB/RDMA/RoCE Networks
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact
  • Experience adhering to and implementing responsible, ethical AI practices
  • Demonstrated ongoing AI skill development

Other signals

  • supporting AI goals
  • network for AI training and Inference
  • integrate AI tools to optimize/redesign workflows