Data Engineer II

Data Engineer II role focused on migrating data pipelines and AI/ML workloads to a federated AI platform, developing distributed AI workflows for preprocessing, training, and inference, and providing technical support for data science teams. Experience with cloud environments (AWS, GCP, Azure) and federated learning frameworks is required.

What you'd actually do

  1. Work directly with data scientists and AI model developers to migrate data pipelines and training/inference workloads onto a federated AI platform, ensuring scalability and performance.
  2. Research, design, develop, and test distributed AI workflows, including data preprocessing, model training, and inference across AWS, GCP, and Azure environments.
  3. Provide hands-on technical support to data science teams, including troubleshooting VM, OS-level, and cloud infrastructure issues impacting model execution.
  4. Develop and maintain expertise in federated learning frameworks, including NVIDIA FLARE, to enable secure and efficient distributed model development.

Skills

Required

  • Bachelor's degree
  • Must be legally authorized to work in the United States without the need for employer sponsorship, now or at any time in the future
  • Must be able to obtain and maintain the required clearance for this role
  • 2+ years of experience working in agile environments with hands-on involvement in testing and validating data, ML, or distributed workflows
  • 2+ years of experience implementing solutions in cloud environments (AWS, GCP, or Azure), including deploying data pipelines or AI/ML workloads
  • 2+ years of experience with cloud networking fundamentals, including configuring and troubleshooting network components supporting distributed systems
  • 2+ years of experience supporting or contributing to AI/ML or GenAI solutions, with exposure to model architecture, training, or inference workflows
  • 2+ years of experience developing in Python for data engineering, automation, or machine learning use cases

Nice to have

  • 2+ years of experience refactoring and optimizing code for performance, scalability, and maintainability in cloud or distributed environments
  • 2+ years of experience with identity and access management concepts in cloud or federated environments
  • 2+ years of experience working with modular cloud platforms, orchestration tools, or model control processes supporting ML pipelines

What the JD emphasized

  • 2+ years of experience supporting or contributing to AI/ML or GenAI solutions, with exposure to model architecture, training, or inference workflows

Other signals

  • migrate data pipelines and training/inference workloads onto a federated AI platform
  • Research, design, develop, and test distributed AI workflows, including data preprocessing, model training, and inference
  • Provide hands-on technical support to data science teams, including troubleshooting VM, OS-level, and cloud infrastructure issues impacting model execution
  • Develop and maintain expertise in federated learning frameworks