Senior, Data Scientist

Walmart Walmart · Retail · Bangalore, KA, India

Senior Data Scientist role focused on enhancing product content quality for improved search accuracy and customer experience within Walmart's Catalog Data Science team. The role involves developing and deploying advanced AI models (NLP, Generative AI) to address challenges like product classification and attribute enrichment. Responsibilities include data sourcing, exploratory analysis, model development (ML, DL, statistical techniques), collaboration, production deployment, visualization, and adopting emerging AI technologies. Requires expertise in Python, PySpark, GCP, big data platforms, and specific ML techniques like embedding generation, RAG, prompt engineering, and fine-tuning.

What you'd actually do

  1. Translate complex business problems into data-driven solutions using advanced analytical and mathematical methods.
  2. Identify and source relevant data sets, ensuring data quality and suitability for modeling purposes.
  3. Develop, test, and validate custom analytical models leveraging machine learning, deep learning, and statistical techniques.
  4. Collaborate with cross-functional teams to align data insights with business objectives and support decision-making.
  5. Deploy scalable models into production environments, ensuring sustainability and performance monitoring.

Skills

Required

  • Python
  • PySpark
  • Google Cloud platform
  • model deployment
  • big data platforms
  • machine learning
  • supervised and unsupervised learning
  • NLP
  • Classification
  • Data/Text Mining
  • Multi-modal models
  • Neural Networks
  • Deep Learning Algorithms
  • Embedding generation
  • Vector Databases
  • LLM gateways
  • Retrieval augmented generation
  • LLM agents
  • model selection
  • prompt engineering
  • finetuning

Nice to have

  • Bachelor's with > 7 years of relevant experience OR Masters with > 5 years of relevant experience OR PHD in Comp Science/Statistics/Mathematics with > 3 years of relevant experience

What the JD emphasized

  • Embedding generation from training materials, storage and retrieval from Vector Databases, set-up and provisioning of managed LLM gateways, development of Retrieval augmented generation based LLM agents, model selection, prompt engineering and finetuning based on accuracy and user-feedback, monitoring and governance.

Other signals

  • Developing and deploying advanced AI models
  • NLP and Generative AI
  • ML Ops
  • Embedding generation
  • Vector Databases
  • Retrieval augmented generation based LLM agents
  • model selection
  • prompt engineering
  • finetuning