Software Engineer II

Microsoft Microsoft · Big Tech · Hyderabad, TS, IN · Software Engineering

Software Engineer to work on a petabyte-scale commercial data platform that turns raw operational and business signals into grounded intelligence, powering agents that help leadership and every persona make faster, better decisions. The role involves designing and building pipelines, curated semantic models, and agents that convert raw data into trusted KPIs, recommendations, and deep-linked actions. The engineer will work across the full stack of a modern data estate, from ingestion and large-scale compute, to curated logical models, to the agent orchestration layer.

What you'd actually do

  1. Design & Build petabyte-scale data pipelines
  2. Convert raw data into intelligence
  3. Design & build agents for leadership and other personas
  4. Deliver KPIs that measure business outcomes
  5. Own engineering fundamentals

Skills

Required

  • 3-5 years of software/data engineering experience building production data pipelines at scale
  • Strong programming skills (e.g., Python, and/or Scala) and expert SQL
  • Hands-on experience with large-scale distributed processing (Spark/Databricks, Synapse) and cloud data lakes (ADLS Gen2 / Fabric OneLake)
  • Experience designing curated/semantic data models and delivering KPIs to business stakeholders
  • Solid grounding in data governance, security, lineage, and data quality
  • Ability to translate ambiguous business questions into robust, well-modeled data products

Nice to have

  • Experience building LLM/agentic applications — orchestration, RAG/grounding, NL2SQL/NL2DAX, skill routing, or MCP-based integrations
  • Familiarity with Azure AI Foundry, Copilot Studio, and evaluation/fine-tuning workflows for agents

What the JD emphasized

  • petabyte-scale commercial data platform
  • agents that help leadership and every persona make faster, better decisions
  • design and build the pipelines, curated semantic models, and agents
  • convert raw data from 200+ transactional sources into trusted KPIs, recommendations, and deep-linked actions
  • work across the full stack of a modern data estate — from ingestion and large-scale compute, to curated logical models, to the agent orchestration layer that answers natural-language questions with facts, reasoning, and recommendations.

Other signals

  • design and build pipelines
  • curated semantic models
  • agents that convert raw data
  • agent orchestration layer