Senior Language Engineer

Microsoft Microsoft · Big Tech · Redmond, WA +2 · Applied Sciences

Senior Language Engineer for Microsoft Copilot, focusing on improving LLMs through prompt engineering, data collection, and evaluation frameworks. The role involves crafting prompts, creating evaluation metrics for both technical and EQ performance, researching novel prompting techniques, and analyzing language datasets to enhance model response quality. While not primarily a coding role, it requires foundational software engineering knowledge and the ability to contribute to code repositories for implementing improvements and writing scorers.

What you'd actually do

  1. Craft and refine the context and prompts used to steer, train and evaluate the language models that power Microsoft Copilot for emotional and intellectual use cases.
  2. Create evaluations and establish evaluation frameworks to measure both technical/practical performance and non-deterministic performance like EQ
  3. Research & implement novel prompting techniques
  4. Spin the data flywheel to extract insights from extensive language datasets to uncover issues and opportunities to improve model response quality.
  5. Accountable to own the status of key projects, proactively identifying risks and proposing solutions to ensure timely delivery.

Skills

Required

  • Bachelor’s Degree in Computer Science, Philosophy, Linguistics, Psychology, Literature, or related discipline
  • 1+ years experience in context engineering with foundational software engineering knowledge
  • 3+ years experience working in a fast-paced environment, managing multiple priorities, and adapting to changing requirements and deadlines
  • Hands-on experience with prompt design, context window management, and model evaluation

Nice to have

  • 2+ years of experience shipping consumer-facing products
  • 1+ years of experience in building LLM applications with familiarity in agent and orchestration frameworks, tool use, LLM evaluations and driving efficiency improvements
  • Technical depth in software development, data science and machine learning
  • able to navigate and contribute to code repositories to implement context engineering improvements and write LLM scorers and classifiers to evaluate model response quality

What the JD emphasized

  • hands-on experience with prompt design
  • context window management
  • model evaluation
  • building LLM applications
  • agent and orchestration frameworks
  • tool use
  • LLM evaluations
  • driving efficiency improvements

Other signals

  • improving Large Language Models
  • evaluating language
  • collecting data
  • craft and refine context and prompts
  • create evaluations and establish evaluation frameworks