Senior Applied Scientist

Oracle Oracle · Enterprise · Seattle, WA +1

The OCI AI Evaluation Science team builds the evidence behind model-selection, product-readiness, and launch decisions. They evaluate frontier foundation models and AI systems across various capabilities. The Senior Applied Scientist will independently own complex evaluation work from problem definition to recommendation to executive leadership, translating ambiguous product and customer questions into measurable hypotheses, selecting or creating benchmarks, designing experiments, building evaluation pipelines, validating data and metrics, analyzing failure modes, and communicating conclusions. This is hands-on applied science involving high-quality code, large datasets, developing and calibrating automated evaluators, and turning one-off analyses into reproducible evaluation protocols and reusable infrastructure. The role examines factors beyond aggregate benchmark scores, including statistical validity, data provenance, contamination, robustness, cost, latency, reliability, safety, and operational constraints. The work bridges research results and product decisions, requiring scientific rigor, engineering judgment, clear writing, and progress in evolving environments. Collaboration with scientists, engineers, product teams, data/human-annotation teams, and external partners is key. The role also involves developing novel benchmarks and evaluation methodologies for publication.

What you'd actually do

  1. Independently own end-to-end evaluations of foundation models, AI agents, and enterprise AI systems, from initial question and experiment design through analysis, reporting, and stakeholder review.
  2. Translate customer, product, and business needs into testable hypotheses, evaluation criteria, datasets, metrics, baselines, and acceptance thresholds.
  3. Design, implement, and maintain benchmarks and evaluation methods for areas such as but not limited to reasoning, coding, agentic workflows, RAG, NL2SQL, multimodal systems, multilingual performance, and responsible AI.
  4. Publish original research in top-tier peer-reviewed conferences and journals, and translate relevant evaluation advances into reusable methods, technical reports, or production capabilities for OCI.
  5. Write production-quality evaluation code; build reproducible pipelines, test suites, automated checks, and integrations with shared evaluation platforms.

Skills

Required

  • ML model evaluation
  • benchmark development
  • experimental design
  • data analysis
  • statistical analysis
  • Python
  • software engineering
  • communication skills

Nice to have

  • LLM-as-a-judge
  • VLM-as-a-judge
  • RAG
  • NL2SQL
  • multimodal systems
  • responsible AI
  • publication record
  • human evaluation design

What the JD emphasized

  • independently own complex evaluation work
  • publish original research
  • production-quality evaluation code
  • evaluate model and system behavior across quality, cost, latency, reliability, safety, robustness, and domain fit rather than relying only on aggregate scores
  • develop and validate automated evaluators, including LLM-as-a-judge and VLM-as-a-judge methods

Other signals

  • evaluation frameworks
  • foundation models
  • AI agents
  • responsible AI
  • benchmarks