Core Infrastructure Engineer 2

Oracle Oracle · Enterprise · BENGALURU, KARNATAKA, India

This role focuses on building and scaling intelligent operational platforms (AIOps) using machine learning and LLMs. Responsibilities include designing ML models for anomaly detection, developing operational copilots and chatbots, building feature pipelines, implementing predictive operations, integrating LLMs into observability, developing RAG pipelines, and creating automated remediation workflows. The role requires experience with ML/LLM systems, automation tools, observability platforms, and streaming systems.

What you'd actually do

  1. Design and build AIOps models using LLMs or classical ML for anomaly detection, correlation, root-cause identification, and intelligent event clustering.
  2. Develop operational copilots and chatbots capable of responding to incidents, surfacing insights, and driving automation through natural language.
  3. Build knowledge-grounding systems for operational copilots using runbooks, incident data, historical patterns, service maps, and topology.
  4. Build automated workflows for incident triage, diagnostics, collaboration, and remediation.
  5. Integrate AIOps models with observability platforms handling logs, metrics, traces, events, and topology data.

Skills

Required

  • System Design
  • Platform & reliability engineering
  • ML engineering
  • Data engineering
  • AIOps
  • Python
  • Java
  • PyTorch
  • TensorFlow
  • LLM frameworks
  • Automation workflows
  • StackStorm
  • Rundeck
  • Airflow
  • Jenkins
  • Cloud-native orchestration platforms
  • Observability data (logs, metrics, traces)
  • Datadog
  • Splunk
  • Prometheus
  • Grafana
  • ELK
  • RAG pipelines
  • Embeddings
  • Intent models
  • Operational chatbots
  • Streaming systems
  • Kafka
  • Kinesis
  • Pub/Sub
  • Cloud-native systems
  • Kubernetes
  • Microservices
  • Claude Code
  • Codex
  • GitHub Copilot
  • Context engineering
  • Agentic harness frameworks
  • MCP server

Nice to have

  • LLMs
  • observability
  • automation
  • service reliability
  • predictive operations
  • capacity forecasting
  • early warning systems
  • noisy-neighbor detection
  • knowledge-grounding systems
  • runbooks
  • incident data
  • historical patterns
  • service maps
  • topology
  • LLM-based reasoning
  • retrieval systems
  • intent classification
  • incident triage
  • diagnostics
  • collaboration
  • remediation
  • closed-loop automation
  • alerts
  • insights
  • actions
  • verification
  • reusable automation modules
  • unified observability platforms
  • cloud platforms
  • orchestration systems
  • telemetry
  • logs
  • metrics
  • traces
  • events
  • topology data
  • real-time inference systems
  • high-volume telemetry streams
  • data contracts
  • instrumentation
  • AIOps onboarding patterns
  • enablement models
  • implementation guidelines
  • AIOps adoption
  • architecture reviews
  • data modeling discussions
  • SRE transformation initiatives

What the JD emphasized

  • Must-Have Skills
  • Hands-on experience with at least one of the following tools: Claude Code, Codex, GitHub Copilot
  • Good understanding of context engineering
  • Understanding of agentic harness frameworks
  • Experience building at least one MCP server

Other signals

  • AIOps
  • LLM
  • observability
  • automation
  • incident management
  • anomaly detection
  • chatbots
  • remediation