Senior Agentic System & Application Engineer

AMD AMD · Semiconductors · Helsinki, Finland · Engineering

This role focuses on designing and developing agentic AI systems for automating complex HPC and enterprise workflows. The engineer will build production-grade agents capable of planning, invoking tools, managing state, launching jobs, diagnosing failures, optimizing execution, and operating safely. Key responsibilities include designing agents with planning, tool use, memory, retrieval, and governance, building production systems for orchestration and automation, automating environments, and implementing safe execution controls. The role requires experience delivering production agentic AI or LLM systems with orchestration, tool use, memory, evaluations, and long-horizon reliability, along with strong Python and systems programming skills.

What you'd actually do

  1. Design agents for HPC and enterprise workflows, incorporating planning, tool use, memory, retrieval, governance, and human approval mechanisms.
  2. Build production systems for job orchestration, enterprise workflow automation, monitoring, debugging, optimization, reproducibility, and rollback.
  3. Automate environments, containers, toolchains, schedulers, CI/CD, enterprise systems, telemetry, workflow composition, and launch processes.
  4. Implement safe execution through authentication, authorization, secrets handling, audit logging, sandboxing, compliance controls, and policy enforcement.
  5. Research and prototype agentic methods for long-running workflows, failure recovery, optimization loops, evaluation, and human-in-the-loop reliability.

Skills

Required

  • Python
  • systems programming language (C, C++, Rust, or Go)
  • integrating automation with real-world tooling (code execution, build/test systems, schedulers, enterprise systems, telemetry, CI/CD, or deployment pipelines)
  • building reliable production systems with observability, fallback mechanisms, regression gates, incident debugging, security controls, and operational ownership

Nice to have

  • HPC workflows (Slurm, Kubernetes, multi-node GPU execution, containers, distributed launch, shared clusters)
  • automating enterprise workflows, developer platforms, IT operations, business systems, approvals, audit trails, or governed tool execution
  • GPU profiling and performance analysis tools and trace-based workflows
  • kernel authoring, tuning, or code generation (Triton, CUDA, HIP, MLIR, LLVM, XLA-like flows)
  • LLM training, supervised fine-tuning, reinforcement learning, evaluation workflows, inference serving, model optimization, or agent evaluation

What the JD emphasized

  • production agentic AI or LLM systems with orchestration, tool use, memory, evaluations, and long-horizon reliability
  • reliable production systems with observability, fallback mechanisms, regression gates, incident debugging, security controls, and operational ownership

Other signals

  • design and develop agentic AI systems
  • production-grade agents
  • orchestration, tool use, memory, retrieval, governance
  • safe execution through authentication, authorization, secrets handling, audit logging, sandboxing, compliance controls, and policy enforcement