Machine Learning Engineer

DocuSign DocuSign · Enterprise · Dublin, Ireland · Engineering

Machine Learning Engineer responsible for designing, developing, and optimizing large-scale data processing systems and integrating LLMs into production applications. Focuses on distributed computing, data pipeline architecture, and LLM operationalization.

What you'd actually do

  1. Architect and maintain high-performance, fault-tolerant distributed systems to ensure scalability and availability
  2. Design, build, and optimize scalable data pipelines and ETL processes using Python, Spark, and Ray to process complex, unstructured datasets
  3. Develop batch data processing workflows to support analytics and machine learning initiatives
  4. Integrate and operationalize Large Language Models (LLMs) such as OpenAl, Azure OpenAl, or similar platforms into production applications, including PII detection and redaction workflows
  5. Design and maintain prompt engineering strategies, fine-tuning pipelines, and evaluation frameworks for LLM-based solutions

Skills

Required

  • Python
  • Spark
  • Ray
  • LLM integration
  • Prompt engineering
  • Distributed systems
  • Data pipelines
  • ETL
  • REST APIs
  • Docker
  • Kubernetes
  • SQL
  • Git
  • CI/CD

Nice to have

  • Unstructured data processing
  • Cloud platforms (AWS, Azure, GCP)
  • MLOps
  • LLM orchestration frameworks (LangChain, LlamaIndex)
  • Responsible AI practices
  • Open-source contributions

What the JD emphasized

  • 5+ years of experience in data engineering or related roles
  • Experience with Python for data engineering and automation
  • Experience with distributed computing frameworks like Apache Spark or Ray
  • Experience integrating LLMs (e.g., OpenAl, Azure OpenAl, Anthropic, or similar) into applications via APIs, including prompt design, response parsing, and error handling
  • Experience with distributed systems concepts including data partitioning and sharding strategies, fault tolerance and replication, consistency models and distributed consensus, load balancing and resource management
  • Experience with distributed file systems (Azure Data Lake, S3, HDFS)
  • Experience implementing REST APIs using Python frameworks such as FastAPI, Flask, or similar
  • Experience with containerization and orchestration (Docker, Kubernetes)
  • Experience with SQL and database optimization techniques
  • Experience with data structures, algorithms, and software design patterns
  • Experience with version control systems (Git) and CI/CD pipelines

Other signals

  • LLM integration
  • data pipelines
  • distributed systems