Company Overview Docusign brings agreements to life. Over 1.5 million customers and more than a billion people in over 180 countries use Docusign solutions to accelerate the process of doing business and simplify people’s lives. With intelligent agreement management, Docusign unleashes business-critical data that is trapped inside of documents. Until now, these were disconnected from business systems of record, costing businesses time, money, and opportunity. Using Docusign’s Intelligent Agreement Management platform, companies can create, commit, and manage agreements with solutions created by the #1 company in e-signature and contract lifecycle management (CLM). What you'll do We are seeking an experienced Machine Learning Engineer to join our team and drive the design, development, and optimization of large-scale data processing systems. This role requires deep expertise in distributed computing, data pipeline architecture, modern big data technologies, and applied Large Language Model (LLM) integration. This position is an individual contributor role reporting to the Machine Learning Engineering Manager. Responsibility Architect and maintain high-performance, fault-tolerant distributed systems to ensure scalability and availability Design, build, and optimize scalable data pipelines and ETL processes using Python, Spark, and Ray to process complex, unstructured datasets Develop batch data processing workflows to support analytics and machine learning initiatives Integrate and operationalize Large Language Models (LLMs) such as OpenAl, Azure OpenAl, or similar platforms into production applications, including PII detection and redaction workflows Design and maintain prompt engineering strategies, fine-tuning pipelines, and evaluation frameworks for LLM-based solutions Optimize data processing jobs and system performance for cost-efficiency and reliability Collaborate with applied scientists, SMEs, and engineering teams to understand data requirements and deliver robust solutions Implement data quality frameworks and monitoring systems to ensure data integrity Troubleshoot and resolve complex data processing issues in production environments Job Designation Hybrid: Employee divides their time between in-office and remote work. Access to an office location is required. (Frequency: Minimum 2 days per week; may vary by team but will be weekly in-office expectation) Positions at Docusign are assigned a job designation of either In Office, Hybrid or Remote and are specific to the role/job. Preferred job designations are not guaranteed when changing positions within Docusign. Docusign reserves the right to change a position's job designation depending on business needs and as permitted by local law. What you bring Basic 5+ years of experience in data engineering or related roles Experience with Python for data engineering and automation Experience with distributed computing frameworks like Apache Spark or Ray Experience integrating LLMs (e.g., OpenAl, Azure OpenAl, Anthropic, or similar) into applications via APIs, including prompt design, response parsing, and error handling Experience with distributed systems concepts including data partitioning and sharding strategies, fault tolerance and replication, consistency models and distributed consensus, load balancing and resource management Experience with distributed file systems (Azure Data Lake, S3, HDFS) Experience implementing REST APIs using Python frameworks such as FastAPI, Flask, or similar Experience with containerization and orchestration (Docker, Kubernetes) Experience with SQL and database optimization techniques Experience with data structures, algorithms, and software design patterns Experience with version control systems (Git) and CI/CD pipelines Bachelor's or Master's degree in Computer Science, Engineering, or related field, or equivalent practical experience Preferred Experience working with unstructured data (text, images, video, audio) and associated processing techniques Experience with multiple cloud platforms (e.g., AWS, Azure, GCP) Knowledge of data governance and security best practices Experience with machine learning pipelines and MLOps Experience with LLM orchestration frameworks (e.g., LangChain, LlamaIndex) Familiarity with responsible Al practices, including bias mitigation, content filtering, and token cost optimization for LLM-based applications Contributions to open-source projects Excellent problem-solving skills and ability to work with complex, ambiguous requirements Strong communication skills and ability to collaborate across teams Life at Docusign Working here Docusign is committed to building trust and making the world more agreeable for our employees, customers and the communities in which we live and work. You can count on us to listen, be honest, and try our best to do what’s right, every day. At Docusign, everything is equal. We each have a responsibility to ensure every team member has an equal opportunity to succeed, to be heard, to exchange ideas openly, to build lasting relationships, and to do the work of their life. Best of all, you will be able to feel deep pride in the work you do, because your contribution helps us make the world better than we found it. And for that, you’ll be loved by us, our customers, and the world in which we live. Accommodation Docusign is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need such an accommodation, or a religious accommodation, during the application process, please contact us at accommodations@docusign.com. If you experience any issues, concerns, or technical difficulties during the application process please get in touch with our Talent organization at taops@docusign.com for assistance. Applicant and Candidate Privacy Notice #LI-Hybrid #LI-IC1
Machine Learning Engineer
Machine Learning Engineer responsible for designing, developing, and optimizing large-scale data processing systems and integrating LLMs into production applications. Focuses on distributed computing, data pipeline architecture, and LLM operationalization.
What you'd actually do
- Architect and maintain high-performance, fault-tolerant distributed systems to ensure scalability and availability
- Design, build, and optimize scalable data pipelines and ETL processes using Python, Spark, and Ray to process complex, unstructured datasets
- Develop batch data processing workflows to support analytics and machine learning initiatives
- Integrate and operationalize Large Language Models (LLMs) such as OpenAl, Azure OpenAl, or similar platforms into production applications, including PII detection and redaction workflows
- Design and maintain prompt engineering strategies, fine-tuning pipelines, and evaluation frameworks for LLM-based solutions
Skills
Required
- Python
- Spark
- Ray
- LLM integration
- Prompt engineering
- Distributed systems
- Data pipelines
- ETL
- REST APIs
- Docker
- Kubernetes
- SQL
- Git
- CI/CD
Nice to have
- Unstructured data processing
- Cloud platforms (AWS, Azure, GCP)
- MLOps
- LLM orchestration frameworks (LangChain, LlamaIndex)
- Responsible AI practices
- Open-source contributions
What the JD emphasized
- 5+ years of experience in data engineering or related roles
- Experience with Python for data engineering and automation
- Experience with distributed computing frameworks like Apache Spark or Ray
- Experience integrating LLMs (e.g., OpenAl, Azure OpenAl, Anthropic, or similar) into applications via APIs, including prompt design, response parsing, and error handling
- Experience with distributed systems concepts including data partitioning and sharding strategies, fault tolerance and replication, consistency models and distributed consensus, load balancing and resource management
- Experience with distributed file systems (Azure Data Lake, S3, HDFS)
- Experience implementing REST APIs using Python frameworks such as FastAPI, Flask, or similar
- Experience with containerization and orchestration (Docker, Kubernetes)
- Experience with SQL and database optimization techniques
- Experience with data structures, algorithms, and software design patterns
- Experience with version control systems (Git) and CI/CD pipelines
Other signals
- LLM integration
- data pipelines
- distributed systems