Hardware Analytics Engineer

Cerebras Cerebras · Semiconductors · Headquarters +1 · Hardware Departments

The Hardware Analytics Engineer at Cerebras Systems will design and optimize data pipelines for hardware telemetry, reliability, and performance. This role involves architecting ETL processes, developing anomaly detection systems using ML, and analyzing hardware performance to improve efficiency and sustainability. The engineer will also collaborate on next-generation AI platforms and silicon products, focusing on large language model training and inference workloads.

What you'd actually do

  1. Design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry, reliability analytics, and performance optimization.
  2. Architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes to aggregate, process, and analyze multi-terabyte hardware performance and telemetry streams, including utilization, power, thermal, acoustic, and reliability metrics across heterogeneous compute, storage, and AI server platforms, ensuring hardware performance compliance and operational reliability.
  3. Design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark to forecast hardware failure curves, identify performance bottlenecks, and generate prescriptive recommendations for hardware and system optimization.
  4. Lead hardware characterization experiments and thermal/cooling A/B studies to evaluate operational envelopes, delivering validated strategies that reduce carbon footprint, improve water usage efficiency, and maintain or enhance system reliability.
  5. Engineer telemetry ingestion, monitoring, and visualization systems to provide real-time, high-fidelity hardware health data to hardware, firmware, and datacenter operations teams, enabling data-driven decision-making at scale.

Skills

Required

  • Large-scale data pipeline architecture and ETL
  • distributed data processing (Hive, Spark)
  • dashboard development
  • Python
  • SQL
  • Tableau
  • Linux
  • automation scripting
  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction
  • Predictive modeling
  • statistical analysis
  • A/B testing
  • anomaly detection
  • data visualization in hardware reliability and performance
  • Hardware analytics for compute, storage, and AI servers
  • power and thermal optimization
  • GPU burn-in efficiency optimization
  • reliability modeling for AI hardware systems and components including CPU, GPU, DRAM, and SSD

What the JD emphasized

  • Design, training, and deployment of machine learning models for hardware performance optimization and failure prediction

Other signals

  • design and optimize scalable data pipeline architectures for multi-terabyte hardware telemetry
  • architect, develop, and optimize hyperscale data pipeline frameworks and ETL processes
  • design and implement hardware performance analysis and anomaly detection systems using Python, SQL, Tableau, Hive, and Spark
  • define, operationalize, and maintain custom efficiency and reliability metrics
  • perform root cause analysis of systemic failures using large-scale statistical and machine learning methods
  • support the evolution and optimization of next-generation AI platforms and silicon products