Über diese Machine Learning Engineer, Predictive Maintenance Stelle bei AssetWatch, Inc.
AssetWatch serves global manufacturers by powering manufacturing uptime through the delivery of an unparalleled condition monitoring experience, with a passion to care about the assets our customers care for every day. We are a devoted and capable team that includes world-renowned engineers and distinguished business leaders united by a common goal – To build the future of predictive maintenance. As we enter the next phase of rapid growth, we are seeking people to help lead the journey.
JOB SUMMARY
We are seeking an experienced Machine Learning Engineer with a specialized background in Signal Processing and Predictive Maintenance. The ideal candidate will have a talent for extracting diagnostic features and fault signatures from IIoT sensors, including accelerometers, temperature sensors, and electrical signals, as well as other static and contextual machine data such as images and equipment metadata. The successful candidate will utilize this data to develop and implement algorithms tailored for the diagnosis, prognosis, anomaly detection, and health assessment of critical machinery components, including bearings, gearboxes, shafts, motors, pumps, fans, belts, and more.
This crucial role requires a thorough understanding of classical Machine Learning, Artificial Neural Networks, and Deep Learning algorithms, including CNN and Sequence-based Models such as RNNs and LSTMs enhanced with attention mechanisms. In addition, expertise in supervised and unsupervised learning, anomaly detection, ranking and decision models, feature engineering, and leveraging Large Language Models (LLMs) and agentic AI is a must. This is a hands-on role that also carries solution-architect responsibility: the candidate is expected to design end-to-end ML solutions — from data access and large-scale feature computation through training, inference, and production monitoring — on AWS, using services such as SageMaker (Training, Processing, Pipelines, Endpoints), S3, Athena, AWS Glue, EMR/Spark, Lambda, Step Functions, ECS/Fargate, Timestream, Aurora/RDS, Amazon Bedrock, and CloudWatch. Strong software engineering fundamentals and data/ML engineering at scale (distributed processing with Spark or equivalent) are also essential for building maintainable, testable, and production-ready ML systems. While mastery in algorithm and model development is a must, experience developing and deploying these solutions within the AWS cloud environment and collaborating closely with MLOps and Data Engineering teams is required.
ESSENTIAL RESPONSIBILITIES
- Extract, preprocess, and analyze data from various sensor sources, primarily machine vibration, while also working with temperature, electrical, image, and machine-context data to identify and enhance diagnostic features and fault-specific signatures.
- Determine the appropriate modeling techniques for each problem, including classical Machine Learning, supervised and unsupervised learning, anomaly detection, Artificial Neural Networks, Deep Learning, signal-processing methods, and hybrid or rules-based approaches.
- Train, validate, and deploy predictive maintenance models to accurately identify and predict machinery faults such as bearing, gearbox, shaft, motor, pump, fan, and belt faults. This may include CNNs, RNNs, LSTMs, attention-based models, and other time-series or representation-learning approaches.
- Develop and evaluate time-domain, frequency-domain, and contextual features, including harmonics, sidebands, envelope-spectrum features, running-speed characteristics, bearing and gear-mesh fault frequencies, and other fault-specific evidence.
- Act as the solution architect for assigned ML initiatives: define the end-to-end design — data sources and contracts, batch vs. near-real-time execution, feature computation strategy, storage layout (S3 partitioning, Timestream, Aurora/MySQL), orchestration, inference pattern, and cost/scalability tradeoffs — and document the design before large-scale execution, since backfills across tens of thousands of assets are expensive to rerun.
- Build and scale distributed data and feature pipelines over large historical sensor datasets using Spark (EMR, Glue, or SageMaker Processing), Athena, and Parquet-based data layouts, with attention to partitioning, throughput, and compute cost.
- Design and run data mining and annotation workflows that connect Condition Monitoring Engineers with Data Science — including case discovery, labeling interfaces and conventions, fault-type and time-window labeling, label QA, and turning expert feedback into reusable supervised training data.
- Build and maintain benchmark, training, validation, and regression datasets using real machine cases; work with Condition Monitoring Engineers and other domain experts to establish reliable ground truth and measurable model acceptance criteria.
- Design experiments and compare new algorithms against existing production methods, investigating false positives, false negatives, suppression behavior, data-quality issues, and performance across different machines, operating conditions, sampling configurations, and sensor characteristics.
- Develop robust, modular, reusable, and maintainable ML software in Python, including production-quality model code, shared libraries, configuration, unit and regression tests, code reviews, refactoring, debugging, and version-controlled development workflows.
- Use modern AI-assisted and agentic coding approaches to accelerate data preparation, feature engineering, model prototyping, testing, debugging, pipeline development, and documentation; develop or integrate LLM- and agent-based decision-support capabilities where they provide measurable value.
- Collaborate with the research and engineering team to constantly refine and improve model architectures, algorithms, and signal-processing approaches, ensuring high accuracy, explainability, scalability, and maintainability.
- Work closely with MLOps and Data Engineering teams to ensure smooth deployment of models, feature pipelines, data contracts, and other ML solutions in production environments, including CI/CD for ML, model versioning and registry, IaC (CloudFormation/CDK/Terraform), containerization (Docker/ECR), and production monitoring and alerting via CloudWatch, and participate in production validation and troubleshooting.
- Stay updated with the latest advancements in Machine Learning, Deep Learning, Signal Processing, agentic AI, and industrial condition monitoring, ensuring our solutions remain at the forefront of the industry.
- Present findings, model performance, strategies, tradeoffs, and solutions to other teams and stakeholders in a clear and concise manner, ensuring that insights drive actionable outcomes.
REQUIREMENTS
- Master's or Ph.D. in Mechanical Engineering, Electrical Engineering, Computer Science/Engineering, or a related field.
- Proven experience in Predictive Maintenance and Condition Monitoring with a focus on Signal Processing techniques such as FFT, Short Time Fourier Transform, Time Synchronous Averaging, Wavelet Transform, Spectral Kurtosis, Spectral Correlation, Envelope Analysis, Hilbert Transform, and related methods.
- Strong foundations in Machine Learning and model development, including classical Machine Learning algorithms, supervised and unsupervised learning, anomaly detection, feature engineering, Artificial Neural Networks, Convolutional Neural Networks, Sequence-based Models such as RNNs and LSTMs, and Attention Mechanisms.
- Strong understanding of time-series modeling and evaluation, including the ability to select appropriate techniques based on the physical problem, data characteristics, available ground truth, and operational requirements.
- Hands-on experience architecting and delivering production ML solutions on AWS is required, including several of: SageMaker (Training, Processing, Pipelines, Model Registry, Endpoints), S3, Athena, AWS Glue, EMR, Lambda, Step Functions, ECS/Fargate, Timestream, Aurora/RDS, Amazon Bedrock, IoT Core, CloudWatch, and IAM.
- Demonstrated data and ML engineering at scale, including distributed processing with Spark (PySpark) or equivalent, efficient handling of large time-series and binary signal data, and pragmatic reasoning about compute cost, runtime, and reprocessing risk.
- Experience designing data mining and annotation workflows in partnership with domain experts, including label taxonomy design, labeling tooling, label quality control, and converting noisy expert feedback into model-ready ground truth.
- Strong software engineering and programming skills in Python, with experience developing modular, testable, maintainable production code; proficiency with Git, code review, unit/integration/regression testing, debugging, CI/CD, and reproducible development practices on noisy, large, imperfect industrial datasets.
- Proficiency in scientific and ML/DL frameworks and libraries such as NumPy, Pandas, SciPy, scikit-learn, TensorFlow, PyTorch, or similar tools.
- Experience using Large Language Models, AI-assisted coding, and agentic AI tools for software and ML development; familiarity with designing or integrating LLM- or agent-based workflows for data processing, analysis, decision support, or automation is highly desirable.
- Working knowledge of SQL (MySQL/Aurora, Athena) and the ability to independently explore, validate, and troubleshoot large and imperfect industrial datasets.
- Experience building representative benchmark datasets, validating models against reviewed ground truth, analyzing false positives and false negatives, and creating repeatable evaluation or regression workflows is highly desirable.
- Physics-informed ML and experience with additional modalities (electrical/current signature analysis, ultrasound, oil analysis) are a plus.
- Excellent communication and teamwork skills, including the ability to collaborate with condition-monitoring domain experts, MLOps engineers, Data Engineers, Product, and other stakeholders.
- Demonstrated research work in high-quality journals or significant applied work on related subjects is a plus.
The base salary range for this full-time position is posted below, plus equity and benefits. Variable pay, bonuses, and other cash compensation will be discussed throughout the interview process.
The salary range was determined by role, level, and location. Individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range applicable to your location during the hiring process.
AssetWatch is a remote-first company that puts people at the center of everything we do. We want our team members to thrive - that’s why we offer a range of benefits and perks designed to support your well-being, growth, and work-life balance.
- Competitive compensation package including stock options
- Flexible work schedule
- Comprehensive benefits including retirement plan match
- Opportunity to make a real impact every day
- Work with a dynamic and growing team
- Unlimited PTO