About the Role
Job Description
We are looking for competence in designing, configuring, deploying, and maintaining machine learning solutions for anomaly detection and early warning in a data center environment. The solution must run on-premises and integrate with existing operational data streams, including InfluxDB as the time-series database and Kafka as part of the data pipeline. The role requires practical experience with machine learning, time-series analytics, Linux operations, and production deployment in industrial or infrastructure environments.
Required Competence
- Experience with time-series data analysis and anomaly detection for infrastructure systems such as cooling, power, UPS, environmental sensors, network equipment, and IT load.
- Ability to evaluate and select suitable ML models for anomaly detection, forecasting, and early warning use cases.
- Knowledge of statistical anomaly detection, Isolation Forest, Autoencoders, LSTM models, Prophet, ARIMA, streaming anomaly detection, and hybrid rule/ML-based systems.
- Ability to explain trade-offs related to accuracy, explainability, training data requirements, operational complexity, false positives, and real-time performance.
Programming and Development Competence
- Experience with recognized programming languages commonly used for machine learning and data engineering such as Python, Go, Java, Scala, Rust, or C++.
- Understanding of frameworks and tooling for machine learning pipelines, data processing, and model serving.
- Experience building APIs and backend services for operational environments.
- Ability to work with containerized and distributed systems.
Technical Environment
- Ubuntu Linux configuration, hardening, and operations.
- InfluxDB integration for time-series storage and retrieval.
- Kafka integration for streaming pipelines and event-driven architectures.
- Docker or Kubernetes-based deployment.
- REST API development using frameworks such as FastAPI, Spring Boot, Go services, or similar technologies.
- Logging, monitoring, alerting, and production operations.
Data Pipeline and Integration
- Reading real-time or near-real-time data from Kafka.
- Querying historical data from InfluxDB for model training.
- Writing anomaly scores and prediction results back to InfluxDB.
- Exposing results through APIs, dashboards, or Grafana integrations.
- Supporting integration with ITSM or operational monitoring systems.
Deployment Responsibilities
- Installing required dependencies and runtime environments.
- Configuring services using systemd, Docker, or Kubernetes.
- Managing model artifacts and versioning.
- Scheduling retraining and inference jobs.
- Implementing health checks, logging, and operational metrics.
- Documenting deployment, rollback, and maintenance procedures.
Operational Requirements
- Designing robust and explainable anomaly detection systems.
- Reducing false positives while ensuring meaningful early warning capability.
- Defining training data requirements and retention periods.
- Handling missing data, sensor errors, and seasonal operational patterns.
- Distinguishing between normal operational variation and real anomalies.
Desired Deliverables
- Recommended ML model architecture.
- Data pipeline design using Kafka and InfluxDB.
- Ubuntu deployment guide.
- Configuration files and service definitions.
- Model training and inference implementation.
- Alerting and anomaly scoring logic.
- Operational and maintenance documentation.
Requirements
Machine Learning Experience
Experience with machine learning solutions for anomaly detection.
Time-Series Analytics
Proficiency in analyzing time-series data for infrastructure systems.
Programming Skills
Experience with programming languages like Python, Go, Java, Scala, Rust, or C++.
Linux Operations
Knowledge of Ubuntu Linux configuration and operations.
Nice to Have
Familiarity with Docker or Kubernetes for deployment.
Experience in building REST APIs using frameworks like FastAPI or Spring Boot.
Benefits
Remote Work
Flexible remote work options available.
Learning Opportunities
Access to professional development and training resources.