Senior Data Scientist
At Global IT Hub, we are more than a technology organization. We are a community of innovators, problem-solvers, and digital pioneers dedicated to accelerating Atlas Copco Group’s business potential through world-class technology solutions. Join us to shape the future of digital transformation, work on global challenges, and create impact for businesses and customers worldwide.
CLASSICAL MACHINE LEARNING | JOB DESCRIPTION
|
Role purpose Lead the design, validation, deployment and lifecycle ownership of production-grade classical machine learning solutions that convert business problems into measurable outcomes. The role combines deep algorithmic expertise, disciplined CRISP-DM execution, MLOps practices and senior technical leadership. |
Role at a glance
|
Team |
AI Team, Global IT Hub |
|
Primary stakeholders |
Business product owners, IT, data engineering, analytics and domain teams |
|
Experience |
5+ years in data science / machine learning, including 3+ years owning production ML solutions |
|
Education |
Bachelor’s or Master’s degree in Data Science, Statistics, Mathematics, Computer Science, Engineering or a related quantitative field |
Key responsibilities
- Partner with stakeholders to frame ambiguous business needs as well-defined machine learning problems, including decision context, baselines, constraints, risks and measurable success criteria.
- Lead end-to-end delivery using CRISP-DM: business understanding, data understanding, data preparation, modeling, evaluation and deployment, with iteration and traceability across phases.
- Design, train and compare classical supervised and unsupervised learning models; select the simplest approach that meets performance, interpretability, latency, scalability and cost requirements.
- Own feature engineering, data leakage prevention, sampling strategy, cross-validation, hyperparameter optimization, threshold selection, calibration and robust error analysis.
- Build reliable evaluation frameworks using fit-for-purpose metrics, statistical tests, confidence intervals, business-cost functions and champion-challenger comparisons.
- Take models from experimentation to production with engineering teams, covering packaging, APIs or batch scoring, versioning, CI/CD, monitoring, retraining and rollback.
- Own post-production performance, including drift detection, data-quality controls, incident analysis, model refresh decisions and communication of model limitations.
- Apply explainability and responsible AI practices through model cards, reproducibility, bias and fairness checks, audit trails, privacy-aware design and human oversight.
- Set technical standards, conduct design and code reviews, mentor data scientists, create reusable components and influence the classical ML roadmap.
- Communicate recommendations, uncertainty, assumptions and trade-offs clearly to technical teams, business leaders and governance forums.
Mandatory Technical Skills
- Advanced Python and SQL, with strong hands-on use of NumPy, Pandas, scikit-learn, SciPy and statsmodels; ability to write modular, tested and maintainable code.
- Deep knowledge of regression and classification: linear and logistic regression, regularization, decision trees, Random Forest, Extra Trees, XGBoost, LightGBM, CatBoost, support vector machines, k-nearest neighbors, Naive Bayes and appropriate ensemble methods.
- Strong understanding of unsupervised learning: k-means and hierarchical clustering, DBSCAN, PCA and other dimensionality-reduction methods, anomaly or outlier detection and segmentation evaluation.
- Practical knowledge of forecasting and statistical modeling, including ARIMA/SARIMA, exponential smoothing, trend and seasonality, lag features, back-testing and prediction intervals.
- Expertise in feature selection and engineering, missing-value treatment, categorical encoding, class imbalance, resampling, pipeline design and prevention of target leakage.
- Rigorous evaluation skills across ROC-AUC, PR-AUC, precision, recall, F1, log loss, RMSE, MAE, MAPE/WAPE, lift, gain, calibration and business-impact metrics, selected according to the use case.
- Hands-on optimization using grid, random or Bayesian search, with sound cross-validation design for grouped, temporal and imbalanced datasets.
- Practical MLOps experience with experiment tracking, model registry, source control, automated testing, Docker, CI/CD, deployment and model or data monitoring.
- Working knowledge of explainability techniques such as feature importance, permutation importance, partial dependence and SHAP, including their limitations.
- Experience handling large structured datasets and collaborating on scalable data pipelines in a cloud or distributed environment.
Senior-level capabilities
- Independently owns technical direction and delivery for complex ML use cases from discovery through production adoption.
- Challenges weak problem formulations and prevents unnecessary use of ML when rules, analytics or simpler statistical methods are more suitable.
- Makes defensible trade-offs among accuracy, interpretability, maintainability, latency and cost.
- Reviews experimental design and model evidence before approving production release.
- Mentors team members and raises standards for coding, documentation, reproducibility and stakeholder communication.
Good-to-have experience
- Domain exposure in manufacturing, industrial operations, supply chain, service, pricing, finance, sales or customer analytics.
- Optimization, causal inference, survival analysis, uplift modeling, recommendation systems, geospatial analytics or graph-based techniques.
- Azure Machine Learning, Databricks, Microsoft Fabric, MLflow, Spark and enterprise data sources such as SAP.
- Experience integrating classical ML models into applications, APIs, decision-support workflows or Power BI solutions.
- Awareness of deep learning and generative AI, with the judgment to benchmark them against classical ML rather than defaulting to higher-complexity approaches.
Core algorithm and method coverage
Candidates are expected to demonstrate depth in several categories and sound model-selection judgment across the full landscape. This is not a requirement to have used every library or algorithm in production.
|
Capability area |
Expected knowledge |
|
Supervised learning |
Regression, classification, tree ensembles, boosting, SVM, nearest-neighbor and probabilistic baselines. |
|
Unsupervised learning |
Clustering, dimensionality reduction, anomaly detection and segmentation validation. |
|
Time-dependent modeling |
Forecasting, temporal validation, lag and rolling features, seasonality and uncertainty. |
|
Data and features |
EDA, data-quality assessment, transformations, encoding, imbalance handling and leakage control. |
|
Evaluation |
Metric selection, baselines, validation design, statistical significance, calibration and business value. |
|
Production lifecycle |
Reproducibility, versioning, deployment, observability, drift, retraining and retirement. |
|
Governance |
Explainability, documentation, bias and fairness assessment, privacy, approvals and auditability. |
What you can expect from us
- Meaningful global business problems and access to cross-functional expertise.
- Freedom to choose fit-for-purpose methods and challenge unnecessary complexity.
- A collaborative environment focused on engineering quality, responsible AI and measurable value.
- Opportunities to mentor, build reusable capabilities and shape the enterprise AI practice.
At Global IT Hub, we believe exceptional talent drives exceptional transformation. As a strategic partner for next-generation digital transformation, we empower our people to innovate boldly, collaborate globally, and turn ideas into solutions that create lasting business value. Join a team where learning never stops, innovation is part of our DNA, and your contribution helps shape the future of a global industry leader.
Role Evolution – Next 3–4 Years
- Evolve from predictive modelling toward decision intelligence, combining prediction, causal analysis, simulation and optimization.
- Build hybrid AI solutions integrating classical ML with Generative AI, foundation models and AI agents.
- Establish systematic benchmarking across classical ML, deep learning and emerging AI models to select the best approach based on performance, explainability, scalability and cost.
- Drive automation of the end-to-end ML lifecycle, including feature engineering, experimentation, deployment, monitoring, drift detection, retraining and model retirement.
- Develop advanced model evaluation frameworks covering technical performance, business impact, robustness, explainability, fairness and operational reliability.
- Advance model observability and continuous learning, proactively managing data drift, concept drift and model degradation.
- Strengthen Responsible AI and model governance, including lineage, reproducibility, explainability, bias assessment, human oversight and auditability.
- Develop reusable ML frameworks, components and engineering standards to accelerate enterprise-wide AI development.