A public healthcare institution in Belgium is expanding its data and AI capabilities to analyse large volumes of national healthcare data. As a Senior Data Scientist, you will create and industrialise machine learning solutions in Azure Databricks, using Python and PySpark to detect unusual billing patterns, prioritise risk and give healthcare experts explainable evidence for medical and administrative controls.
The mission
The Data Office combines data scientists, data engineers, architects, analysts, project managers and healthcare experts. Its current programme develops anomaly detection solutions that analyse Belgian healthcare data, identify unusual billing patterns and prioritise high-risk cases for review. The platform uses Azure Databricks, PySpark, Delta Lake, Unity Catalog, MLflow, Databricks Workflows, Azure Data Factory, Azure DevOps and Azure Storage. The work supports regulated healthcare controls, so model accuracy, explainability, privacy, security and maintainability are part of the delivery outcome.
You will lead analytical solutions from use-case selection with domain and policy stakeholders through feature engineering, model development, evaluation and production integration. Your scope covers supervised and unsupervised methods, including anomaly detection and risk modelling, classification and clustering on large-scale data. You will connect model outputs to data and ML pipelines, define operational-value measures and make results understandable to non-technical users. You will also set standards and coach colleagues while remaining accountable for delivery quality.
Your responsibilities
- Lead the design and industrialisation of scalable machine learning solutions in Azure Databricks, taking models from exploration to maintainable production use.
- Develop supervised and unsupervised models for anomaly detection, classification and risk scoring that help prioritise medical and administrative controls.
- Translate healthcare and policy requirements into data-driven solutions by selecting high-value use cases with domain experts, analysts, engineers and architects.
- Establish evaluation approaches that balance predictive accuracy, explainability and operational value, and communicate findings clearly to stakeholders.
- Strengthen feature engineering, data and ML pipelines, integration and MLOps using PySpark, MLflow and Microsoft Azure services.
- Define standards for security, privacy, governance and responsible AI, while coaching team members in quality, testing and maintainable delivery.
Your profile
Essential skills
- Bring at least 10 years of experience in data science, machine learning or advanced analytics, with ownership of complex analytical delivery.
- Build and assess anomaly detection, classification, clustering, risk scoring and explainable machine learning models.
- Work fluently with Python, SQL, Azure Databricks, PySpark and the Microsoft Azure ecosystem for large-scale data processing.
- Use MLflow, Delta Lake, Unity Catalog and Databricks Workflows, and apply Git, automated testing, CI/CD and MLOps in production-oriented environments.
- Convert business needs into scalable, secure and maintainable solutions while working pragmatically, autonomously and in an agile setting.
- Communicate clearly with healthcare and policy stakeholders, take technical ownership and provide leadership or coaching to colleagues.
Languages
- Dutch: C2 or native-level proficiency if it is your mother tongue, with B1 passive understanding if it is your second national language.
- French: C2 or native-level proficiency if it is your mother tongue, with B1 passive understanding if it is your second national language.
- English: B2 professional working proficiency.
Education
- Master’s or Ph.D. in Data Science, Computer Science, AI, Statistics, Mathematics or Engineering, or equivalent professional experience.