MLflow, Kubeflow, BentoML, data drift: master the deployment and monitoring of your ML models in production. A complete MLops guide by Smile experts.
The vast majority of machine learning models developed by data scientists never make it into production. Among those that do, a large proportion silently degrade their performance in the weeks following deployment, without anyone detecting it in time.
This isn't a problem with model quality. It's an engineering problem. MLOps is the discipline that solves it. This guide explains how to structure a reliable deployment architecture and how to maintain model performance in production over time.
- ML failure: The majority of projects fail when moving to production, as the infrastructure is much more complex than the algorithm itself (Google Research).
- MLOps speed: Mature AI engineering reduces model deployment from several weeks to a few hours (Gartner, 2024).
What is MLOps?
Machine Learning Operations ( MLOps ) encompasses the practices, processes, and tools used to deploy, monitor, and maintain machine learning models in production to ensure the reliability and reproducibility of predictions. It applies DevOps and DataOps principles to the ML model lifecycle, while adding AI-specific features such as training data management, model versioning, drift monitoring, and automated retraining.
The distinction with DataOps is important. DataOps automates the data pipelines that produce and transform data. MLOps automates the lifecycle of models that consume this data to produce predictions. The two are complementary and interdependent: a machine learning model is only as reliable as the data that feeds it.
The ML lifecycle: from notebook to production
An ML project goes through six stages before producing value in production.
- Exploration and experimentation : the data scientist explores data, tests algorithms, and evaluates performance on historical datasets.
- Data preparation : cleaning, transformation and feature engineering to produce training data
- Training : training the model on prepared data, with tracking of metrics and hyperparameters
- Validation : performance evaluation on independent test data, comparison with previous versions
- Deployment : putting the model into production, exposing it via an API, or integrating it into a processing pipeline
- Monitoring : continuous monitoring of performance and input data to detect deviations
The 3 levels of MLOps maturity
Google has defined three levels of maturity that structure the progression of an organization.
Level 0 : Manual process. Data scientists train and deploy models manually. No CI/CD, no automated monitoring. Suitable for one-off projects.
Level 1 : Automation of the ML pipeline. Training is automated and triggered by new data or a schedule. Retraining is continuous. Basic monitoring is in place.
Level 2 : CI/CD pipeline automation. The ML pipeline code is automatically versioned, tested, and deployed. Models are promoted to production through an automated validation process.
The key components of an MLOps architecture
Feature store: centralized feature management
The feature store is the central repository that allows data to be collected, stored, versioned and served as features to training and production models.
It ensures consistency between features used in training and those used in production, eliminating training-serving skew, one of the most frequent causes of performance problems in production.
Experiment tracking: reproducibility of experiments
Experiment tracking automatically records all parameters, metrics, and artifacts from each training experiment and provides valuable information for comparing experiments.
It helps teams reproduce any version of a model and audit the history of experiments. Without experiment tracking, finding out which model produced which results quickly becomes impossible.
Model registry: model versioning and governance
The model registry is the centralized catalog of all trained models, with their version, performance metrics, status (staging, production, archived) and metadata.
It is the equivalent of the data catalog for ML models and constitutes the central piece of model governance in any production IT system.
CI/CD for ML models
CI/CD ML automatically triggers the training, validation, and deployment of a new model whenever the pipeline code changes or new data arrives. This eliminates risky manual interventions during deployments and ensures that the code in production has always been validated. Each model promoted to production has passed a battery of automated tests: superior performance compared to the previous model, no bias detected, and compliance with business constraints.
Model serving: deployment and inference
Model serving is the infrastructure that exposes trained models to respond to real-time queries (online serving) or process batches of data (batch inference). It manages the scalability, latency, and availability of the model in production to ensure an optimal user experience.
Monitoring and observability of models in production
This is the most often overlooked and most critical dimension of MLOps. A model deployed without monitoring is a blind system that silently degrades the user experience and end-user trust.
Data drift: drift in input data
Data drift occurs when the distribution of data received in real time in production deviates from that of the training data. User behavior changes, data sources evolve, and market conditions shift. If the model was trained on 2023 data and the patterns in 2026 are different, its predictions will degrade.
Model drift: performance degradation
Model drift (or concept drift) occurs when the relationship between the input variables and the target variable changes in the real world. The model was correct yesterday, but it's no longer correct today. Without monitoring performance metrics in production, these performance issues can go unnoticed for weeks.
Automated retraining strategies
Three retraining strategies are commonly used.
Scheduled retraining : the model is updated at fixed intervals (weekly, monthly) regardless of drift signals. Simple to implement, but not always optimal.
Drift-triggered retraining : a monitoring system detects significant drift and automatically triggers retraining. More responsive but requires a well-calibrated trigger threshold.
Continuous retraining : the model is continuously retrained on new data using online learning techniques. Reserved for use cases where model freshness is critical.
MLOps reference tools in 2026
MLflow: experiment tracking and model registry
MLflow is both an experience observability platform and the open source reference for experiment tracking, model registry and model deployment.
Its ease of integration with the main ML frameworks (scikit-learn, TensorFlow, PyTorch) and its compatibility with major clouds make it the recommended starting point for any organization structuring its MLOps practices.
Kubeflow: ML pipelines on Kubernetes
Kubeflow is an open-source framework designed for cloud-native applications that orchestrates machine learning pipelines on Kubernetes. It covers the entire lifecycle: data preparation, distributed training, hyperparameter tuning, and deployment. Its Kubernetes-native architecture makes it particularly well-suited for organizations that already have a Kubernetes infrastructure in production.
BentoML and Seldon: model serving
BentoML is an open-source model serving framework that simplifies the packaging and deployment of ML models as REST APIs. Seldon Core is an advanced serving platform on Kubernetes, with A/B testing, canary deployment, and model explainability capabilities.
Managed cloud services
AWS SageMaker, Google Vertex AI, and Azure ML offer managed MLOps platforms that cover the entire ML lifecycle. They reduce operational complexity but impose a dependency on the cloud provider and potentially high costs at scale.
Comparative table
Tool | Category | Open source | Key point |
MLflow | Tracking and registry | ✅ | Simplicity, broad ecosystem |
Kubeflow | ML Pipelines | ✅ | Native Kubernetes, scalability |
BentoML | Model serving | ✅ | Simplified packaging |
Seldon Core | Advanced Serving Model | ✅ | A/B testing, explainability |
SageMaker | Complete platform | ❌ | AWS Integration |
Vertex AI | Complete platform | ❌ | GCP Integration |
Azure ML | Complete platform | ❌ | Azure Integration |
MLOps and DataOps: Essential Synergies
MLOps and DataOps are not competing disciplines. They are two sides of the same coin in a data-mature organization.
Data pipelines feed the ML models
A machine learning model is only as reliable as the data that feeds it. DataOps data pipelines produce the training features and production data that the model consumes.
Any serious MLOps observability tool relies on three types of metrics to detect drift: direct performance metrics, input data distribution metrics, and prediction distribution metrics to detect abnormal changes in system performance.
Common governance of data and models
Data governance defines the rules that apply to data. Model governance defines who deploys what, with what validation, and according to what approval process. Both must be coordinated in a coherent data and AI governance program at the organization's IT system level.
To learn more about the DataOps practices that underpin the industrialization of ML models, consult our MLOps guide with Smile .
Smile and MLOps: feedback from experience
At Smile, we support organizations in the industrialization of their ML models in production. Our expertise covers the entire end-to-end MLOps lifecycle : structuring training pipelines, setting up the feature store, deploying the model registry with MLflow, serving with BentoML, monitoring drift, and retraining strategies.
Our approach is consistently pragmatic. We begin by assessing the organization's MLOps maturity level and then support the implementation of practices tailored to the specific context. A solo data scientist needs to implement different tools than a team of 20 ML engineers.
We work on both sovereign on-premise infrastructures and managed cloud platforms, with particular attention to GDPR compliance and data sovereignty issues in production AI systems.
Do you want to industrialize your ML models with an MLOps approach? Discover our DataOps and MLOps practices .
Frequently Asked Questions about MLOps
What is the difference between MLOps and DataOps?
DataOps automates data pipelines: ingestion, transformation, data quality, and delivery to consuming systems. MLOps automates the lifecycle of machine learning models: training, validation, deployment, and monitoring.
The two are complementary. A reliable DataOps pipeline is a prerequisite for a high-performing MLOps system: if the training data is of poor quality, the model will be of poor quality regardless of the sophistication of the MLOps pipeline.
Do you need DevOps skills to practice MLOps?
An understanding of DevOps concepts (CI/CD, containerization, orchestration) is helpful but not essential to get started. Tools like MLflow, BentoML, and managed cloud platforms are designed to be accessible to data scientists without in-depth DevOps expertise. For large-scale deployments with high availability requirements, close collaboration with ML engineers or platform engineers is necessary.
How can we detect when a model is deviating in production?
Three types of metrics can be used to detect drift. Direct performance metrics (accuracy, F1, AUC) are used when labels are readily available in production. Input data distribution metrics (statistical tests such as KS test, PSI) are used to detect data drift before performance degrades.
Prediction distribution metrics are used to detect abnormal changes in model outputs. A dedicated observability platform like Monte Carlo or Elementary automates this monitoring without manual intervention.
Which MLOps tool should I choose to get started?
MLflow is the recommended starting point for the vast majority of teams. It is open source, easy to install, compatible with all major ML frameworks, and covers essential needs: experiment tracking, model registry, and basic deployment.
Once the fundamental practices are in place, adding Kubeflow for orchestration or BentoML for advanced serving is done naturally as needed.