Incremental Drift-Aware Learning: Simulate drift-aware learning using Simulink model template
A new model template for incremental drift-aware learning is available on the Simulink® start page under Statistics and Machine Learning. It combines the use of the following blocks into a drift-aware incremental learning workflow in Simulink:
IncrementalClassificationLinear Fit
Detect Drift
Per Observation Loss
The model templates are compatible with all incremental fit blocks in Statistics and Machine Learning Toolbox™, except the IncrementalNaiveBayes blocks. For an example of usage, see Configure Simulink Template for Drift-Aware Incremental Learning.
Incremental Naive Bayes Blocks: Simulate a model and generate code in Simulink
You can now integrate the fit and predict functions of
incremental naive Bayes objects into Simulink using the new incremental naive Bayes fit and predict blocks. Import a
model and configure data types using the Block Parameters dialog box.
IncrementalClassificationNaiveBayes Fit — This block fits a configured
naive Bayes classification model for incremental learning (incrementalClassificationNaiveBayes) to new data.
IncrementalClassificationNaiveBayes Predict — This block returns classified labels and predicted scores (optional) for new data using a naive Bayes classification model for incremental learning trained in Simulink.
The blocks support C/C++ code generation and fixed-point conversion. You can generate C and C++ code using Simulink Coder™, and design and simulate fixed-point systems using Fixed-Point Designer™.
Incremental Learning: Calculate per observation loss in Simulink
You can now integrate the perObservationLoss function into Simulink using the new Per Observation
Loss block. Import an incremental model (e.g. incrementalClassificationLinear) and configure data types using the Block
Parameters dialog box.
The block supports C/C++ code generation and fixed-point conversion. You can generate C and C++ code using Simulink Coder, and design and simulate fixed-point systems using Fixed-Point Designer.
Generate C/C++ code for performing incremental learning using multiclass classification model functions (requires MATLAB Coder)
You can generate C/C++ code that trains a multiclass classification model in real
time, as batches of observations become available. You can save an incrementalClassificationECOC model by using saveLearnerForCoder, and then load the model in the entry-point function
for code generation by using loadLearnerForCoder (see Introduction to Code Generation). These incremental learning functions
support code generation:
The fit
function fits the model to incoming data by updating the coefficients.
The predict function predicts responses given incoming data.
The loss
function returns the classification loss of the model for incoming data.
The updateMetrics function updates performance metrics given incoming
data.
The updateMetricsAndFit function updates performance metrics and fits
the model to incoming data.
Incremental Fit Blocks: Reset ports to reset the model parameters
A new reset inport in incremental fit blocks, such as IncrementalClassificationLinear Fit, allows resetting of the learned parameters and hyperparameters of a trained incremental learner. Provide a reset signal to the inport upon drift detection during drift-aware learning. The following blocks have a reset inport:
IncrementalClassificationLinear Fit
IncrementalClassificationECOC Fit
IncrementalClassificationKernel Fit
IncrementalClassificationNaiveBayes Fit
IncrementalRegressionLinear Fit
IncrementalRegressionKernel Fit
Update Metrics
Example of In-Place Model Update for Offline Models
A new example, In-Place Model Update of Offline Linear Model Using IncrementalClassificationLinear Predict Block, shows how to update offline models like linear classification using the incremental Simulink block of the model. Applying in-place model updates reduces the effort required to regenerate, redeploy, and reverify C/C++ code when you retrain a model with new data or configurations.
Machine Learning Apps: Export trained model to MATLAB Coder to generate C/C++ code
After you train a supported machine learning model in the Classification Learner or Regression Learner app, you can now export the model directly to MATLAB® Coder to generate C/C++ code. On the Learn tab, in the Export section, click Export Model and select Export Model to Coder. For more information and a list of supported models, see Export Classification Model to MATLAB Coder to Generate C/C++ Code and Export Regression Model to MATLAB Coder to Generate C/C++ Code.
Machine Learning Apps: Interpret trained models using permutation importance plot
The permutation importance plot allows you to interpret which predictors have the largest (or smallest) average impact on model predictions. The plot displays a horizontal bar graph of mean permutation importance values for the top predictors in the model, sorted from highest to lowest mean importance value.
For more information, see the Algorithms section of permutationImportance.
Machine Learning Apps: Export kernel approximation models to Simulink
You can now export the following model types from the learner apps to Simulink:
Regression Learner — All kernel approximation regression models
Classification Learner — All kernel approximation models except those trained on multiclass data
Machine Learning Apps: Check coder model size and sort models by size
After you train a supported model in Classification Learner or Regression Learner,
you can check the approximate size of the model (in bytes) in C/C++ code generated by
MATLAB
Coder. Select the model in the Models pane, and then select
the Model size (Coder) value in the Training
Results section of the model Summary tab. The app
displays a coder model size of NaN for model types that are not
supported for code generation. For more information on coder model size and a list of
supported model types, see learnersize,
Export Classification Model to MATLAB
Coder to Generate C/C++ Code, and Export Regression Model to MATLAB
Coder to Generate C/C++ Code.
To compare the coder model size of multiple models, use the results table. On the Learn or Test tab, in the Plots and Results section, click Results Table. Click the Select columns button at the top right of the table. In the Select Columns to Display dialog box, select Coder Model Size (bytes) and click OK. In the table, sort the models by clicking the arrows for the Coder Model Size (bytes) column.
You can also sort trained models by size in the Models pane. Select Model Size (Coder) or Model Size (Compact) in the Sort by list.
Classification Learner App: Compare receiver operating characteristic (ROC) curves
In the Classification Learner app, you can now compare the ROC curves of up to four trained models in a single plot. Create a Compare ROC Curves plot in one of two ways:
In the Plots and Results section of the Learn tab, click the arrow to open the gallery, and then click Compare ROC curves.
In the Plots and Results section of the Test tab, click Compare ROC curves.
Use the Models options to display the ROC curves of up to four models. You can also display macro, micro, and weighted average ROC curves for each model.
Distribution Fitter App: New bin rules for histogram of a data set
The new default rule for data comprising of only integers is Bins centered on integers. Previously, the default rule was the Freedman - Diaconis rule. The new default ensures a more accurate histogram for data containing only integers. You can now also select the default rule as Automatic in the Set Default Bin Rules dialog box. If selected, the default bin rule for data comprising of only integers is Bins centered on integers and the default bin rule for other types of data is the Freedman - Diaconis rule. For more information, see Set Bin Rules.
Cluster Data Live Editor Task: Specify dendrogram plot options
When you perform hierarchical clustering in the Cluster Data Live Editor Task, you can now:
Color a dendrogram plot according to cluster assignments.
Arrange dendrogram nodes in optimal leaf order. The optimal leaf order for a binary tree maximizes the sum of the similarities between adjacent leaves by flipping tree branches without dividing the clusters.
Display a dashed line that shows where the tree is cut to produce each colored leaf node assignment in a dendrogram plot.
Show a marker for each leaf node. Click a marker to display information about the row numbers (and cluster assignments, when applicable) for the leaf node.
The task automatically generates code that becomes part of your live script.
To use the Cluster Data task in the Live Editor, select Task > Cluster Data on the Live Editor tab. Alternatively, in a code block in a live script, begin typing the task name and then select the task from the suggested command completions.
Experiment Manager: Randomly sample parameter values from probability distributions
In the Experiment Manager (Deep Learning Toolbox) app, you can randomly sample parameter values for your experiment from probability distributions using the Random Sampling strategy. On the experiment definition tab, in the Parameters section:
Select the Random Sampling strategy.
For each parameter in the table, choose a probability distribution from the list.
Set the value for each distribution property.
Then, in the Random Sampling Options section, specify the number of trials.
Experiment Manager: Tune additional hyperparameters
When you create an experiment from a trained model in the Classification Learner or Regression Learner app, you can now add hyperparameters to tune from a suggested list. In Experiment Manager, on the tab for your experiment, click the arrow next to Add in the Hyperparameters section and select Add From Suggested List. In the Add From Suggested List dialog box, select the hyperparameters you want to add and click Add. For more information, see Export Model from Classification Learner to Experiment Manager and Export Model from Regression Learner to Experiment Manager.
Functionality being removed or changed
Regression Learner App: Train an SVM model with an automatic box constraint
Behavior change
In Regression Learner, when you train a support vector machine (SVM) model with
Box constraint mode set to Auto, the
app now sets the box constraint equal to 1 when you select a
non-Gaussian kernel function. For a Gaussian kernel function, the app uses a
heuristic procedure to set the box constraint. In previous releases, the app used the
same heuristic procedure for all kernel functions. For more information, see SVM Model Hyperparameter Options.
Incremental Clustering: Create a k-means model for incremental clustering
Create an incrementalKMeans model object for incremental k-means
clustering with a fixed number of clusters by using the
incrementalKMeans function. After creating the model, you can
train it, assign cluster indices, and calculate performance metrics either per
individual observation or per specified batch size, in real time.
The fit
function fits the model by updating the cluster parameters, given an incoming
batch of data. After training, if the model is warm, fit
optionally returns cluster indices for the incoming observations.
When the model is warm, the updateMetrics function computes performance metrics.
The assignClusters function returns cluster indices and cluster
distances.
The reset
function resets all learned parameters and estimated hyperparameters of the
model.
Incremental Dynamic Clustering: Create a k-means model for incremental dynamic clustering
Create an incrementalDynamicKMeans model object for incremental
k-means clustering with a dynamic number of clusters by using the
incrementalDynamicKMeans function. After creating the model, you
can train it, assign cluster indices, and calculate performance metrics either per
individual observation or per specified batch size, in real time.
The fit
function fits the model by updating the cluster parameters, given an incoming
batch of data. After training, if the model is warm, fit
optionally returns cluster indices for the incoming observations.
When the model is warm, the updateMetrics function computes performance metrics.
The assignClusters function returns cluster indices and cluster
distances.
The reset
function resets all learned parameters and estimated hyperparameters of the
model.
Neural Networks: Specify custom neural network architecture (requires Deep Learning Toolbox)
For neural networks with complex architecture (such as, neural networks with skip
connections), specify the architecture for the fitrnet and
fitcnet
functions using the Network name-value argument with a layer array
or dlnetwork (Deep Learning Toolbox) object.
For a list of available layers, see List of Deep Learning Layers (Deep Learning Toolbox).
Synthetic Data Generation: Test similarity of data sets using k-nearest neighbor statistic
After you synthesize data, you can use the knntest
function to test whether the new data set comes from the same distribution as the
original data set. The function indicates how well separated the two data sets are,
based on whether the observations' nearest neighbors tend to be in the same set as the
observations.
cvpartition Function: Create group k-fold
cross-validation
The cvpartition function supports the creation
of group k-fold cross-validation partitions. To create this type of partition, specify
GroupingVariables=, where
groupingVariablesgroupingVariables contains one or more grouping variables. You
cannot specify grouping variables when you use a stratification variable stratvar.
The cvpartition object has two additional properties: IsGrouped
and IsStratified. When either property value is 1
(true), you can use the summary
object function to display summary information about the partition.
Quantile Regression: Optimize, cross-validate, or reduce the size of quantile regression models
You can optimize, cross-validate, or reduce the size of quantile regression models
created using the fitrqlinear
and fitrqnet functions.
To optimize the hyperparameters of a quantile regression model, specify the
OptimizeHyperparameters name-value argument in the call
to fitrqlinear or fitrqnet.
To cross-validate a quantile regression model, specify one of these
name-value arguments in the call to fitrqlinear or
fitrqnet: CrossVal,
CVPartition, Holdout,
KFold, or Leaveout. Alternatively,
create a full quantile regression model and then use the crossval object function. The resulting model is a RegressionPartitionedQuantileModel object. You can estimate the
quality of the model by using the kfoldPredict, kfoldLoss, and kfoldfun object functions.
To reduce the size of a quantile regression model, create a full quantile
regression model and then use the compact object function. The resulting model is a CompactRegressionQuantileLinear or CompactRegressionQuantileNeuralNetwork object.
Additionally, you can:
Train a quantile linear regression model in parallel by using the
UseParallel name-value argument in the call to
fitrqlinear.
Specify the prediction values for observations with missing predictor values
by using the PredictionForMissingValues name-value
argument in the call to the loss and
predict object functions.
fitcknn Function: Use fast Euclidean distances
In the call to the fitcknn function, you can specify the
"fasteuclidean" or "fastseuclidean" distance
metric (Distance) to accelerate the
computation of Euclidean distances using a cache and an alternative algorithm. Set the
size of the cache by using the CacheSize
name-value argument in the call to fitcknn, or the CacheSize
property of the model object after you create it.
GPU Support: Specify GPU arrays for Gaussian kernel models (requires Parallel Computing Toolbox)
The following functions now support GPU arrays, enabling you to execute the functions on a GPU:
The fitrkernel and fitckernel functions accept GPU array input arguments.
Most object functions of the models RegressionKernel and ClassificationKernel now support GPU array input arguments. The
functions that do not are incrementalLearner, lime, and shapley.
For a full list of Statistics and Machine Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays).
GPU Support: Specify GPU arrays for one-class support vector machine (SVM) models (requires Parallel Computing Toolbox)
You can now fit a one-class support vector machine (SVM) model for anomaly detection
on a GPU by supplying GPU array predictor data to the ocsvm
function. The isanomaly
object function of OneClassSVM
supports GPU array input arguments so that the function can execute on a GPU. The
incrementalLearner object function does not support GPU arrays.
For a full list of Statistics and Machine Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays).
Quantile Regression: Prevent quantile crossing
A new example, Regularize Quantile Regression Model to Prevent Quantile Crossing, shows how to use regularization to prevent quantile crossing in quantile regression models.
Quantile Regression: Create prediction intervals using split conformal prediction
A new example, Create Prediction Intervals Using Split Conformal Prediction, shows how to use split conformal prediction (SCP) to create prediction intervals with a quantile regression model.
Functionality being removed or changed
Hyperparameter Optimization: Compute in serial when parallel optimization is unavailable
Behavior change
Starting in R2025a, bayesopt and supported fitting functions
run computations serially when you specify parallel hyperparameter optimization and
the software is unable to open a parallel pool. For a list of supported fitting
functions, see AggregateBayesianOptimization.
To specify parallel hyperparameter optimization, set
UseParallel=true for the bayesopt function or
set UseParallel=true within the
HyperparameterOptimizationOptions argument of a supported
fitting function.
In previous releases, the software throws an error if it cannot open a parallel pool under these conditions.
Censored Linear Regression: Create a CensoredLinearModel or
CompactCensoredLinearModel object to perform linear regression on
censored data
To perform linear regression on censored data, create a CensoredLinearModel object using the fitlmcens
function. After creating the object, you can use it to predict responses, analyze model
coefficients and residuals, and create plots.
Use the compact function to create a CompactCensoredLinearModel object.
Use the plotResiduals, plotSlice, and plotPartialDependence functions to create visualizations of the
fitted model and its summary statistics.
Use the predict, feval, and
random
functions to generate predictions.
Use the partialDependence, coefCI, and
coefTest functions
to analyze the model predictors and fitted coefficients.
Design of Experiments: Create a taguchiDOE object to generate a
Taguchi experiment design
Generate a Taguchi experiment design by using the taguchiDOE
function to create an object of the same name. The object properties include information
about the design type, model, and factors used to generate the design. Use the taguchiTypes
function to return a list of valid Taguchi design types for a set of factors and levels.
After creating a taguchiDOE object, you can use the fitlm
function to fit a linear model to the design points. Use the snr function
with response data to calculate signal-to-noise ratios for each factor-level
combination, and plot signal-to-noise ratios using the plotsnr
function.
Calculate and plot profile loglikelihoods for nonlinear models
Two new object functions, profileLikelihood and plotProfileLikelihood, are available for NonLinearModel objects. Use the profileLikelihood
function to calculate the profile loglikelihood and likelihood-ratio confidence interval
for a coefficient in a nonlinear model. Use the
plotProfileLikelihood to plot the profile loglikelihood, Wald
approximation, coefficient estimate, and likelihood-ratio and Wald confidence
intervals.
Include missing counts, control cross-tabulation table format, and specify table grouping variables
The crosstab function has two new name-value
arguments, IncludeMissingGroups and
OutputFormat, which allow you to count missing values and
specify the format of the cross-tabulation table, respectively. The new input argument
datatbl lets you specify table grouping variables.
Design of Experiments: Use extensions to the Gage repeatability and reproducibility
study gagerr function
gagerr has several extensions and new syntaxes:
You can now specify input data as a table that contains a variable for the measurements, a variable for the parts, and (optionally) a variable for the operators.
When you specify input data as a table, or set the new name-value argument
OutputFormat as "table", the
results output is returned as a table instead of a
matrix.
You can now specify the SigmaMultiplier name-value
argument, which gagerr uses to calculate the study
variance and the precision-to-tolerance ratio. The default value
5.15 is used in previous releases, and corresponds to the
number of standard deviations that span the middle 99% of a normally
distributed population.
You can now specify DisplayANOVA="on" to display the
ANOVA results in a figure window.
The Model name-value argument has a new value,
"part-nested", which specifies an ANOVA model where the
parts are nested within the operators.
Pearson Distribution: Create a pearsonDistribution object function
and evaluate Pearson inverse with pearsinv
You can now use makedist to create a PearsonDistribution object. The object can be used with generic functions
such as pdf, cdf and
random. The new distribution specific function pearsinv can
be used to compute the Pearson inverse cumulative distribution.
Empirical Distribution: Create an empiricalDistribution object
function
You can now fit data using fitdist to create a EmpiricalDistribution object. The object can be used with generic functions
such as pdf, cdf and
random.
Improved performance of binomial inverse algorithm for large number of trials
The binoinv function shows improved
performance when computing the binomial inverse cdf for number of trials greater than
100. For example, the binomial inverse computation in this code is about 14x faster than
in the previous release:
function timingTest() nval = 500; pval = 0.6; yval = 0.6; x = binoinv(yval,nval,pval); end timeit(@timingTest)
The approximate execution times are:
R2025a: 0.0015 s
R2024b: 0.022 s
The code was timed on a Windows® 11, Intel®
Xeon® E5-1650 v3 CPU @ 3.5 GHz test system by calling the function
timeit.
controlchart Function: Return graphics handles to individual
plots
The controlchart function can now return
graphics handles to the individual plots as an array of Axes objects.
For details, see the function reference page.
confusionchart Function: Customize confusion chart
display
You can now customize the titles and labels of ConfusionMatrixChart objects:
Specify labels for the row (y-axis) and
column (x-axis) values using the RowDisplayLabels and ColumnDisplayLabels properties.
Interpret text labels, such as titles and axis labels, with TeX markup,
LaTeX markup, or no markup using the Interpreter property.
Display labels for the true positive rate (TPR), false negative rate (FNR),
positive predictive value (PPV), and false discovery rate (FDR) on the row and
column summaries using the SummaryLabelsVisible property.
To create a confusion chart, use the confusionchart function. For more information about how to set these
properties, see ConfusionMatrixChart Properties.
Functionality being removed or changed
multivarichart Function: Return graphics handles to
individual plots as Axes objects
Behavior change
The multivarichart function now returns
graphics handles to the individual plots as an array of Axes
objects. For details, see the function reference page.