Experiment Manager: Monitor model performance while you run experiments
During training, the results table now displays the intermediate values for standard training and validation metrics for built-in training experiments. These metrics include loss, accuracy of classification experiments, and root mean squared error of regression experiments.
Experiment Manager: Inspect execution environment for each trial
In built-in training experiments, the new Execution Environment
column of the results table displays whether each trial runs on a single CPU, a single
GPU, multiple CPUs, or multiple GPUs. To show this information, click the show or hide
columns button
located above the results table and select
Execution Environment.
Experiment Manager: Restart multiple trials
Restart multiple trials of your experiment by opening the Restart
list, selecting one or more restarting criteria, and clicking Restart
. The restarting criteria include All
Canceled, All Stopped, All
Error, and All Discarded. For more information,
see Stop and Restart Training.
Experiment Manager: Reduce storage size by discarding unwanted results
To reduce the size of your experiments, you can now discard the results of trials that
are no longer relevant. In the Actions column of the
results table, click the Discard button
for a trial. Experiment Manager deletes the training
plot, confusion matrix, trained network, training information, and training output from
your project.
Experiment Manager: New properties for experiments.Monitor
objects
The experiments.Monitor object has these new properties:
InfoData — Information column values for trial
MetricData — Metric column values for trial
An experiments.Monitor object has the same properties and object
functions as a TrainingProgressMonitor object. Therefore, you can easily adapt your custom
training loop plotting code for use in an Experiment Manager setup script. For more information, see Prepare Plotting Code for Custom Training Experiment.
Deep Network Designer: New layer groups and colors
Deep Network Designer has improved layer colors and layer groups. Smaller layer groups mean you can easily find the layer you need.
Visualization: Monitor and plot custom training loop progress
Track custom training loop progress and generate training plots using a
TrainingProgressMonitor object.
Create a TrainingProgressMonitor object using the trainingProgressMonitor function. Use the
TrainingProgressMonitor object to:
Track training progress.
Display animated plots of metric values during training.
Display and monitor information values during training.
Log metric and information values.
After you create a TrainingProgressMonitor object, you can use these
object functions:
recordMetrics — Record metric values during training
updateInfo — Update information values during training
groupSubPlot — Group metrics in training plot
For more information, see Monitor Custom Training Loop Progress.
A TrainingProgressMonitor object has the same properties and object
functions as an experiments.Monitor object. Therefore, you can easily adapt your plotting code
for use in an Experiment Manager setup script. For more information, see Prepare Plotting Code for Custom Training Experiment.
Visualization: Compute and plot ROC curve
Compute and plot classification performance metrics including receiver operating characteristic (ROC) curves.
Use rocmetrics to evaluate the performance of classification models with
performance metrics. You can create a rocmetrics object by passing true
labels, classification scores, and class names. By default, rocmetrics
computes the true positive rates (TPR), false positive rates (FPR), and area under the ROC
curve for each class. Additionally, you can specify more supported performance metrics by
using the AdditionalMetrics name-value argument.
After you create a rocmetrics object, you can use these object
functions:
plot — Plot ROC or other classifier performance curves. The
plot function returns a ROCCurve graphics
object for each curve. You can control the appearance of each curve by modifying the
properties of the objects. For details, see ROCCurve Properties.
average — Compute performance metrics for an average ROC curve for
multiclass problems.
addMetrics — Compute additional classification performance
metrics.
For an example that shows how to use ROC curves to compare the performance of deep learning models, see Compare Deep Learning Models Using ROC Curves.
In R2022a, rocmetrics was introduced in Statistics and Machine Learning Toolbox™. You can now use rocmetrics without a Statistics and Machine Learning Toolbox license. Some options, for example, rocmetrics using
confidence intervals, still require a Statistics and Machine Learning Toolbox license. For more information, see rocmetrics.
Interpretability: Create Grad-CAM maps with 1-D convolutional networks for sequence and time-series data
Compute Grad-CAM interpretability maps for deep neural networks that are trained on
sequence and time-series data. To compute Grad-CAM interpretability maps, use the
gradCAM function. For an example that shows how to use a Grad-CAM map to
interpret the predictions of a network trained on time series data, see Interpret Deep Learning Time-Series Classifications Using Grad-CAM.
Complex Numbers: Train networks using complex-valued data
To pass complex-valued data to a neural network, you can use the input layer to split the complex values into their real and imaginary parts before the network passes the data to the subsequent layers. When the layer splits complex-valued data, the layer outputs the data in twice as many channels as the input data.
The ImageInputLayer, SequenceInputLayer, and FeatureInputLayer objects support splitting input data into its real and
imaginary components. To input complex-valued data into a network by splitting it into its
real and imaginary parts, set the SplitComplexInputs option of the
network input layer to 1 (true).
To input complex data into a network, the SplitComplexInputs option
must be 1.
For an example that shows how to train a network with complex-valued data, see Train Network with Complex-Valued Data.
Gaussian Error Linear Unit (GELU) Activation: Create and train networks with GELU activation
The Gaussian error linear unit (GELU) operation weights the input by its probability under a Gaussian distribution. This operation is given by
where erf denotes the error function.
To apply the GELU operation in a layer array or layer graph, use a geluLayer object.
To apply the GELU operation in a custom layer or deep learning model function, use the
gelu function.
Spatio-Temporal Data: 2-D average and max pooling layers support data with both spatial and time dimensions
AveragePooling2DLayer and MaxPooling2DLayer objects now support input data with these data formats:
One spatial dimension and a time dimension
Two spatial dimensions and a time dimension
For an example that shows how to create a 2-D CNN-LSTM network for speech classification tasks by combining a 2-D convolutional neural network (CNN) with a long short-term memory (LSTM) layer, see Sequence Classification Using CNN-LSTM Network.
Sequence Networks: Make predictions in parallel
The predict, classify, and activations functions now support making predictions in parallel for
networks with sequence input. To make predictions in parallel, set the
ExecutionEnvironment option to "parallel" or
"multi-gpu". Making predictions in parallel requires a Parallel Computing Toolbox™ license.
When you make predictions in parallel for networks with recurrent layers, the
SequenceLength option must be "longest" or
"shortest".
Networks with custom layers that contain State parameters do not
support making predictions in parallel.
LSTM Projected Layer: Perform LSTM operations with fewer learnable parameters
To compress a deep learning network by reducing the number of learnable parameters, you can use projected layers. A projected layer is a variant of a deep learning layer that enables compression by reducing the number of stored learnable parameters by replacing multiplications of the form , where W is a learnable matrix, with the multiplication , where Q is a projector matrix. Instead of storing W, the layer instead stores Q and . Projecting into a lower-dimensional space with Q typically requires less memory and can have similarly strong prediction accuracy.
To create a long short-term memory (LSTM) projected layer, use lstmProjectedLayer. For an example that shows how to train a network with an
LSTM projected layer and compare the prediction accuracy and number of learnable
parameters to a network without projection, see Train Network with LSTM Projected Layer.
Custom Layers: Define Custom Layer Learnable Parameter Initialization
When you create a custom layer, specify a custom learnable parameter initialization function by implementing a function with this syntax:
function layer = initialize(layer,layout1,...,layoutN) ... end
layer is an instance of the custom layer
and layout1,...,layoutN are networkDataLayout objects corresponding to each of the N
inputs of the layer.The software uses this initialize function when you include the
layer in a dlnetwork object or call the initialize function on the dlnetwork object.
For an example that shows how to define a custom layer that uses automatic learnable parameter initialization, see Define Custom Deep Learning Layer with Learnable Parameters.
For more information about defining custom layers, see Define Custom Deep Learning Layers.
Customization: Add, remove, and replace layers of dlnetwork
objects
Add, remove, and replace layers in dlnetwork objects using these functions:
addInputLayer — Add input layer to dlnetwork
object
addLayers — Add layers to dlnetwork object
removeLayers — Remove layers from dlnetwork
object
connectLayers — Connect layers in dlnetwork
object
disconnectLayers — Disconnect layers in dlnetwork
object
replaceLayer — Replace layer in dlnetwork
object
Customization: Plot and view summary of dlnetwork objects
Customization: Initialize dlnetwork objects using only input size
and format information
Initialize the learnable parameters of a dlnetwork object by passing a networkDataLayout object to the dlnetwork and initialize functions.
In previous versions, the dlnetwork and initialize
functions require you to pass examples of input data. You can now use
networkDataLayout objects, which only contain the data size and
dlarray format.
Create an unformatted networkDataLayout object using the syntax
layout = networkDataLayout(sz), where sz specifies
the size of the data. Create a formatted networkDataLayout object using
the syntax layout = networkDataLayout(sz,fmt), where
fmt specifies the dlarray format of the data.
Attention: Apply attention operation to dlarray input
The attention operation focuses on parts of the input using weighted multiplication
operations. To apply the attention operation to a set of queries, keys, and values, use
the attention function. You can specify the scale, dropout probability, and
padding and attention masks used in the operation.
You can use the attention function when you work with custom training loops and custom layers to implement:
Dot product attention
Scaled dot product attention
Multihead attention
Luong attention
For an example that shows how to use attention for sequence-to-sequence translation, see Sequence-to-Sequence Translation Using Attention.
Automatic Differentiation: Use more functions with dlarray
input
Use these functions with a dlarray object as input when you work with custom training loops and custom layers.
Attention — Apply the attention operation to a set of queries, keys, and values,
using attention.
GELU — Apply the GELU activation function to input elements using gelu.
Error function — Compute the error function of input elements using erf.
String – Convert dlarray objects to string
objects using string.
Categorical – Convert dlarray objects to
categorical objects using categorical.
For dlarray input, the string and
categorical functions convert to these respective data types only and
do not support automatic differentiation.
For a full list of functions that support dlarray input, see List of Functions with dlarray Support.
Function Layer: Accelerate layer functions
You can now specify that a layer function supports acceleration using dlaccelerate by setting the Acceleratable property to
1
(true) when you create a functionLayer. Setting Acceleratable to
1
(true) can improve the performance of training and inference
(prediction) when you use a dlnetwork object. For example, calling
predict on a dlnetwork object containing a number
of functionLayer objects in this test is about 3x faster than in the
previous release:
function timeFunctionLayer % Prepare input data X = dlarray(rand(500,500,3,"single","gpuArray"),"SSCB"); % Prepare a convolution and ReLU block convBlock = [convolution2dLayer(4,20); reluLayer()]; if version("-release") == "2022b" % Create a network using functionLayer ReLU with Acceleratable convBlock(2) = functionLayer(@(x) relu(x),Acceleratable=true); layers = repmat(convBlock,80,1); net = dlnetwork(layers,X); else % Create a network using functionLayer ReLU convBlock(2) = functionLayer(@(x) relu(x)); layers = repmat(convBlock,80,1); net = dlnetwork(layers,X); end % Time the predict function gputimeit(@()predict(net,X)) end
The approximate execution times are:
R2022a: 0.12 seconds
R2022b: 0.04 seconds
The code was timed on a Windows® 10, Intel®
Xeon® W-2133 @ 3.60 GHz test system with an NVIDIA® RTX A5000 GPU by calling the timeFunctionLayer
function.
Bayesian Neural Networks: Create and train Bayesian neural networks
A Bayesian neural network (BNN) is a type of deep learning network that uses Bayesian methods to quantify the uncertainty in the predictions of a deep learning network.
You can train a BNN to generate a distribution of weights and biases, rather than a single set. You can then use these distributions to measure the uncertainty of the network predictions.
To learn how to train a BNN using the Bayes by backpropagation method, see Train Bayesian Neural Network.
Background Dispatch: Use DispatchInBackground on thread
pools
You can now use background dispatch (asynchronous prefetch queuing) for reading training data from datastores on thread-based parallel pools.
To use background dispatch when training a network using
trainNetwork, set the DispatchInBackground
training option to 1
(true) using the trainingOptions function and open a thread-based parallel pool using
parpool("Threads").
To use background dispatch when training a network using a custom training loop,
create a minibatchqueue object and set the DispatchInBackground
property to 1
(true), and open a thread-based parallel pool using
parpool("Threads").
As an alternative to opening a thread-based parallel pool using
parpool("Threads"), you can set the default parallel environment on
your local machine from the MATLAB desktop Home tab, in the
Environment area, by selecting Parallel > Select Parallel Environment > Threads, or by calling parallel.defaultProfile("Threads").
Network Creation: Improved performance
Assembling a DAGNetwork object using assembleNetwork and assembling a dlnetwork object using
dlnetwork show improved performance. Improvements are greater for networks
that contain more layers. For example, assembling a DAGNetwork object
in this test is about 1.9x faster than in the previous release:
function timeAssembleNetwork % Create a layer graph lgraph = layerGraph(resnet50); % Time assembling a DAGNetwork object from the layer graph tic net = assembleNetwork(lgraph); toc end
The approximate execution times are:
R2022a: 3.80 seconds
R2022b: 2.01 seconds
Assembling a dlnetwork in this test is about 1.9x faster than in
the previous release:
function timeDlnetwork % Create a layer graph lgraph = layerGraph(resnet50); % Remove the output layer from the layer graph lgraph = removeLayers(lgraph, lgraph.Layers(end).Name); % Time assembling a dlnetwork object from the layer graph tic dlnet = dlnetwork(lgraph); toc end
The approximate execution times are:
R2022a: 4.09 seconds
R2022b: 2.13 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeAssembleNetwork and
timeDlnetwork functions.
dlarray Constructor: Improved performance
Creating dlarray data shows improved performance. For example, creating formatted
dlarray objects in this test is about 3.2x faster than in the
previous release:
function timeDlarray % Prepare data params = arrayfun(@(i)randn(5,5,20,20,'single'),1:10000,UniformOutput=0); % Time the dlarray constructor tic for i = 1:numel(params) params{i} = dlarray(params{i},'SSCB'); end toc end
The approximate execution times are:
R2022a: 0.55 seconds
R2022b: 0.17 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDlarray function.
replaceLayer Function: Improved performance
The replaceLayer function shows improved performance. For example, replacing the
final classification layer of a network containing a classification layer with different
classes in this test is about 1.5x faster than in the previous release:
function timeReplaceLayer % Load a pretrained network and get the name of the layer to be replaced net = squeezenet; cLayer = net.Layers(end); layerName = cLayer.Name; % Set the classes for the replacement layer cLayer.Classes = string(0:1000); % Convert the network to a layer graph lgraph = layerGraph(net); % Time the layer replacement tic lgraph = replaceLayer(lgraph,layerName,cLayer); toc end
The approximate execution times are:
R2022a: 0.41 seconds
R2022b: 0.28 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeReplaceLayer
function.
Parameter Updates: Improved performance of parameter updates using a GPU
Element-wise operations on large numbers of gpuArray objects show
improved performance, significantly speeding up parameter updates using the adamupdate, sgdmupdate, and rmspropupdate functions. Improvements are greater for networks that contain
more layers. For example, updating parameters using adamupdate in this
test is about 2.7x faster than in the previous release:
function timeAdamUpdate % Create a layer graph net = resnet101; lgraph = layerGraph(net); % Remove the output layer from the layer graph and create a dlnetwork lgraph = removeLayers(lgraph,lgraph.Layers(end).Name); net = dlnetwork(lgraph); % Convert the learnable parameters to gpuArray objects net = dlupdate(@gpuArray,net); % Initialize the update variables fakeG = net.Learnables; avgG = []; avgsqG = []; [net, avgG, avgsqG] = adamupdate(net,fakeG,avgG,avgsqG,1); % Time adamupdate gputimeit(@() adamupdate(net,fakeG,avgG,avgsqG,2)) end
The approximate execution times are:
R2022a: 0.81 seconds
R2022b: 0.30 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeAdamUpdate
function.
Custom Layers: Improved performance in dlnetwork
Custom layers in a dlnetwork object show improved performance for
training and inference. For example, computing the network outputs for training using
forward for a network containing a number of instances of functionLayer in this test is about 3.7x faster than in the previous
release:
function timeCustomLayer % Create a dlnetwork object containing custom layers layer = functionLayer(@(x) plus(x,1)); layerArray = repelem(layer,100); dlnet = dlnetwork(layerArray,dlarray(1,"SSCB")); % Prepare input data X = dlarray(1,"SSCB"); % Warm-up iterations for i=1:10 forward(dlnet,X); end % Timed iterations tic for i=1:100 forward(dlnet,X); end toc end
The approximate execution times are:
R2022a: 6.61 seconds
R2022b: 1.79 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeCustomLayer
function.
tan and tanh Functions: Improved performance
within dlgradient call
The tan and tanh functions show improved performance when used within a
dlgradient call. For example, evaluating gradients of a
tan function in this test is about 1.5x faster than in the previous
release:
function timeTan % Prepare input data and gradient function x = dlarray(randn(5000)); Tan = @(x) dlgradient(sum(tan(x),"all"),x); % Warm-up iterations for i = 1:10 dlfeval(Tan,x); end % Timed iterations tic for i = 1:10 dlfeval(Tan,x); end toc end
The approximate execution times are:
R2022a: 2.31 seconds
R2022b: 1.58 seconds
Evaluating gradients of a tanh function in this test is about 1.5x
faster than in the previous release:
function timeTanh % Prepare input data and gradient function x = dlarray(randn(5000)); Tanh = @(x) dlgradient(sum(tanh(x),"all"),x); % Warm-up iterations for i = 1:10 dlfeval(Tanh,x); end % Timed iterations tic for i = 1:10 dlfeval(Tanh,x); end toc end
The approximate execution times are:
R2022a: 3.10 seconds
R2022b: 2.03 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeTan and
timeTanh functions.
dlode45 Function: Improved performance within
dlgradient call
The dlode45 function shows improved performance when computing gradients using
the adjoint gradient mode within a dlgradient call.
In particular, the first call to dlode45 shows a significant
performance improvement. For example, evaluating gradients of a function calling
dlode45 for the first time in this test is about 6.1x faster than in
the previous release and subsequent calls are about 1.2x faster than in the previous
release:
function timeDlode45 % Prepare input data tspan = [0 1]; nChannels = 5; pp = cell(1,2); pp{1} = 0.01*dlarray(rand(nChannels)); pp{2} = 0.01*dlarray(rand(nChannels)); x = dlarray(rand(nChannels,100)); % Prepare gradient function ODE = @(~,y,p) p{2}*sin(p{1}*y); sol = @(x,pp,tspan) dlode45(ODE,tspan,x,pp,DataFormat="CB",GradientMode="adjoint"); gradsAdj = @(x,pp,tspan) dlgradient(sum(sol(x,pp,tspan),"all"),x,pp); % Time only the first call to dlode45 for i=1:5 % Clear any previously cached traces of the accelerated function clear functions tic dlfeval(gradsAdj,x,pp,tspan); time(i) = toc; end % Calculate the mean time for the first call to dlode45 timeFirstCall = mean(time) % Time many calls to dlode45 % Warm-up iterations clear functions for i = 1:10 dlfeval(gradsAdj,x,pp,tspan); end % Timed iterations tic for i=1:100 dlfeval(gradsAdj,x,pp,tspan); end toc end
The approximate execution times for the first call are:
R2022a: 44.5 seconds
R2022b: 7.3 seconds
The approximate execution times for the 100 subsequent calls are:
R2022a: 6.78 seconds
R2022b: 5.80 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDlode45 function.
DAGNetwork Inference: Improved performance of GPU inference on
networks performing operations with a stride
GPU inference using a DAGNetwork object containing layers
performing operations with a stride shows improved performance. Layers that perform
operations with a stride include pooling and convolution layers, such as a convolution2dLayer and a maxPooling2dLayer, when any element of the stride property
is greater than 1. The performance improvement is greater for networks
containing more layers that perform operations with a stride. For example, making
predictions using resnet101 on the GPU in this test is about 1.1x
faster than in the previous release:
function timeStrideInference % Load a trained network and prepare input data net = resnet101; X = rand(224,224,3,500,"gpuArray"); % Time the predict function gputimeit(@() predict(net,X)) end
The approximate execution times are:
R2022a: 0.52 seconds
R2022b: 0.49 seconds
The code was timed on a Windows 10, Intel
Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeStrideInference
function.
Pretrained Model on GitHub: Deep Speech speech-to-text model
To learn how to load a pretrained Deep Speech model into MATLAB®, see the Speech-to-Text Transcription Using Deep Speech repository. The Deep Speech model is suitable for transfer learning, code generation, and streaming.
To find the latest pretrained models for deep learning in MATLAB, see MATLAB Deep Learning Model Hub.
Functionality being removed or changed
trainNetwork pads mini-batches to length of longest sequence
before splitting when you specify SequenceLength training option as
an integer
Behavior change
Starting in R2022b, when you train a network with sequence data using the trainNetwork function and the SequenceLength option is an integer, the software pads sequences to the
length of the longest sequence in each mini-batch and then splits the sequences into
mini-batches with the specified sequence length. If SequenceLength
does not evenly divide the sequence length of the mini-batch, then the last split
mini-batch has a length shorter than SequenceLength. This behavior
prevents the network training on time steps that contain only padding values.
In previous releases, the software pads mini-batches of sequences to have a length
matching the nearest multiple of SequenceLength that is greater than
or equal to the mini-batch length and then splits the data. To reproduce this behavior,
use a custom training loop and implement this behavior when you preprocess mini-batches
of data.
Prediction functions pad mini-batches to length of longest sequence before
splitting when you specify SequenceLength option as an
integer
Behavior change
Starting in R2022b, when you make predictions with sequence data using the predict, classify, predictAndUpdateState, classifyAndUpdateState, and activations functions and the SequenceLength option is
an integer, the software pads sequences to the length of the longest sequence in each
mini-batch and then splits the sequences into mini-batches with the specified sequence
length. If SequenceLength does not evenly divide the sequence length
of the mini-batch, then the last split mini-batch has a length shorter than
SequenceLength. This behavior prevents time steps that contain only
padding values from influencing predictions.
In previous releases, the software pads mini-batches of sequences to have a length
matching the nearest multiple of SequenceLength that is greater than
or equal to the mini-batch length and then splits the data. To reproduce this behavior,
manually pad the input data such that the mini-batches have the length of the
appropriate multiple of SequenceLength. For sequence-to-sequence
workflows, you may also need to manually remove time steps of the output that correspond
to padding values.
Deep Learning Toolbox Converter for PyTorch Models: Support for importing networks from PyTorch
You can now import a pretrained PyTorch® model for image classification as a MATLAB network by using the importNetworkFromPyTorch function. The
importNetworkFromPyTorch function imports the PyTorch model as an uninitialized dlnetwork object. For an example
that shows how to add an input image to the imported network, initialize the network, and
use the network for image classification, see Import Network from PyTorch and Classify Image.
The importNetworkFromPyTorch function requires the new support
package Deep Learning Toolbox™ Converter for PyTorch Models. If this support package is not installed, then the function provides a
download link.
Export to TensorFlow Model: Save MATLAB network or layer graph as TensorFlow model
You can now export a Deep Learning Toolbox network or layer graph to TensorFlow™ by using the exportNetworkToTensorFlow function. The
exportNetworkToTensorFlow function saves the exported TensorFlow model in a regular Python® package. You can load the exported model and use it for prediction or
training. You can also share the exported model by saving it to
SavedModel or HDF5 format.
TensorFlow Import Operator Support: Import models that include
Assert, GreaterEqual, and Size
operators
You can now import a TensorFlow model that includes Assert,
GreaterEqual, and Size operators by using the
importTensorFlowNetwork and importTensorFlowLayers functions. For a list of the TensorFlow operators that the functions support for conversion into MATLAB functions with dlarray support, see Supported TensorFlow Operators.
Import and Export Workflows: New help and tips for interoperability between Deep Learning Toolbox, TensorFlow, PyTorch, and ONNX
These topics help you import networks from and export networks to external deep learning platforms:
Network Projection: Compress neural networks using neuron principal component analysis (October 2022; version 22.2.1)
Compress neural networks using projection with the compressNetworkUsingProjection function. The
compressNetworkUsingProjection function reduces the number of
learnables in a network using principal component analysis (PCA) to identify the subspace
of learnable parameters that result in the highest variance in neuron activations by
analyzing the network activations using a data set representative of the training data.
After the analysis, the function replaces supported layers with projected
layers. Forward passes of a projected deep neural network are typically
faster when you deploy the network to embedded hardware using library-free C or C++ code
generation.
A projected layer is a variant of a deep learning layer that enables compression by reducing the number of stored learnable parameters by replacing multiplications of the form , where W is a learnable matrix, with the multiplication , where Q is a projector matrix. Instead of storing W, the layer instead stores Q and . Projecting into a lower-dimensional space with Q typically requires less memory and can have similarly strong prediction accuracy.
The PCA step can be computationally intensive. If you expect to compress the same
network multiple times (for example, when exploring different levels of compression), then
you can perform the PCA step up front using a neuronPCA object.
These functions require the Deep Learning Toolbox Model Quantization Library support package. This support package is a free add-on that you can download using the Add-On Explorer. Alternatively, see Deep Learning Toolbox Model Quantization Library.
If you prune or quantize your network, then use compression using projection after pruning and before quantization. Network compression using projection supports projecting LSTM layers only.
For an example showing how to compress a network using projection, see Compress Neural Network Using Projection.
Quantization: Independently select calibration, simulation, and validation environments
The prerequisites required for each step of the quantization workflow now depend on your selection at each stage of the quantization workflow. For details, see Quantization Workflow Prerequisites.
In previous versions, the prerequisites required for validation of a quantized network on target hardware are also required for the calibration step of quantization. You can now choose the calibration environment to use independent of the selected execution environment and can choose to simulate the quantized network in MATLAB rather than quantizing and validating on hardware.
Quantization: Calibrate on host GPU or CPU
You can now choose whether to calibrate your network using the host GPU or host CPU.
By default, the calibrate function and the Deep Network Quantizer app calibrate
on the host GPU if one is available.
In previous versions, the execution environment must be the same as the instrumentation environment you use for the calibration step of quantization.
Quantization: Quantize dlnetwork objects
The dlquantizer object and Deep Network Quantizer now support dlnetwork objects for quantization with the calibrate and validate functions.
Quantization: Prepare for quantization with layer equalization
Use the equalizeLayers function to equalize the layer parameters of a deep neural
network. Equalizing layer parameters before quantization can improve the accuracy of the
quantized network and does not require data or retraining of the network.
Quantization: Specify mini-batch size for calibration
Use the MiniBatchSize argument of the calibrate function to specify the size of mini-batches for calibration.
Larger mini-batch sizes require more memory, but can lead to faster calibration.
Quantization: Simulate quantized network for FPGA execution environment
You can now use the quantize function to create a quantized network for simulation when you set
the ExecutionEnvironment property of dlquantizer to FPGA. The quantized network enables
visibility of the quantized layers, weights, and biases of the network, as well as
quantized inference behavior for simulation.
TensorFlow Lite: Generate C++ code for pretrained models and deploy on Windows platforms
Use the loadTFLiteModel function to load a pretrained TensorFlow Lite model into a TFLiteModel object. Use this object with the predict function in your MATLAB code to perform inference in MATLAB execution, code generation, or inside MATLAB Function blocks
in Simulink® models.
To use this functionality, you must install the Deep Learning Toolbox Interface for TensorFlow Lite. For more information, see Prerequisites for Deep Learning with TensorFlow Lite Models.
For examples, see:
Verification: AI Verification Library for Deep Learning Toolbox (October 2022; Version 22.2.1)
AI Verification Library for Deep Learning Toolbox enables testing of the robustness properties of deep learning networks. Use this library to verify whether a deep learning network is robust to adversarial examples and to compute the output bounds for a set of input bounds.
Use the verifyNetworkRobustness function to verify network robustness to
adversarial examples. A network is robust to adversarial examples if the class that
the network predicts does not change when the input is perturbed between the lower
and upper input bounds that you specify. For a set of input bounds, the function
checks whether the network is robust to adversarial examples between those input
bounds and returns verified, violated, or
unproven. For more information, see Verify Robustness of Deep Learning Neural Network .
Use the estimateNetworkOutputBounds function to estimate the range of output
values that the network returns when the input is between the lower and upper bounds
that you specify. Use this function to estimate how sensitive the network
predictions are to input perturbation.
Deep Learning Workflows: New and updated examples and topics
New examples and topics help you progress with deep learning:
Interpret Deep Learning Time-Series Classifications Using Grad-CAM
Train Latent ODE Network with Irregularly Sampled Time-Series Data
Multivariate Time Series Anomaly Detection Using Graph Neural Network
Export Image Classification Network from Deep Network Designer to Simulink
Improve Performance of Deep Learning Simulations in Simulink
Image Processing and Computer Vision: New examples
New examples for image processing and computer vision tasks, including the new Medical Imaging Toolbox™:
Signal Processing: New examples
New examples for signal processing tasks include:
Human Health Monitoring Using Continuous Wave Radar and Deep Learning
Detect Air Compressor Sounds in Simulink Using Wavelet Scattering (DSP System Toolbox)
Maritime Clutter Removal with Neural Networks (Radar Toolbox)
Signal Recovery with Differentiable Scalograms and Spectrograms (Signal Processing Toolbox)
Signal Source Separation Using W-Net Architecture (Signal Processing Toolbox)
Audio Processing: New examples
Wireless Communications: New examples
New examples for wireless applications include:
Deployment: New examples
New deployment examples include:
Deploy Object Detection Model as Microservice (MATLAB Compiler SDK)