R2022b

New Features, Bug Fixes, Compatibility Considerations

Apps and Visualization

Experiment Manager: Monitor model performance while you run experiments

During training, the results table now displays the intermediate values for standard training and validation metrics for built-in training experiments. These metrics include loss, accuracy of classification experiments, and root mean squared error of regression experiments.

Experiment Manager: Inspect execution environment for each trial

In built-in training experiments, the new Execution Environment column of the results table displays whether each trial runs on a single CPU, a single GPU, multiple CPUs, or multiple GPUs. To show this information, click the show or hide columns button located above the results table and select Execution Environment.

Experiment Manager: Restart multiple trials

Restart multiple trials of your experiment by opening the Restart list, selecting one or more restarting criteria, and clicking Restart . The restarting criteria include All Canceled, All Stopped, All Error, and All Discarded. For more information, see Stop and Restart Training.

Experiment Manager: Reduce storage size by discarding unwanted results

To reduce the size of your experiments, you can now discard the results of trials that are no longer relevant. In the Actions column of the results table, click the Discard button for a trial. Experiment Manager deletes the training plot, confusion matrix, trained network, training information, and training output from your project.

Experiment Manager: New properties for experiments.Monitor objects

The experiments.Monitor object has these new properties:

  • InfoData — Information column values for trial

  • MetricData — Metric column values for trial

An experiments.Monitor object has the same properties and object functions as a TrainingProgressMonitor object. Therefore, you can easily adapt your custom training loop plotting code for use in an Experiment Manager setup script. For more information, see Prepare Plotting Code for Custom Training Experiment.

Deep Network Designer: New layer groups and colors

Deep Network Designer has improved layer colors and layer groups. Smaller layer groups mean you can easily find the layer you need.

Visualization: Monitor and plot custom training loop progress

Track custom training loop progress and generate training plots using a TrainingProgressMonitor object.

Create a TrainingProgressMonitor object using the trainingProgressMonitor function. Use the TrainingProgressMonitor object to:

  • Track training progress.

  • Display animated plots of metric values during training.

  • Display and monitor information values during training.

  • Log metric and information values.

After you create a TrainingProgressMonitor object, you can use these object functions:

For more information, see Monitor Custom Training Loop Progress.

A TrainingProgressMonitor object has the same properties and object functions as an experiments.Monitor object. Therefore, you can easily adapt your plotting code for use in an Experiment Manager setup script. For more information, see Prepare Plotting Code for Custom Training Experiment.

Visualization: Compute and plot ROC curve

Compute and plot classification performance metrics including receiver operating characteristic (ROC) curves.

Use rocmetrics to evaluate the performance of classification models with performance metrics. You can create a rocmetrics object by passing true labels, classification scores, and class names. By default, rocmetrics computes the true positive rates (TPR), false positive rates (FPR), and area under the ROC curve for each class. Additionally, you can specify more supported performance metrics by using the AdditionalMetrics name-value argument.

After you create a rocmetrics object, you can use these object functions:

  • plot — Plot ROC or other classifier performance curves. The plot function returns a ROCCurve graphics object for each curve. You can control the appearance of each curve by modifying the properties of the objects. For details, see ROCCurve Properties.

  • average — Compute performance metrics for an average ROC curve for multiclass problems.

  • addMetrics — Compute additional classification performance metrics.

For an example that shows how to use ROC curves to compare the performance of deep learning models, see Compare Deep Learning Models Using ROC Curves.

In R2022a, rocmetrics was introduced in Statistics and Machine Learning Toolbox™. You can now use rocmetrics without a Statistics and Machine Learning Toolbox license. Some options, for example, rocmetrics using confidence intervals, still require a Statistics and Machine Learning Toolbox license. For more information, see rocmetrics.

Interpretability: Create Grad-CAM maps with 1-D convolutional networks for sequence and time-series data

Compute Grad-CAM interpretability maps for deep neural networks that are trained on sequence and time-series data. To compute Grad-CAM interpretability maps, use the gradCAM function. For an example that shows how to use a Grad-CAM map to interpret the predictions of a network trained on time series data, see Interpret Deep Learning Time-Series Classifications Using Grad-CAM.

Algorithms

Complex Numbers: Train networks using complex-valued data

To pass complex-valued data to a neural network, you can use the input layer to split the complex values into their real and imaginary parts before the network passes the data to the subsequent layers. When the layer splits complex-valued data, the layer outputs the data in twice as many channels as the input data.

The ImageInputLayer, SequenceInputLayer, and FeatureInputLayer objects support splitting input data into its real and imaginary components. To input complex-valued data into a network by splitting it into its real and imaginary parts, set the SplitComplexInputs option of the network input layer to 1 (true).

To input complex data into a network, the SplitComplexInputs option must be 1.

For an example that shows how to train a network with complex-valued data, see Train Network with Complex-Valued Data.

Gaussian Error Linear Unit (GELU) Activation: Create and train networks with GELU activation

The Gaussian error linear unit (GELU) operation weights the input by its probability under a Gaussian distribution. This operation is given by

GELU(x)=x2(1+​erf(x2)),

where erf denotes the error function.

To apply the GELU operation in a layer array or layer graph, use a geluLayer object.

To apply the GELU operation in a custom layer or deep learning model function, use the gelu function.

Spatio-Temporal Data: 2-D average and max pooling layers support data with both spatial and time dimensions

AveragePooling2DLayer and MaxPooling2DLayer objects now support input data with these data formats:

  • One spatial dimension and a time dimension

  • Two spatial dimensions and a time dimension

For an example that shows how to create a 2-D CNN-LSTM network for speech classification tasks by combining a 2-D convolutional neural network (CNN) with a long short-term memory (LSTM) layer, see Sequence Classification Using CNN-LSTM Network.

Sequence Networks: Make predictions in parallel

The predict, classify, and activations functions now support making predictions in parallel for networks with sequence input. To make predictions in parallel, set the ExecutionEnvironment option to "parallel" or "multi-gpu". Making predictions in parallel requires a Parallel Computing Toolbox™ license.

When you make predictions in parallel for networks with recurrent layers, the SequenceLength option must be "longest" or "shortest".

Networks with custom layers that contain State parameters do not support making predictions in parallel.

LSTM Projected Layer: Perform LSTM operations with fewer learnable parameters

To compress a deep learning network by reducing the number of learnable parameters, you can use projected layers. A projected layer is a variant of a deep learning layer that enables compression by reducing the number of stored learnable parameters by replacing multiplications of the form Wx, where W is a learnable matrix, with the multiplication WQQ⊤x, where Q is a projector matrix. Instead of storing W, the layer instead stores Q and W′=WQ. Projecting into a lower-dimensional space with Q typically requires less memory and can have similarly strong prediction accuracy.

To create a long short-term memory (LSTM) projected layer, use lstmProjectedLayer. For an example that shows how to train a network with an LSTM projected layer and compare the prediction accuracy and number of learnable parameters to a network without projection, see Train Network with LSTM Projected Layer.

Custom Layers: Define Custom Layer Learnable Parameter Initialization

When you create a custom layer, specify a custom learnable parameter initialization function by implementing a function with this syntax:

function layer = initialize(layer,layout1,...,layoutN)
   ...
end
The input layer is an instance of the custom layer and layout1,...,layoutN are networkDataLayout objects corresponding to each of the N inputs of the layer.

The software uses this initialize function when you include the layer in a dlnetwork object or call the initialize function on the dlnetwork object.

For an example that shows how to define a custom layer that uses automatic learnable parameter initialization, see Define Custom Deep Learning Layer with Learnable Parameters.

For more information about defining custom layers, see Define Custom Deep Learning Layers.

Customization: Add, remove, and replace layers of dlnetwork objects

Add, remove, and replace layers in dlnetwork objects using these functions:

Customization: Plot and view summary of dlnetwork objects

Plot dlnetwork architecture using the plot function.

Print a summary of objects using the summary function. The summary shows whether the network is initialized, the total number of learnable parameters, and information about the network inputs.

Customization: Initialize dlnetwork objects using only input size and format information

Initialize the learnable parameters of a dlnetwork object by passing a networkDataLayout object to the dlnetwork and initialize functions.

In previous versions, the dlnetwork and initialize functions require you to pass examples of input data. You can now use networkDataLayout objects, which only contain the data size and dlarray format.

Create an unformatted networkDataLayout object using the syntax layout = networkDataLayout(sz), where sz specifies the size of the data. Create a formatted networkDataLayout object using the syntax layout = networkDataLayout(sz,fmt), where fmt specifies the dlarray format of the data.

Attention: Apply attention operation to dlarray input

The attention operation focuses on parts of the input using weighted multiplication operations. To apply the attention operation to a set of queries, keys, and values, use the attention function. You can specify the scale, dropout probability, and padding and attention masks used in the operation.

You can use the attention function when you work with custom training loops and custom layers to implement:

  • Dot product attention

  • Scaled dot product attention

  • Multihead attention

  • Luong attention

For an example that shows how to use attention for sequence-to-sequence translation, see Sequence-to-Sequence Translation Using Attention.

Automatic Differentiation: Use more functions with dlarray input

Use these functions with a dlarray object as input when you work with custom training loops and custom layers.

  • Attention — Apply the attention operation to a set of queries, keys, and values, using attention.

  • GELU — Apply the GELU activation function to input elements using gelu.

  • Error function — Compute the error function of input elements using erf.

  • String – Convert dlarray objects to string objects using string.

  • Categorical – Convert dlarray objects to categorical objects using categorical.

For dlarray input, the string and categorical functions convert to these respective data types only and do not support automatic differentiation.

For a full list of functions that support dlarray input, see List of Functions with dlarray Support.

Function Layer: Accelerate layer functions

You can now specify that a layer function supports acceleration using dlaccelerate by setting the Acceleratable property to 1 (true) when you create a functionLayer. Setting Acceleratable to 1 (true) can improve the performance of training and inference (prediction) when you use a dlnetwork object. For example, calling predict on a dlnetwork object containing a number of functionLayer objects in this test is about 3x faster than in the previous release:

function timeFunctionLayer

% Prepare input data
X = dlarray(rand(500,500,3,"single","gpuArray"),"SSCB");

% Prepare a convolution and ReLU block
convBlock = [convolution2dLayer(4,20); reluLayer()];

if version("-release") == "2022b"
    % Create a network using functionLayer ReLU with Acceleratable
    convBlock(2) = functionLayer(@(x) relu(x),Acceleratable=true);
    layers = repmat(convBlock,80,1);
    net = dlnetwork(layers,X);
else
    % Create a network using functionLayer ReLU
    convBlock(2) = functionLayer(@(x) relu(x));
    layers = repmat(convBlock,80,1);
    net = dlnetwork(layers,X);
end

% Time the predict function
gputimeit(@()predict(net,X))

end

The approximate execution times are:

R2022a: 0.12 seconds

R2022b: 0.04 seconds

The code was timed on a Windows® 10, Intel® Xeon® W-2133 @ 3.60 GHz test system with an NVIDIA® RTX A5000 GPU by calling the timeFunctionLayer function.

Bayesian Neural Networks: Create and train Bayesian neural networks

A Bayesian neural network (BNN) is a type of deep learning network that uses Bayesian methods to quantify the uncertainty in the predictions of a deep learning network.

You can train a BNN to generate a distribution of weights and biases, rather than a single set. You can then use these distributions to measure the uncertainty of the network predictions.

To learn how to train a BNN using the Bayes by backpropagation method, see Train Bayesian Neural Network.

Background Dispatch: Use DispatchInBackground on thread pools

You can now use background dispatch (asynchronous prefetch queuing) for reading training data from datastores on thread-based parallel pools.

To use background dispatch when training a network using trainNetwork, set the DispatchInBackground training option to 1 (true) using the trainingOptions function and open a thread-based parallel pool using parpool("Threads").

To use background dispatch when training a network using a custom training loop, create a minibatchqueue object and set the DispatchInBackground property to 1 (true), and open a thread-based parallel pool using parpool("Threads").

As an alternative to opening a thread-based parallel pool using parpool("Threads"), you can set the default parallel environment on your local machine from the MATLAB desktop Home tab, in the Environment area, by selecting Parallel > Select Parallel Environment > Threads, or by calling parallel.defaultProfile("Threads").

Network Creation: Improved performance

Assembling a DAGNetwork object using assembleNetwork and assembling a dlnetwork object using dlnetwork show improved performance. Improvements are greater for networks that contain more layers. For example, assembling a DAGNetwork object in this test is about 1.9x faster than in the previous release:

function timeAssembleNetwork

% Create a layer graph
lgraph = layerGraph(resnet50);

% Time assembling a DAGNetwork object from the layer graph 
tic
net = assembleNetwork(lgraph);
toc

end

The approximate execution times are:

R2022a: 3.80 seconds

R2022b: 2.01 seconds

Assembling a dlnetwork in this test is about 1.9x faster than in the previous release:

function timeDlnetwork

% Create a layer graph
lgraph = layerGraph(resnet50);

% Remove the output layer from the layer graph
lgraph = removeLayers(lgraph, lgraph.Layers(end).Name);

% Time assembling a dlnetwork object from the layer graph 
tic
dlnet = dlnetwork(lgraph);
toc

end

The approximate execution times are:

R2022a: 4.09 seconds

R2022b: 2.13 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeAssembleNetwork and timeDlnetwork functions.

dlarray Constructor: Improved performance

Creating dlarray data shows improved performance. For example, creating formatted dlarray objects in this test is about 3.2x faster than in the previous release:

function timeDlarray

% Prepare data
params = arrayfun(@(i)randn(5,5,20,20,'single'),1:10000,UniformOutput=0);

% Time the dlarray constructor 
tic
for i = 1:numel(params)
    params{i} = dlarray(params{i},'SSCB');
end
toc

end

The approximate execution times are:

R2022a: 0.55 seconds

R2022b: 0.17 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDlarray function.

replaceLayer Function: Improved performance

The replaceLayer function shows improved performance. For example, replacing the final classification layer of a network containing a classification layer with different classes in this test is about 1.5x faster than in the previous release:

function timeReplaceLayer

% Load a pretrained network and get the name of the layer to be replaced
net = squeezenet;
cLayer = net.Layers(end);
layerName = cLayer.Name;

% Set the classes for the replacement layer
cLayer.Classes = string(0:1000);

% Convert the network to a layer graph
lgraph = layerGraph(net);

% Time the layer replacement
tic
lgraph = replaceLayer(lgraph,layerName,cLayer);
toc

end

The approximate execution times are:

R2022a: 0.41 seconds

R2022b: 0.28 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeReplaceLayer function.

Parameter Updates: Improved performance of parameter updates using a GPU

Element-wise operations on large numbers of gpuArray objects show improved performance, significantly speeding up parameter updates using the adamupdate, sgdmupdate, and rmspropupdate functions. Improvements are greater for networks that contain more layers. For example, updating parameters using adamupdate in this test is about 2.7x faster than in the previous release:

function timeAdamUpdate

% Create a layer graph
net = resnet101;
lgraph = layerGraph(net);

% Remove the output layer from the layer graph and create a dlnetwork
lgraph = removeLayers(lgraph,lgraph.Layers(end).Name);
net = dlnetwork(lgraph);

% Convert the learnable parameters to gpuArray objects
net = dlupdate(@gpuArray,net);

% Initialize the update variables
fakeG = net.Learnables;
avgG = [];
avgsqG = [];
[net, avgG, avgsqG] = adamupdate(net,fakeG,avgG,avgsqG,1);

% Time adamupdate
gputimeit(@() adamupdate(net,fakeG,avgG,avgsqG,2))

end

The approximate execution times are:

R2022a: 0.81 seconds

R2022b: 0.30 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeAdamUpdate function.

Custom Layers: Improved performance in dlnetwork

Custom layers in a dlnetwork object show improved performance for training and inference. For example, computing the network outputs for training using forward for a network containing a number of instances of functionLayer in this test is about 3.7x faster than in the previous release:

function timeCustomLayer

% Create a dlnetwork object containing custom layers
layer = functionLayer(@(x) plus(x,1));
layerArray = repelem(layer,100);
dlnet = dlnetwork(layerArray,dlarray(1,"SSCB"));

% Prepare input data
X = dlarray(1,"SSCB");

% Warm-up iterations
for i=1:10
    forward(dlnet,X);
end

% Timed iterations
tic
for i=1:100
    forward(dlnet,X);
end
toc

end

The approximate execution times are:

R2022a: 6.61 seconds

R2022b: 1.79 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeCustomLayer function.

tan and tanh Functions: Improved performance within dlgradient call

The tan and tanh functions show improved performance when used within a dlgradient call. For example, evaluating gradients of a tan function in this test is about 1.5x faster than in the previous release:

function timeTan

% Prepare input data and gradient function
x = dlarray(randn(5000));
Tan = @(x) dlgradient(sum(tan(x),"all"),x);

% Warm-up iterations
for i = 1:10
    dlfeval(Tan,x);
end

% Timed iterations
tic
for i = 1:10
    dlfeval(Tan,x);
end
toc

end

The approximate execution times are:

R2022a: 2.31 seconds

R2022b: 1.58 seconds

Evaluating gradients of a tanh function in this test is about 1.5x faster than in the previous release:

function timeTanh

% Prepare input data and gradient function
x = dlarray(randn(5000));
Tanh = @(x) dlgradient(sum(tanh(x),"all"),x);

% Warm-up iterations
for i = 1:10
    dlfeval(Tanh,x);
end

% Timed iterations
tic
for i = 1:10
    dlfeval(Tanh,x);
end
toc

end

The approximate execution times are:

R2022a: 3.10 seconds

R2022b: 2.03 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeTan and timeTanh functions.

dlode45 Function: Improved performance within dlgradient call

The dlode45 function shows improved performance when computing gradients using the adjoint gradient mode within a dlgradient call. In particular, the first call to dlode45 shows a significant performance improvement. For example, evaluating gradients of a function calling dlode45 for the first time in this test is about 6.1x faster than in the previous release and subsequent calls are about 1.2x faster than in the previous release:

function timeDlode45

% Prepare input data 
tspan = [0 1];
nChannels = 5;
pp = cell(1,2);
pp{1} = 0.01*dlarray(rand(nChannels));
pp{2} = 0.01*dlarray(rand(nChannels));
x  = dlarray(rand(nChannels,100));

% Prepare gradient function
ODE = @(~,y,p) p{2}*sin(p{1}*y);
sol = @(x,pp,tspan) dlode45(ODE,tspan,x,pp,DataFormat="CB",GradientMode="adjoint");
gradsAdj = @(x,pp,tspan) dlgradient(sum(sol(x,pp,tspan),"all"),x,pp);

% Time only the first call to dlode45
for i=1:5
    % Clear any previously cached traces of the accelerated function
    clear functions 

    tic
    dlfeval(gradsAdj,x,pp,tspan);
    time(i) = toc;
end

% Calculate the mean time for the first call to dlode45
timeFirstCall = mean(time)


% Time many calls to dlode45
% Warm-up iterations
clear functions 
for i = 1:10
    dlfeval(gradsAdj,x,pp,tspan);
end

% Timed iterations
tic
for i=1:100
    dlfeval(gradsAdj,x,pp,tspan);
end
toc

end

The approximate execution times for the first call are:

R2022a: 44.5 seconds

R2022b: 7.3 seconds

The approximate execution times for the 100 subsequent calls are:

R2022a: 6.78 seconds

R2022b: 5.80 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDlode45 function.

DAGNetwork Inference: Improved performance of GPU inference on networks performing operations with a stride

GPU inference using a DAGNetwork object containing layers performing operations with a stride shows improved performance. Layers that perform operations with a stride include pooling and convolution layers, such as a convolution2dLayer and a maxPooling2dLayer, when any element of the stride property is greater than 1. The performance improvement is greater for networks containing more layers that perform operations with a stride. For example, making predictions using resnet101 on the GPU in this test is about 1.1x faster than in the previous release:

function timeStrideInference

% Load a trained network and prepare input data
net = resnet101;
X = rand(224,224,3,500,"gpuArray");

% Time the predict function
gputimeit(@() predict(net,X))

end

The approximate execution times are:

R2022a: 0.52 seconds

R2022b: 0.49 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeStrideInference function.

Pretrained Model on GitHub: Deep Speech speech-to-text model

To learn how to load a pretrained Deep Speech model into MATLAB®, see the Speech-to-Text Transcription Using Deep Speech repository. The Deep Speech model is suitable for transfer learning, code generation, and streaming.

To find the latest pretrained models for deep learning in MATLAB, see MATLAB Deep Learning Model Hub.

 Functionality being removed or changed

trainNetwork pads mini-batches to length of longest sequence before splitting when you specify SequenceLength training option as an integer

Behavior change

Starting in R2022b, when you train a network with sequence data using the trainNetwork function and the SequenceLength option is an integer, the software pads sequences to the length of the longest sequence in each mini-batch and then splits the sequences into mini-batches with the specified sequence length. If SequenceLength does not evenly divide the sequence length of the mini-batch, then the last split mini-batch has a length shorter than SequenceLength. This behavior prevents the network training on time steps that contain only padding values.

In previous releases, the software pads mini-batches of sequences to have a length matching the nearest multiple of SequenceLength that is greater than or equal to the mini-batch length and then splits the data. To reproduce this behavior, use a custom training loop and implement this behavior when you preprocess mini-batches of data.

Prediction functions pad mini-batches to length of longest sequence before splitting when you specify SequenceLength option as an integer

Behavior change

Starting in R2022b, when you make predictions with sequence data using the predict, classify, predictAndUpdateState, classifyAndUpdateState, and activations functions and the SequenceLength option is an integer, the software pads sequences to the length of the longest sequence in each mini-batch and then splits the sequences into mini-batches with the specified sequence length. If SequenceLength does not evenly divide the sequence length of the mini-batch, then the last split mini-batch has a length shorter than SequenceLength. This behavior prevents time steps that contain only padding values from influencing predictions.

In previous releases, the software pads mini-batches of sequences to have a length matching the nearest multiple of SequenceLength that is greater than or equal to the mini-batch length and then splits the data. To reproduce this behavior, manually pad the input data such that the mini-batches have the length of the appropriate multiple of SequenceLength. For sequence-to-sequence workflows, you may also need to manually remove time steps of the output that correspond to padding values.

Interoperability

Deep Learning Toolbox Converter for PyTorch Models: Support for importing networks from PyTorch

You can now import a pretrained PyTorch® model for image classification as a MATLAB network by using the importNetworkFromPyTorch function. The importNetworkFromPyTorch function imports the PyTorch model as an uninitialized dlnetwork object. For an example that shows how to add an input image to the imported network, initialize the network, and use the network for image classification, see Import Network from PyTorch and Classify Image.

The importNetworkFromPyTorch function requires the new support package Deep Learning Toolbox™ Converter for PyTorch Models. If this support package is not installed, then the function provides a download link.

Export to TensorFlow Model: Save MATLAB network or layer graph as TensorFlow model

You can now export a Deep Learning Toolbox network or layer graph to TensorFlow™ by using the exportNetworkToTensorFlow function. The exportNetworkToTensorFlow function saves the exported TensorFlow model in a regular Python® package. You can load the exported model and use it for prediction or training. You can also share the exported model by saving it to SavedModel or HDF5 format.

TensorFlow Import Operator Support: Import models that include Assert, GreaterEqual, and Size operators

You can now import a TensorFlow model that includes Assert, GreaterEqual, and Size operators by using the importTensorFlowNetwork and importTensorFlowLayers functions. For a list of the TensorFlow operators that the functions support for conversion into MATLAB functions with dlarray support, see Supported TensorFlow Operators.

Import and Export Workflows: New help and tips for interoperability between Deep Learning Toolbox, TensorFlow, PyTorch, and ONNX

These topics help you import networks from and export networks to external deep learning platforms:

Deployment

Network Projection: Compress neural networks using neuron principal component analysis (October 2022; version 22.2.1)

Compress neural networks using projection with the compressNetworkUsingProjection function. The compressNetworkUsingProjection function reduces the number of learnables in a network using principal component analysis (PCA) to identify the subspace of learnable parameters that result in the highest variance in neuron activations by analyzing the network activations using a data set representative of the training data. After the analysis, the function replaces supported layers with projected layers. Forward passes of a projected deep neural network are typically faster when you deploy the network to embedded hardware using library-free C or C++ code generation.

A projected layer is a variant of a deep learning layer that enables compression by reducing the number of stored learnable parameters by replacing multiplications of the form Wx, where W is a learnable matrix, with the multiplication WQQ⊤x, where Q is a projector matrix. Instead of storing W, the layer instead stores Q and W′=WQ. Projecting into a lower-dimensional space with Q typically requires less memory and can have similarly strong prediction accuracy.

The PCA step can be computationally intensive. If you expect to compress the same network multiple times (for example, when exploring different levels of compression), then you can perform the PCA step up front using a neuronPCA object.

These functions require the Deep Learning Toolbox Model Quantization Library support package. This support package is a free add-on that you can download using the Add-On Explorer. Alternatively, see Deep Learning Toolbox Model Quantization Library.

If you prune or quantize your network, then use compression using projection after pruning and before quantization. Network compression using projection supports projecting LSTM layers only.

For an example showing how to compress a network using projection, see Compress Neural Network Using Projection.

Quantization: Independently select calibration, simulation, and validation environments

The prerequisites required for each step of the quantization workflow now depend on your selection at each stage of the quantization workflow. For details, see Quantization Workflow Prerequisites.

In previous versions, the prerequisites required for validation of a quantized network on target hardware are also required for the calibration step of quantization. You can now choose the calibration environment to use independent of the selected execution environment and can choose to simulate the quantized network in MATLAB rather than quantizing and validating on hardware.

Quantization: Calibrate on host GPU or CPU

You can now choose whether to calibrate your network using the host GPU or host CPU. By default, the calibrate function and the Deep Network Quantizer app calibrate on the host GPU if one is available.

In previous versions, the execution environment must be the same as the instrumentation environment you use for the calibration step of quantization.

Quantization: Quantize dlnetwork objects

The dlquantizer object and Deep Network Quantizer now support dlnetwork objects for quantization with the calibrate and validate functions.

Quantization: Prepare for quantization with layer equalization

Use the equalizeLayers function to equalize the layer parameters of a deep neural network. Equalizing layer parameters before quantization can improve the accuracy of the quantized network and does not require data or retraining of the network.

Quantization: Specify mini-batch size for calibration

Use the MiniBatchSize argument of the calibrate function to specify the size of mini-batches for calibration. Larger mini-batch sizes require more memory, but can lead to faster calibration.

Quantization: Simulate quantized network for FPGA execution environment

You can now use the quantize function to create a quantized network for simulation when you set the ExecutionEnvironment property of dlquantizer to FPGA. The quantized network enables visibility of the quantized layers, weights, and biases of the network, as well as quantized inference behavior for simulation.

TensorFlow Lite: Generate C++ code for pretrained models and deploy on Windows platforms

Use the loadTFLiteModel function to load a pretrained TensorFlow Lite model into a TFLiteModel object. Use this object with the predict function in your MATLAB code to perform inference in MATLAB execution, code generation, or inside MATLAB Function blocks in Simulink® models.

To use this functionality, you must install the Deep Learning Toolbox Interface for TensorFlow Lite. For more information, see Prerequisites for Deep Learning with TensorFlow Lite Models.

For examples, see:

Verification

Verification: AI Verification Library for Deep Learning Toolbox (October 2022; Version 22.2.1)

AI Verification Library for Deep Learning Toolbox enables testing of the robustness properties of deep learning networks. Use this library to verify whether a deep learning network is robust to adversarial examples and to compute the output bounds for a set of input bounds.

  • Use the verifyNetworkRobustness function to verify network robustness to adversarial examples. A network is robust to adversarial examples if the class that the network predicts does not change when the input is perturbed between the lower and upper input bounds that you specify. For a set of input bounds, the function checks whether the network is robust to adversarial examples between those input bounds and returns verified, violated, or unproven. For more information, see Verify Robustness of Deep Learning Neural Network .

  • Use the estimateNetworkOutputBounds function to estimate the range of output values that the network returns when the input is between the lower and upper bounds that you specify. Use this function to estimate how sensitive the network predictions are to input perturbation.

Application Examples

Image Processing and Computer Vision: New examples

Wireless Communications: New examples

New examples for wireless applications include:

Deployment: New examples

New deployment examples include: