R2026b

New Features, Bug Fixes, Compatibility Considerations

Parallel Language and Cluster Computing

 Explain parfor errors with MATLAB Copilot

When the MATLAB® Code Analyzer app reports warnings or errors in a parfor-loop, you can now use a MATLAB Copilot action to get a more detailed explanation of the issue. Copilot explains the parfor requirement, guideline, or limitation that the code violates, and it suggests changes to help resolve the error. To use MATLAB Copilot, you must have a MATLAB Copilot license. For more information, see Resolve Issues in parfor-loops Using MATLAB Copilot.

myParallelCode file in the MATLAB Editor. The Code Analyzer message displays the error "Line 2: Unable to classify the variable 'x' in the body of the parfor loop." and Detail and Explain buttons

 Parallel Panel: Improved interface for managing your parallel computing environment

Use the Parallel panel to manage cluster profiles, discover clusters on your network or in the cloud, start and stop parallel pools, and select and monitor GPU devices. If the Parallel panel icon is not in the sidebar, click the Open more panels button . In the Open Panel dialog box that appears, select the Parallel panel.

For more information about using the panel to manage cluster profiles, see Discover Clusters and Use Cluster Profiles.

The Parallel panel in the MATLAB desktop sidebar, showing the default cluster profile, cluster management options, and GPU environment.

Parallel Pools: Improved resilience to worker loss

Parallel pools can recover from worker loss without interrupting your computation. Parallel pool recovery is now supported on all cluster types. Previously, it was not supported on third-party scheduler clusters.

When you create an interactive or batch parallel pool, MATLAB starts the pool without Single Program Multiple Data (SPMD) communication enabled, allowing the pool to recover from worker loss. MATLAB enables SPMD communication automatically when required, such as when you use spmd blocks or distributed arrays.

For more information, see How Parallel Pool Recovery Works.

Pool Dashboard: Select time range in Timeline

The Timeline graph of the Pool Dashboard now supports pan and zoom interactions that enable you to select regions of interest. When you zoom or pan to a time range in the Timeline graph, the Parallel Constructs, Summary, and Worker Summary tables display information about pool activity in the selected time range.

Run Signal Processing and Wavelet Functions on Thread Workers

Most Signal Processing Toolbox™ and Wavelet Toolbox™ functions now support thread-based environments. For full lists of supported functions, see Signal Processing Toolbox Functions with Thread Support (Signal Processing Toolbox) and Wavelet Toolbox Functions with Thread Support (Wavelet Toolbox).

For more information, see Run MATLAB Functions in Thread-Based Environment.

Improved memory usage for batch processing figures

When you create multiple figure windows using parallel workers, the figure windows now consume less memory than in previous releases. To further reduce memory usage, close any figures that you no longer need.

Distributed Arrays: Use functions with new and enhanced distributed array support

These functions have new and enhanced distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

Tall Arrays: Use functions with new and enhanced tall array support

This function has enhanced tall array support:

  • unique — You can now use the TreatMissingAsDistinct name-value argument.

For more information, see Tall Arrays for Out-of-Memory Data.

MATLAB Job Scheduler: Control scripts accept double-dash option syntax

All MATLAB Job Scheduler control scripts now accept a double dash (--) for command-line options, in addition to the existing single dash (-) syntax.

Intel MPI Upgrade to Version "2021.17.0"

Parallel Computing Toolbox™ and MATLAB Parallel Server™ now include an upgraded Intel® MPI Library, version "2021.17.0", for use on Linux® for third-party schedulers.

To use Intel MPI, set the MPIImplementation additional property for your third-party cluster profile or object to "IntelMPI". To learn more about setting additional properties, see Set Additional Properties (MATLAB Parallel Server).

  Java Runtime no longer installed with MATLAB Parallel Server

MATLAB Parallel Server no longer includes Java® as part of its installation. Before R2026b, MATLAB Parallel Server installations on Windows® and Linux platforms included a Java Runtime Environment (JRE™).

 Compatibility Considerations
  • MATLAB Job Scheduler Clusters

    To start and run scheduler processes, MATLAB Job Scheduler requires a supported version of the JRE software. You can set MATLAB Job Scheduler to use your own JRE installation. Alternatively, you can allow MATLAB Job Scheduler to download and install a compatible OpenJDK® version by using the MATLAB Support for OpenJDK add-on. For more details, see Java Configuration for MATLAB Job Scheduler (MATLAB Parallel Server).

  • MATLAB Parallel Server Cluster Workers

    MATLAB Parallel Server cluster workers no longer have access to JRE software. If you want cluster workers to run code that requires Java, you can configure your MATLAB Parallel Server installation to use any compatible JRE version installed on the cluster nodes. For details, see Configure MATLAB Parallel Server Workers to Use Java (MATLAB Parallel Server).

 Admin Center: Simplified Tests for MATLAB Job Scheduler Cluster Connectivity

The Admin Center cluster connectivity tests for MATLAB Job Scheduler clusters are simplified. Admin Center now runs only Client to Nodes and Node to Nodes connectivity tests. Additionally, when you receive test failures, Admin Center now provides clear reasons for the failure and actionable next steps.

For example, these images show the connectivity tests in progress in each release.

R2026aR2026b

R2026a Admin Center Running Tests dialog box with four test sections: Client, Client to Nodes, Nodes to Nodes, and Nodes to Client, showing multiple connectivity checks in each section.

R2026b Admin Center Running Tests dialog box with two test sections: Client to Nodes and Nodes to Nodes, each showing a single connectivity check.

 Compatibility Considerations

The Admin Center connectivity tests Client and Nodes to Client have been removed.

 Functionality being removed or changed

parfor and mapreduce now throw an error when automatic pool creation fails

Behavior change

If no parallel pool is open and automatic pool creation is enabled in your parallel settings, then when automatic pool creation with the default parallel environment fails, the parfor and mapreduce functions now throw an error. You can use the error message to troubleshoot issues with pool creation. In previous releases, the parfor and mapreduce functions run in serial when automatic pool creation fails.

In most cases, you do not need to make changes to your code if you automatically start a parallel pool with the parfor or mapreduce functions. However, if you do not want parfor or mapreduce to throw an error when MATLAB cannot automatically create a pool using the default parallel environment, use one of these workarounds.

FunctionExample CodeWorkaround
parfor

parfor(idx = 1:5)
    idx
end

try
    % Attempt to return current pool object or create a new pool object
    pool = gcp; 
catch Error
    warning("Failed to automatically create a parallel pool " + ...
        "with the default parallel environment. " + ...
        "Running serially on the client instead.")
    pool = gcp("nocreate"); % Otherwise create an empty pool object
end
parfor(idx = 1:5,pool)
    idx
end

mapreduce
outds = mapreduce(ds,mapfun,reducefun);

try
   % Attempt to return current pool object or create a new pool object
   pool = gcp;
   mr = mapreducer(pool); % Create ParallelMapReducer object
catch Error
    warning("Failed to automatically create a parallel pool " + ...
        "with the default parallel environment. " + ...
        "Running serially on the client instead.")
    mr = mapreducer(0); % Otherwise, create a SerialMapReducer object
end
outds = mapreduce(ds,mapfun,reducefun,mr);

Support for Windows Compute Cluster Server 2003, Windows HPC Server 2008, Windows HPC Server 2008 R2, Microsoft HPC Pack 2012, and Microsoft HPC Pack 2012 R2 has been removed

Errors

Support for these Microsoft® HPC schedulers has been removed:

  • Windows Compute Cluster Server 2003

  • Windows HPC Server 2008

  • Windows HPC Server 2008 R2

  • Microsoft HPC Pack 2012

  • Microsoft HPC Pack 2012 R2

If you use one of these schedulers, then to continue running MATLAB jobs on your cluster, update your HPC scheduler to Microsoft HPC Pack 2019.

For more information about configuring supported HPC Pack clusters to run MATLAB jobs, see Configure for Microsoft HPC Pack (MATLAB Parallel Server).

Additional configuration required for multi-release MATLAB Job Scheduler clusters

Behavior change

When you configure a MATLAB Job Scheduler cluster to run jobs with multiple MATLAB releases, supporting releases R2022b and earlier now requires additional configuration.

Before R2026b, you configure a MATLAB Job Scheduler cluster to support multiple MATLAB releases by using only the MJS_ADDITIONAL_MATLABROOTS parameter. Starting in R2026b, to support releases R2022b and earlier, you must also provide values for these parameters in the mjs_def file on the head node:

  • MJS_ADDITIONAL_SUPPORTED_RELEASES — Specify the supported older releases.

  • MJS_LEGACY_JAVA — Specify a Java Runtime Environment (JRE) path for scheduler and worker processes. The JRE must be version 8 or 11.

For additional configuration steps, see Use Multiple MATLAB Parallel Server Releases in Cluster (MATLAB Parallel Server).

GPU Computing

 Monitor GPU utilization, memory usage, and processes in real time

The GPU Monitor is an interactive tool for monitoring your GPU utilization, memory usage, and processes in real time. Monitor your GPU to verify that your code is running on the GPU and identify inefficiencies and bottlenecks in your GPU computing code.

With the GPU Monitor, you can:

  • Observe GPU utilization and memory usage during MATLAB computations. The monitor can plot these metrics for one GPU or for multiple GPUs simultaneously.

  • See which processes on your machine are using your GPUs.

  • Access GPU device information, including the state, capabilities, and driver version.

For information about how to interpret and respond to the metrics you see in the GPU Monitor, see Improve Performance Using GPU Monitor Metrics.

Pass GPU data to Python to improve performance of workflows that call Python from MATLAB​

You can now use these functions to convert gpuArray data in MATLAB to Python® GPU data types. These functions improve performance of GPU workflows that call Python from MATLAB​ by removing the need to gather GPU data to host memory.

  • pycupyarray — converts an array in MATLAB to a CuPy array, if the CuPy module is installed in the Python environment.​

  • pydlpack — converts a gpuArray to a DLPack capsule.​

You can also send GPU data from Python back to MATLAB by calling the gpuArray function on a CuPy array or a DLPack capsule.

For more information about transferring GPU data between MATLAB and Python, see Pass GPU Data Between MATLAB and Python.

GPU Functionality in MATLAB: Use functions with new and enhanced gpuArray support

These MATLAB and Parallel Computing Toolbox functions have new and enhanced gpuArray support:

  • arrayfun — You can now use functions defined in a namespace with arrayfun. Using a namespace helps organize code and creates more robust names for the items contained inside it.

  • isapprox

  • mustBeNonzeroLengthText

  • mustBeSorted

  • mustBeText

  • mustBeTextScalar

  • norm — You can now calculate the generalized vector p-norm for 0 < p < 1.

  • parallel.gpu.RandStream.list — You can now specify an output argument to return a table of all available GPU generator algorithms. The table provides detailed information for each algorithm, including the generator name, multiple-stream support, and description.

For more information, see Run MATLAB Functions on a GPU. For a full list of MATLAB functions that accept GPU arrays, see Function List (GPU Arrays).

GPU Functionality in Signal Processing and Wireless Communications: Use functions with new and enhanced gpuArray support

Signal Processing Toolbox

These Signal Processing Toolbox functions have new and enhanced gpuArray support:

  • emd (Signal Processing Toolbox)

  • vmd (Signal Processing Toolbox)

  • chirp (Signal Processing Toolbox)

  • tsa (Signal Processing Toolbox)

  • tfridge (Signal Processing Toolbox) — Improved performance when extracting time-frequency ridges without a penalty for changing frequency.

  • signalTimeFrequencyFeatureExtractor (Signal Processing Toolbox) — You can now use the emd and vmd time-frequency analysis methods.

For a full list of Signal Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Signal Processing Toolbox).

Communications Toolbox

This Communications Toolbox™ function has new gpuArray support:

For a full list of Communications Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Communications Toolbox).

DSP System Toolbox

These DSP System Toolbox™ System objects have new and enhanced gpuArray support:

For a full list of DSP System Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (DSP System Toolbox).

5G Toolbox

These 5G Toolbox™ functions have new gpuArray support:

For a full list of 5G Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (5G Toolbox).

GPU Functionality in Image Processing Toolbox: Use functions with new and enhanced gpuArray support

These Image Processing Toolbox™ functions have new and enhanced gpuArray support:

  • graythresh (Image Processing Toolbox)

  • otsuthresh (Image Processing Toolbox)

  • imerode (Image Processing Toolbox), imdilate (Image Processing Toolbox), imopen (Image Processing Toolbox), imclose (Image Processing Toolbox), imtophat (Image Processing Toolbox), imbothat (Image Processing Toolbox) — You can now use these functions to apply 3-D morphological operations using 3-D structuring elements.

For a full list of Image Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Image Processing Toolbox).

GPU Functionality in Deep Learning Toolbox: Speed up training using automatic mixed precision

This Deep Learning Toolbox™ function has enhanced GPU support:

  • trainnet (Deep Learning Toolbox) — You can now accelerate training and reduce memory usage on a GPU using automatic mixed precision. Set the TrainingPrecision option to "automatic-mixed" using the trainingOptions (Deep Learning Toolbox) function to use half-precision floating-point arithmetic.

For a full list of Deep Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Deep Learning Toolbox).

GPU Functionality in Reinforcement Learning Toolbox: Set the UseGPUForLearning agent property to calculate gradients using GPU

Agents in Reinforcement Learning Toolbox™ have this new property:

  • UseGPUForLearning — Set to "on" or "auto" to configure agents for GPU usage during training. This allows you to more easily configure the learnables, targets, and optimizers of an agent for GPU usage during training.

For more information, see Train Agents Using Parallel Computing and GPUs (Reinforcement Learning Toolbox).

Index into contiguous submatrices of sparse GPU arrays

You can now reference contiguous submatrices, that is, rectangular blocks of adjacent elements of sparse gpuArray objects.

To access a submatrix of a sparse gpuArray, index into the array using consecutive row and column indices. For example, access the central 3-by-3 submatrix of a 10-by-10 sparse gpuArray.

A = gpuArray.speye(10);
A(4:6,4:6)
  3×3 sparse gpuArray double matrix (3 nonzeros)

   (1,1)        1
   (2,2)        1
   (3,3)        1

For more information about indexing sparse GPU arrays, see Work with Sparse Arrays on a GPU.

GPU Workflow Example: Accelerate seismic migration using CUDA code

This new example shows how to accelerate your code by implementing long-running parts in CUDA® code and using advanced CUDA features, including cooperative groups:

Matrix-Matrix Multiplication: Improved performance on Blackwell GPUs

Matrix-matrix multiplication of gpuArray data shows improved performance on Linux with NVIDIA® Blackwell architecture GPUs (compute capability 10.0 to 12.0). For example, multiplying two 8192-by-8192 matrices on a Blackwell GPU is about 10.3x faster than in R2026a.

function t = timeMatrixMultiply

% Reset GPU and set random number generator.
reset(gpuDevice);
gpurng("default");

% Create random matrices on GPU.
n = 8192;
A = rand(n,"gpuArray");
B = rand(n,"gpuArray");

% Time matrix multiplication.
t = gputimeit(@() A*B)
end

The approximate execution times are:

R2026a: 0.71 s

R2026b: 0.069 s

The code was timed on a Linux, Debian® 13.5, AMD® EPYC 9115 16-core processor @ 2.6 GHz test system with 64 GB of memory and an NVIDIA RTX Pro 6000 GPU by calling the timeMatrixMultiply function.

The performance improvement you observe will depend on your GPU hardware.

Sparse arrays: Improved performance when creating sparse GPU arrays

Creating a sparse gpuArray by calling the sparse function on a dense gpuArray shows improved performance. For example, calling sparse on a 10000-by-10000 dense gpuArray is about 12x faster than in the previous release.

function t = timeSparseGpuArray

% Reset GPU and set random number generator.
reset(gpuDevice);
gpurng("default")

% Create dense gpuArray.
sz = 10000;
dense = rand(sz,sz,"gpuArray");

% Time creating sparse gpuArray.
f = @() sparse(dense);
t = gputimeit(f)
end

The approximate execution times are:

R2026a: 0.349 s

R2026b: 0.029 s

The code was timed on a Windows 11, AMD EPYC 7262 8-core processor @ 3.2 GHz test system with 64 GB of memory and an NVIDIA RTX A5000 GPU by calling the timeSparseGpuArray function.

The performance improvement you observe will depend on your GPU hardware.

 GPU computing in MATLAB upgraded to CUDA 13.1

GPU computing in MATLAB now uses CUDA 13.1. You can use new CUDA features when you write custom CUDA code for running in MATLAB via mexcuda and parallel.gpu.CUDAKernel.

 Compatibility Considerations

 Functionality being removed or changed

Support for Maxwell, Pascal, and Volta architecture GPUs is removed

Errors

Starting in R2026b, Maxwell, Pascal, and Volta architecture (compute capability 5.0 to 7.0) GPUs are no longer supported. The supported range of compute capabilities is 7.5 to 12.x.

For more information, see GPU Computing Requirements.

Change to supported CUDA Toolkit

Errors

CUDA Toolkit version 12.8 is no longer supported.

  • Recompile any MEX functions and PTX code that were compiled using the mexcuda function in previous versions of MATLAB.

  • If you previously installed the CUDA Toolkit, install CUDA Toolkit version 13.1 instead.

For more information, see Install CUDA Toolkit (Optional).

Change to parallel.gpu.RandStream.list

Behavior change

Calling parallel.gpu.RandStream.list without an output argument now displays a table of all available GPU generator algorithms, with more detailed information about each algorithm: its name, whether it has multiple-stream support, and a description. In previous releases, parallel.gpu.RandStream.list displayed only the available generator algorithms and their descriptions.

Change to supported PTX files for CUDA kernels

Errors

Creating a CUDAKernel object using the parallel.gpu.CUDAKernel function from a PTX file compiled using PTX ISA version 3.0 or earlier is no longer supported.

Recompile your PTX files using the mexcuda function.

mexcuda -ptx myCUFile.cu

R2026a

New Features, Bug Fixes, Compatibility Considerations

Parallel Language and Cluster Computing

 Parallel Computing Onramp: Free, self-paced, interactive course

Parallel Computing Onramp is a free, self-paced, interactive course that helps you get started with accelerating your MATLAB code using parallel computing.

Parallel Computing Onramp features hands-on exercises using MATLAB in your web browser. The course teaches you to:

  • Create and modify parallel pools.

  • Convert for loops into parfor loops.

  • Explore variable classification in parfor loops.

  • Explore how parallel overhead impacts execution time.

The Parallel Computing Onramp.

parfor-Loops: Use consecutive decreasing integers as loop index variables

You can now use consecutive decreasing integers as loop index variables in a parfor-loop. Using such variables allows you to simplify code that requires reverse iteration.

For example, this parfor-loop uses reverse iteration to square values of n between 10 and 1. The loop index variable n has a step value of -1.

parfor n = 10:-1:1
  out(n) = n^2;
end 
MATLAB continues to execute parfor-loop iterations in a nondeterministic order.

Thread-Based Parallel Pool: Monitor pool activity with Pool Dashboard

Use the Pool Dashboard to collect and visualize monitoring data for thread-based parallel pools. For more information, see Pool Dashboard.

You can also use a command line interface to collect pool activity monitoring data on thread-based interactive parallel pools. For details, see parallel.pool.ActivityMonitor.

Thread-Based Environment: Use new functionality on thread workers

These MATLAB functions and objects now have thread-based support.

canUseGPUmatlab.io.fits.getImgTypematlab.io.fits.readRecordmatlab.io.fits.imgCompressmatlab.io.fits.getNumCols
canUseParallelPoolmatlab.io.fits.insertImgmatlab.io.fits.writeCommentmatlab.io.fits.isCompressedImgmatlab.io.fits.getNumRows
fitsdispmatlab.io.fits.readImgmatlab.io.fits.writeDatematlab.io.fits.setCompressionTypematlab.io.fits.insertCol
fitsinfomatlab.io.fits.setBscalematlab.io.fits.writeKeymatlab.io.fits.setHCompScalematlab.io.fits.insertATbl
fitsreadmatlab.io.fits.writeImgmatlab.io.fits.writeKeyUnitmatlab.io.fits.setHCompSmoothmatlab.io.fits.insertBTbl
fitswritematlab.io.fits.deleteKeymatlab.io.fits.writeHistorymatlab.io.fits.setTileDimmatlab.io.fits.readATblHdr
matlab.io.fits.closeFilematlab.io.fits.deleteRecordmatlab.io.fits.copyHDUmatlab.io.fits.createTblmatlab.io.fits.readBTblHdr
matlab.io.fits.createFilematlab.io.fits.getHdrSpacematlab.io.fits.deleteHDUmatlab.io.fits.deleteColmatlab.io.fits.readCol
matlab.io.fits.deleteFilematlab.io.fits.readCardmatlab.io.fits.getHDUnummatlab.io.fits.deleteRowsmatlab.io.fits.setTscale
matlab.io.fits.fileNamematlab.io.fits.readKeymatlab.io.fits.getHDUtypematlab.io.fits.insertRowsmatlab.io.fits.writeCol
matlab.io.fits.fileModematlab.io.fits.readKeyCmplxmatlab.io.fits.getNumHDUsmatlab.io.fits.getAColParmsmatlab.io.fits.getConstantValue
matlab.io.fits.openFilematlab.io.fits.readKeyDblmatlab.io.fits.movAbsHDUmatlab.io.fits.getBColParmsmatlab.io.fits.getVersion
matlab.io.fits.openDiskFilematlab.io.fits.readKeyLongLongmatlab.io.fits.movNamHDUmatlab.io.fits.getColNamematlab.io.fits.getOpenFiles
matlab.io.fits.createImgmatlab.io.fits.readKeyLongStrmatlab.io.fits.movRelHDUmatlab.io.fits.getColTypeode
matlab.io.fits.getImgSizematlab.io.fits.readKeyUnitmatlab.io.fits.writeChecksummatlab.io.fits.getEqColTypesystem

For more information, see Run MATLAB Functions in Thread-Based Environment.

Pool Dashboard: Run code directly in Pool Dashboard

Run and monitor MATLAB code directly from the Pool Dashboard using the Enter code to run and monitor box. For an example that shows how to use the Enter code to run and monitor box, see Investigate parfor -Loop with Pool Dashboard.

The Pool Dashboard shows the execution timeline, list of constructs and their parent functions, and a summary of worker activity. The run and monitor box is highlighted.

Local Parallel Pools: Support for more than 64 workers on Windows 11

You can now create local parallel pools with more than 64 workers on Windows 11 machines that have more than 64 cores. The default number of workers for local parallel pools you create with the Threads and Processes profiles on a multicore Windows 11 machine is now equal to the number of physical cores. In previous releases, MATLAB creates local pools with only 64 workers even on machines with more cores.

Cluster Profile Manager: View additional NumWorkers property information for local profiles

The Cluster Profile Manager now displays how MATLAB determines the default value of the NumWorkers property of the Processes and Threads profiles, unless you have modified this property. This insight helps you understand the default maximum number of workers available for a local parallel pool, especially on CPUs that have both performance and efficiency cores.

The Cluster Profile Manager now also shows the number of physical cores MATLAB detects on your computer, as well as the number of performance cores, if any. To view this information, point to the information button next to the NumWorkers property value.

For more information about how MATLAB sets the default number of workers for local parallel environments, see Determine Default Pool Size for Local Parallel Environments.

The Cluster Profile Manager shows the Processes profile, with an information box next to the NumWorkers property value that displays "Detected 16 physical cores. Detected 10 high-performance cores."

Cluster Validation: Validate cluster object

You can now validate parallel.Cluster objects using the validate object function. You can specify which validation stages to run, set the number of workers to use, and write the validation results to a file.

Cluster Profiles: Delete profiles programmatically

Use the parallel.deleteProfile function to delete cluster profiles programmatically.

Parallel Workflows: New Examples and Topics

Use these new examples to learn about parallel computing:

Use these new topics to learn about parallel computing:

Cloud Center: Streamlined user interface for creating and using MATLAB Parallel Server clusters

MATLAB Parallel Server on Cloud Center has a new user interface and underlying infrastructure. The operating system is updated to Ubuntu® 24.04, providing improved performance, security, and user experience.

The configuration settings for the cloud machine configuration in Cloud Center are adapted from the MATLAB Parallel Server on Amazon Web Services reference architecture available on GitHub®.

Cloud Clusters: Promote and demote parallel jobs

You can now use the promote and demote functions to manage job priority in the queue of a MATLAB Job Scheduler cluster running in the cloud. This includes MATLAB Parallel Server cloud clusters created using MathWorks® Cloud Center or MathWorks reference architecture templates.

MATLAB Parallel Server on AWS: Customize and deploy your own Amazon Machine Image

You can now customize and build your own Linux Amazon® Machine Image (AMI) for running MATLAB Parallel Server on AWS®, using the same scripts that form the basis of the build process for MathWorks prebuilt images. A HashiCorp® Packer® template generates the machine image.

To build your own machine image using MathWorks scripts, see the build scripts and instructions on GitHub for customizing and deploying your own AMI.

Cluster Administration: Simplify MATLAB Job Scheduler cluster connectivity using SOCKS5 Proxy

You can now use a SOCKS5 proxy server to streamline connectivity between MATLAB clients and MATLAB Job Scheduler clusters. The new parallelserverproxy tool forwards all Parallel Computing Toolbox network traffic through a single endpoint. This capability reduces the need for multiple open ports and simplifies firewall and network configurations, especially for clients outside the cluster’s virtual network or in cloud and hybrid environments. The parallelserverproxy also authenticates clients using mutual TLS and encrypts communication between the MATLAB clients and cluster.

For more information, see Configure SOCKS5 Proxy for MATLAB Job Scheduler (MATLAB Parallel Server).

Support for MPICH: Use MPICH version 4.2.3

Parallel Computing Toolbox and MATLAB Parallel Server now come with MPICH version 4.2.3 for use on Linux and Mac operating systems.

Cluster Administration: New and updated troubleshooting topics

Use these new and updated topics to resolve issues with MATLAB Job Scheduler and MATLAB Parallel Server in third-party scheduler clusters.

Tall Arrays: Use functions with tall tables

You can now use these functions with tall tables: movmad, movmax, movmean, movmedian, movmin, movprod, movstd, movsum, and movvar.

For more information, see Tall Arrays for Out-of-Memory Data.

Tall Arrays: Improved performance when joining tables

The join and innerjoin functions show improved performance when the first input is a tall table. The second input argument can be an in-memory table or the result of a reduction operation on a tall table. For details, see the tall array performance improvements described in the MATLAB release notes.

Distributed Arrays: Use functions with new and enhanced distributed array support

These functions have new distributed array support:

These functions have enhanced distributed array support:

  • isbetween — You can now use the DataVariables and OutputFormat name-value arguments and specify distributed tables and timetables as input arguments.

  • ichol — You can now specify nonsymmetric and non-hermitian matrices as input.

  • movmad, movmax, movmean, movmedian, movmin, movprod, movstd, movsum, and movvar — You can now use these functions with distributed tables and timetables.

For more information, see Run MATLAB Functions with Distributed Arrays.

 Functionality being removed or changed

addAttachedFiles now updates already attached files

Behavior change

The addAttachedFiles function now updates files that are already attached to the parallel pool. In previous releases, the addAttachedFiles function did not update already attached files.

GPU Computing

  Support for Blackwell GPU Architectures: Update to NVIDIA CUDA 12.8

Parallel Computing Toolbox now uses CUDA version 12.8, which supports NVIDIA GPUs with compute capability up to 12.x, including Blackwell architecture GPUs. For more information, see GPU Computing Requirements.

 Compatibility Considerations

CUDA Toolkit version 12.2 is no longer supported. Recompile any MEX functions and PTX code that were compiled using the mexcuda function in previous versions of MATLAB. For more information, see Change to supported CUDA Toolkit.

GPU Functionality in MATLAB: Use functions with new and enhanced gpuArray support

These MATLAB functions have new gpuArray support:

These MATLAB functions have enhanced gpuArray support:

  • besseli — You can now specify negative equation orders.

  • expm — You can now use this function with diagonal sparse gpuArray objects.

  • issorted — You can now determine if the elements of the first column of a gpuArray matrix are sorted in ascending order by using the 'rows' option. If a column contains repeated elements, then the issorted function uses the ordering of the elements in the next column to the right to determine the sorting order.

  • kron — You can now use this function with sparse gpuArray objects.

  • ldivide— You can now divide sparse gpuArray objects by a scalar.

  • max, min — You can now return the maximum or minimum value of a sparse gpuArray object by using the "all" option.

  • plus, + — You can now append a gpuArray object to a string without first calling gather on the gpuArray object.

For more information, see Run MATLAB Functions on a GPU. For a full list of MATLAB functions that accept GPU arrays, see Function List (GPU Arrays).

GPU Functionality in Statistics and Machine Learning Toolbox: Use functions with new and enhanced gpuArray support

These Statistics and Machine Learning Toolbox™ functions have new and enhanced gpuArray support:

  • cvpartition (Statistics and Machine Learning Toolbox) — You can now partition data for cross-validation. The object functions repartition (Statistics and Machine Learning Toolbox), summary (Statistics and Machine Learning Toolbox), test (Statistics and Machine Learning Toolbox), and training (Statistics and Machine Learning Toolbox) accept the resulting cvpartition object and can also execute on a GPU.

  • fitrgp (Statistics and Machine Learning Toolbox) — You can now fit a Gaussian process regression (GPR) model. Most object functions of the models RegressionGP (Statistics and Machine Learning Toolbox), CompactRegressionGP (Statistics and Machine Learning Toolbox), and RegressionPartitionedGP (Statistics and Machine Learning Toolbox) now support GPU array input arguments. The functions that do not are lime (Statistics and Machine Learning Toolbox) and shapley (Statistics and Machine Learning Toolbox).

  • fitdist (Statistics and Machine Learning Toolbox), mle (Statistics and Machine Learning Toolbox) — You can now fit Rician distributions to data and estimate the maximum likelihood of Rician distributions.

  • fitcecoc (Statistics and Machine Learning Toolbox) — You can now specify kernel learners when you create a CompactClassificationECOC (Statistics and Machine Learning Toolbox) or ClassificationPartitionedKernelECOC (Statistics and Machine Learning Toolbox) model object using fitcecoc.

  • ncx2cdf (Statistics and Machine Learning Toolbox), ncx2pdf (Statistics and Machine Learning Toolbox), ncx2inv (Statistics and Machine Learning Toolbox), ncx2stat (Statistics and Machine Learning Toolbox)

  • shapley (Statistics and Machine Learning Toolbox), fit (Statistics and Machine Learning Toolbox) — The shapley (Statistics and Machine Learning Toolbox) and fit (Statistics and Machine Learning Toolbox) functions accept GPU array input arguments when the machine learning model is a regression or binary classification linear model listed below, and the function uses an interventional algorithm (Method="interventional"). The supported models are:

For a full list of Statistics and Machine Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Statistics and Machine Learning Toolbox).

GPU Functionality in Signal Processing: Use functions with new and enhanced gpuArray support

Signal Processing Toolbox

This Signal Processing Toolbox function has new gpuArray support:

For a full list of Signal Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Signal Processing Toolbox).

Wavelet Toolbox

This Wavelet Toolbox function has new gpuArray support:

For a full list of Wavelet Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Wavelet Toolbox).

DSP System Toolbox

This DSP System Toolbox System object has new gpuArray support:

GPU Functionality in Wireless Communications: Use functions with new and enhanced gpuArray support

Communications Toolbox

These Communications Toolbox functions and System objects have new and gpuArray support:

For a full list of Communications Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Communications Toolbox).

5G Toolbox

These 5G Toolbox functions have new and enhanced gpuArray support:

  • nrCDLChannel (5G Toolbox) — You can now enable GPU processing when the ChannelFiltering property is set to false by setting the UseGPU property to 'on' or 'auto'.

  • nrPDSCHPrecode (5G Toolbox)

  • nrResourceGrid (5G Toolbox)

For a full list of 5G Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (5G Toolbox).

GPU arrayfun: Support for eps function with prototype array

You can now use the eps function and specify a prototype array by using the like syntax in the function you apply with arrayfun. Using the eps function with a prototype array p returns the positive distance from 1.0 to the next larger floating-point number of the same precision as p, with the same data type and complexity (real or complex) as p.

For example, this function determines which elements of input array x equal the square root of eps.

function y = compareEps(x)
y = x^2 == eps(like=x);
end

% Generate input gpuArray matrix A.
A = gallery("lauchli",500);
A = gpuArray(A);

% Apply the compareEps function to each element of A.
B = arrayfun(@compareEps,A);

GPU Workflow Examples: New GPU computing examples

  • Write Portable GPU Code — This example shows how to write robust code that can run on machines with or without a GPU.

GPU Device: Driver model property reports MCDM

When you inspect the properties of your GPU using the gpuDevice function, the DriverModel property is now 'MCDM' for devices using the Microsoft Compute Driver Model (MCDM).

GPU find: Improved performance when finding one element

The find function shows improved performance when you use it to find one index of a gpuArray. For example, this code finds the first element of a gpuArray that is above 0.99. The code is about 2.3x faster than in the previous release.

function t = timeFind

% Reset GPU.
reset(gpuDevice)

% Create random input data.
x = rand(1e7,1,"gpuArray");

% Time finding one index.
t = gputimeit(@() find(x>0.99,1))

end

The approximate execution times are:

R2025b: 11.4 ms

R2026a: 4.9 ms

The code was timed on a Windows 11, Intel Xeon® W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeFind function.

GPU discretize: Improved performance

The discretize function shows improved performance when grouping gpuArray data into bins. The improvement is more pronounced for smaller arrays. For example, this code, which groups the data into 10 bins, is about 1.5x faster than in the previous release.

function t = timeDiscretize

% Reset GPU.
reset(gpuDevice)

% Create random input data.
x = rand(1000,"gpuArray");

% Time discretize function.
t = gputimeit(@() discretize(x,10));

end

The approximate execution times are:

R2025b: 699 microseconds

R2026a: 464 microseconds

The code was timed on a Windows 11, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDiscretize function.

GPU diff: Improved performance

The diff function shows improved performance when calculating the second- or higher-order differences between elements of gpuArray data. For example, this code, which calculates the third-order difference between elements of gpuArray data, is about 3.8x faster than in the previous release.

function t = timeDiff

% Reset GPU.
reset(gpuDevice)

% Create random input data.
x = rand(1e8,1,"gpuArray");

% Choose difference order.
N = 3;

% Time diff function.
t = gputimeit(@() diff(x,N));

end

The approximate execution times are:

R2025b: 11.7 ms

R2026a: 3.1 ms

The code was timed on a Windows 11, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeDiff function.

 Functionality being removed or changed

Change to supported CUDA Toolkit

Errors

CUDA Toolkit version 12.2 is no longer supported.

  • Recompile MEX functions and PTX code compiled using the mexcuda function in previous versions of MATLAB.

  • If you previously installed the CUDA Toolkit, install CUDA Toolkit version 12.8 instead.

For more information, see Install CUDA Toolkit (Optional).

R2025b

Bug Fixes, Compatibility Considerations

Quality and stability improvements

R2025b delivers quality and stability improvements, building on the new features introduced in R2025a.

 Functionality being removed or changed

Support for Windows Compute Cluster Server 2003, Windows HPC Server 2008, Windows HPC Server 2008 R2, Microsoft HPC Pack 2012, and Microsoft HPC Pack 2012 R2 will be removed

Still runs

Support for these Microsoft HPC schedulers will be removed in a future release:

  • Windows Compute Cluster Server 2003

  • Windows HPC Server 2008

  • Windows HPC Server 2008 R2

  • Microsoft HPC Pack 2012

  • Microsoft HPC Pack 2012 R2

At that time, if you use one of these schedulers and want to continue running MATLAB jobs on your cluster, update your HPC scheduler to Microsoft HPC Pack 2016 or Microsoft HPC Pack 2019.

For more information about configuring supported HPC Pack clusters to run MATLAB jobs, see Configure for Microsoft HPC Pack.

R2025a

New Features, Bug Fixes, Compatibility Considerations

 Pool Dashboard: Collect and analyze activity in parallel pools

Pool Dashboard is a new tool that you can use to collect and visualize monitoring data for interactive parallel pools. Using the Pool Dashboard, you can:

  • Collect monitoring data on how pool workers execute parallel constructs like parfor, parfeval, and spmd.

  • Track the amount of data (in bytes) the client and workers send and receive.

  • Understand the time each worker spends processing their portion of the parallel code.

  • Examine communication patterns and identify bottlenecks and load-balancing issues.

  • Save pool monitoring results to compare the impact of code improvements.

For more information, see Pool Dashboard.

You can also use a command line interface to collect pool activity monitoring data on interactive and batch pools. For details, see parallel.pool.ActivityMonitor.

The Pool Dashboard shows the execution timeline, list of constructs and their parent functions, and a summary of worker activity.

 Thread-Based Parallel Pool: Profile parallel code on thread workers

You can now use the mpiprofile command to profile parallel code on workers of a thread-based parallel pool. For more information about profiling your parallel code, see Profiling Parallel Code.

Cluster Profile Manager: Customize Threads profile

Manage the local machine Threads profile in the Cluster Profile Manager. You can customize the preconfigured Threads profile by editing the profile properties in the Cluster Profile Manager. You can also create additional Threads profiles.

The Cluster Profile Manager shows the Threads profile selected.

Profile Validation: Programmatically validate profile

Programmatically validate your profile using the parallel.validateProfile function. You can specify which validation stages to run, set the number of workers to use, and write the validation results to a file.

For example, you can validate the default profile.

parallel.validateProfile
Beginning validation for cluster profile 'Processes'
Cluster connection test (parcluster)
   Stage started at 15:35:45.
   Finished at 15:35:45.
..........................................................................PASSED

Job test (createJob)
   Stage started at 15:35:45.
   Finished at 15:36:04.
..........................................................................PASSED

SPMD job test (createCommunicatingJob)
   Stage started at 15:36:13.
   Job ran with 6 workers.
   Finished at 15:36:57.
..........................................................................PASSED

Pool job test (createCommunicatingJob)
   Stage started at 15:37:13.
   Job ran with 6 workers.
   Finished at 15:38:03.
..........................................................................PASSED

Parallel pool test (parpool)
   Stage started at 15:38:20.
   Connected to parallel pool with 6 workers.
   Parallel pool using the 'Processes' profile is shutting down.
   Parallel pool ran with 6 workers.
   Finished at 15:39:24.
..........................................................................PASSED

Finished cluster profile validation with status: PASSED

PollableDataQueue Objects: Enhanced functionality to simplify data queue workflows

You can now create parallel.pool.PollableDataQueue objects that allow the client or any worker in the pool to receive data.

The parallel.pool.PollableDataQueue function has a new Destination name-value argument that enables you to specify the destination behavior of the PollableDataQueue object. If you want to send data to the client or any worker, set the Destination argument to "any". Then any worker with a copy of the PollableDataQueue object can poll it to receive data. This new type of PollableDataQueue object simplifies the process of sending data to workers during asynchronous parfeval computations.

The PollableDataQueue object also has a new close object function to close a PollableDataQueue object. Closing a PollableDataQueue object changes its new isClosed property to true, which prevents you from sending more data to the PollableDataQueue object.

For examples that show how to use the new type of PollableDataQueue object, see Send Messages to Workers Using Pollable Data Queues and Transfer Data Between Workers Using Pollable Data Queues.

Parallel Pools: Partition pools from an existing parallel pool

Use the partition function to divide an existing parallel pool into subset pools that allow you to utilize specific resources from the existing pool. Both the partitioned pools and the input pool schedule work on the same underlying collection of workers.

You can create pools that target specific workers to assign them specific roles or tasks. Additionally, you can create multiple pools to execute multiple parallel workflows simultaneously.

For example, you can partition a parallel pool to assign one worker to each cluster host in the pool.

hostPool = partition(pool,"MaxNumWorkersPerHost",1);
For more details, see Partition Parallel Pools to Optimize Resource Use.

Parallel Pools: Specify pool for parfor, spmd, and Composite functions

The parfor, spmd, and Composite functions now accept parallel.Pool objects as input arguments. You can specify a pool object when you want to run computations on a pool other than the pool the gcp function returns.

parpool Function: Add folders to workers search path

You can now add folders to the MATLAB search path of pool workers at the time of pool creation using the AdditionalPaths name-value argument of the parpool function. This argument ensures that pool workers look for code files, data files, or model files in the correct locations.

 GPU Functionality: Use functions with new and enhanced gpuArray support

MATLAB

These MATLAB functions have new gpuArray support:

These MATLAB functions have enhanced gpuArray support:

For more information, see Run MATLAB Functions on a GPU. For a full list of MATLAB functions that accept GPU arrays, see Function List (GPU Arrays).

Statistics and Machine Learning Toolbox

These Statistics and Machine Learning Toolbox functions have new gpuArray support:

  • fitrkernel (Statistics and Machine Learning Toolbox), fitckernel (Statistics and Machine Learning Toolbox) — These functions and most of the object functions of RegressionKernel (Statistics and Machine Learning Toolbox) and ClassificationKernel (Statistics and Machine Learning Toolbox) models now support GPU array input arguments, allowing you to execute these functions on a GPU. The object functions that do not are incrementalLearner (Statistics and Machine Learning Toolbox), lime (Statistics and Machine Learning Toolbox), and shapley (Statistics and Machine Learning Toolbox).

  • ocsvm (Statistics and Machine Learning Toolbox) — This function and the isanomaly (Statistics and Machine Learning Toolbox) object function of OneClassSVM (Statistics and Machine Learning Toolbox) now support GPU array input arguments, enabling you to execute these functions on a GPU.

For a full list of Statistics and Machine Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Statistics and Machine Learning Toolbox).

Signal Processing Toolbox

These Signal Processing Toolbox functions have new gpuArray support:

For a full list of Signal Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Signal Processing Toolbox).

Wavelet Toolbox

This Wavelet Toolbox function has new gpuArray support:

  • icwt (Wavelet Toolbox)

For a full list of Wavelet Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Wavelet Toolbox).

5G Toolbox

These 5G Toolbox functions have new gpuArray support:

For a full list of 5G Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (5G Toolbox).

Predictive Maintenance Toolbox

These Predictive Maintenance Toolbox™ functions have new gpuArray support:

For a full list of Predictive Maintenance Toolbox functions that accept GPU arrays, see Functions List (GPU Arrays) (Predictive Maintenance Toolbox).

Antenna Toolbox

These Antenna Toolbox™ functions have new gpuArray support:

  • raytrace (Antenna Toolbox), coverage (Antenna Toolbox), sinr (Antenna Toolbox), sigstrength (Antenna Toolbox), link (Antenna Toolbox), pathloss (Antenna Toolbox) — These functions support ray tracing analysis on a GPU when you specify a RayTracing (Antenna Toolbox) propagation model object as input and the UseGPU property of the object is "on" or "auto".

For a full list of Antenna Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Antenna Toolbox).

Phased Array System Toolbox

These Phased Array System Toolbox™ functions have new gpuArray support:

  • ambgfun (Phased Array System Toolbox)

  • pambgfun (Phased Array System Toolbox)

For a full list of Phased Array System Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Phased Array System Toolbox).

GPU Arrays: Reduce memory usage with single-precision sparse GPU arrays

You can now create and use single-precision sparse GPU arrays. Sparse matrices provide efficient storage of data that has a large percentage of zeros and reduce computation time by eliminating operations on zero elements. Single-precision sparse GPU arrays allow you to reduce memory usage and accelerate calculations by taking advantage of your GPU's single-precision floating-point units (FPUs).

You can convert a sparse gpuArray to a single-precision array using the single function, or you can create a sparse, single-precision gpuArray directly by specifying the typename argument as "single" when you create a sparse gpuArray. For example, this code creates a 1000-by-1000 random, single-precision, sparse gpuArray with density 0.1.

R = gpuArray.sprand(1000,1000,0.1,"single");

For more information, see Work with Sparse Arrays on a GPU.

GPU arrayfun: Support for like syntax of intmin, intmax, realmin, and realmax

You can now use the intmin, intmax, realmin, and realmax functions and specify a prototype array using the like syntax in functions you apply using the arrayfun function.

For example, this function uses intmin and intmax to determine whether elements of integer gpuArray x are saturated.

function saturated = findSaturated(x)
saturated = (x == intmin(like=x)) | (x == intmax(like=x));
end

% Generate random integer gpuArray that includes saturated values.
x = randi([-200 200],1e6,1,"int8","gpuArray");

% Find saturated values.
saturated = arrayfun(@findSaturated,x);

GPU Workflow Examples: New and updated GPU computing examples

This updated example shows how to measure key performance characteristics of your GPU hardware:

This new example shows how to use GPU computing to accelerate data preprocessing and deep learning for predictive maintenance workflows:

These new and updated examples show how to use GPU computing to accelerate 5G simulations:

This new example shows how to perform accelerated ray tracing analysis using a GPU:

GPU arrayfun: Support for using P-code files in compiled standalone applications

You can now call the arrayfun function inside P-code files or use arrayfun to evaluate functions obfuscated as a P-code file in standalone applications compiled using MATLAB Compiler™.

For more information about packaging a MATLAB function into a standalone application, see Create Standalone Application from MATLAB (MATLAB Compiler).

GPU Arrays: Gather GPU arrays interactively from workspace

You can now gather gpuArray objects by right-clicking the variable in the workspace, and then selecting Gather from GPU. This action is equivalent to calling the gather function on a gpuArray object.

Context menu for a GPU array variable in the workspace with the pointer on the Gather from GPU option

GPU Statistics Functions: Improved performance for mean, movmean, movstd, and movvar

Mean

These syntaxes of the mean function show improved performance on a GPU:

  • M = mean(A)

  • M = mean(A,"all")

  • M = mean(A,dim), where dim is a scalar double

Improvements are greater when you operate on smaller arrays. For example, computing the mean of a matrix in this code is about 2.4x faster than in the previous release.

function timeMean

% Reset GPU device.
reset(gpuDevice)

% Create random gpuArray data.
X = rand(200,"gpuArray");

% Time mean function.
f = @() mean(X,"all");
gputimeit(f)

end

The approximate execution times are:

R2024b: 63.9 microseconds

R2025a: 26.9 microseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeMean function.

Moving Mean, Standard Deviation, and Variance

The movmean, movstd, and movvar functions show improved performance on a GPU. For example, computing the five-point centered moving average of a matrix in this code is about 1.6x faster than in the previous release.

function timeMovmean

% Reset GPU device.
reset(gpuDevice)

% Create random gpuArray data.
X = rand(1e4,"gpuArray");

% Time movmean function.
f = @() movmean(X,5);
gputimeit(f)

end

The approximate execution times are:

R2024b: 14.8 milliseconds

R2025a: 9.2 milliseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeMovmean function.

GPU histcounts: Improved performance

The histcounts function shows improved performance on a GPU. Improvements are greater when you operate on smaller arrays. For example, partitioning values into 21 bins and returning the bin counts and bin edges in the following code is about 1.7x faster than in the previous release:

function timeHistcounts

% Reset GPU device.
reset(gpuDevice)

% Create random input data.
X = rand(1000,"gpuArray");

% Time histcounts function.
f = @() histcounts(X,21);
gputimeit(f,2)

end

The approximate execution times are:

R2024b: 1.2 milliseconds

R2025a: 0.7 milliseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeHistcounts function.

Tall Arrays: Support for isuniform

The isuniform function has new tall array support. For more information, see Tall Arrays for Out-of-Memory Data.

Distributed Arrays: Reduce memory usage with single-precision sparse distributed arrays

You can now create and use single-precision sparse distributed and codistributed arrays. Sparse matrices provide efficient storage of data that has a large percentage of zeros and reduce computation time by eliminating operations on zero elements. Single-precision sparse distributed and codistributed arrays allow you to reduce memory usage. You can create single-precision sparse distributed and codistributed arrays using these functions:

By default, these functions create a double-precision sparse distributed or codistributed array. To create single-precision sparse distributed or codistributed arrays, specify the new typename argument as "single". For example, this code creates a 50-by-100 single-precision sparse distributed array with density 0.1.

distributed.sprandn(50,100,0.1,"single")

You can also create single-precision sparse distributed or codistributed arrays by providing single-precision distributed or codistributed arrays to the sparse function.

Distributed Arrays: Use functions with new and enhanced distributed array support

These functions have new and enhanced distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

Parallel Workflows: New and Updated Examples and Topics

Use these new examples to progress with parallel computing:

Support for Intel MPI

Parallel computing products now ship with Intel MPI Library for use on Linux for third-party schedulers.

To use Intel MPI, set the MPIImplementation additional property for your third-party cluster profile or object to "IntelMPI". To learn more about setting additional properties, see Set Additional Properties (MATLAB Parallel Server).

Support for MPICH: Upgrade to MPICH 4.2.2

Parallel computing products now ship with MPICH version 4.2.2 for use on Linux for third-party schedulers.

 Functionality being removed or changed

MPICH2 is removed

Errors

Starting in R2025a, parallel computing products no longer ship with MPICH2.

Support for Volta GPUs will be removed

Warns

Support for Volta architecture GPUs with compute capability 7.0 will be removed in a future release. At that time, GPU computing in MATLAB will require a GPU device with compute capability 7.5 or greater.

In R2025a, Volta architecture GPUs are still supported. MATLAB issues a warning the first time you use a Volta GPU.

For more information about supported GPU devices, see GPU Computing Requirements.

R2024b

New Features, Bug Fixes, Compatibility Considerations

 parfor-Loops: Use colon-vector indexing expressions with sliced variables

You can now use colon-vector expressions to index sliced input and output variables in parfor-loop statements. Colon-vector indexing expressions allow you to use more natural forms of indexing for parfor output variables, eliminating the need for nested for-loops. The colon-vector indexing expression must be in the form j:k or j:k:l.

For example, to assign values to columns 3 to 7 of the output variable out, use the colon-vector 3:7 as a subscript when you index the sliced variable.

out = zeros(10);
parfor i = 1:10
    out(i,3:7) = rand(1,5);
end
You can use either simple broadcast variables or scalar integer constants in the colon-vector indexing expressions. Temporary variables or complicated expressions are not supported.

 Compatibility Considerations

Starting in R2024b, indexing a sliced variable with a colon-vector expression no longer throws an error. For more information, see parfor no longer errors when you index sliced variables with colon-vector expressions.

Thread-Based Environment: Use Image Processing Toolbox functionality on thread workers

These Image Processing Toolbox functions can now run in a thread-based environment:

For more information, see Run MATLAB Functions in Thread-Based Environment.

GPU Validation: Verify GPU device setup

Use the validateGPU function to verify that MATLAB can use your GPU and to diagnose issues.

For example, you can validate the currently selected GPU. If no GPU device is selected, then the function validates the default device.

validateGPU
# Beginning GPU validation
# Performing system validation
#    CUDA-supported platform .................................................PASSED
#    CUDA-enabled graphics driver exists .....................................PASSED
#        Version: 537.70
#    CUDA-enabled graphics driver load .......................................PASSED
#    CUDA environment variables ..............................................PASSED
#    CUDA device count .......................................................PASSED
#        Found 2 devices.
#    GPU libraries load ......................................................PASSED
# 
# Performing device validation for device index 1
#    Device exists ...........................................................PASSED
#        NVIDIA RTX A5000
#    Device supported ........................................................PASSED
#    Device available ........................................................PASSED
#        Device is in 'Default' compute mode.
#    Device selectable .......................................................PASSED
#    Device memory allocation ................................................PASSED
#    Device kernel launch ....................................................PASSED
# 
# Finished GPU validation with no failures.

GPU Device: New device properties and updated display

Use the gpuDevice function to inspect these new properties of your GPU device:

  • LastAccessed — the date and time the device was last accessed by the current MATLAB session.

  • SingleDoubleRatio — the ratio of single- to double-precision floating point units (FPUs) on the device.

Creating or querying a GPUDevice object now displays only the Name, Index, ComputeCapability, DriverModel, TotalMemory, AvailableMemory, DeviceAvailable, and DeviceSelected properties. To view all of the properties of a device, create or query a GPUDevice object without suppressing output and click the Show all properties link.

D = gpuDevice
D = 

  CUDADevice with properties:

                 Name: 'NVIDIA RTX A5000'
                Index: 1 (of 2)
    ComputeCapability: '8.6'
          DriverModel: 'TCC'
          TotalMemory: 25544294400 (25.54 GB)
      AvailableMemory: 25120866304 (25.12 GB)
      DeviceAvailable: true
       DeviceSelected: true

  Show all properties.

The SupportsDouble, GPUOverlapsTransfers, and CanMapHostMemory properties are no longer displayed but you can still query these properties using dot notation. There are no plans to remove these properties.

GPU Functionality: Use functions with new and enhanced gpuArray support

MATLAB

These MATLAB functions have new gpuArray support:

These MATLAB functions have enhanced gpuArray support:

For more information, see Run MATLAB Functions on a GPU.

Statistics and Machine Learning Toolbox

These Statistics and Machine Learning Toolbox functions have new and enhanced gpuArray support:

For a full list of Statistics and Machine Learning Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Statistics and Machine Learning Toolbox).

Signal Processing Toolbox

These Signal Processing Toolbox functions and objects have new and enhanced gpuArray support:

For a full list of Signal Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Signal Processing Toolbox).

Wavelet Toolbox

This Wavelet Toolbox function has new gpuArray support:

For a full list of Wavelet Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Wavelet Toolbox).

Communications Toolbox

These Communications Toolbox functions have new and enhanced gpuArray support:

For a full list of Communications Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Communications Toolbox).

5G Toolbox

These 5G Toolbox functions have new and enhanced gpuArray support:

For a full list of 5G Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (5G Toolbox).

Image Processing Toolbox

This Image Processing Toolbox function has new gpuArray support:

For a full list of Image Processing Toolbox functions that accept GPU arrays, see Function List (GPU Arrays) (Image Processing Toolbox).

GPU Sparse Solvers: Use triangular preconditioner matrices to solve sparse linear systems

You can now provide a lower triangular matrix and an upper triangular matrix as preconditioner matrices when you solve sparse linear systems on a GPU using these functions:

Using lower triangular and upper triangular preconditioner matrices can significantly improve convergence of the algorithm.

% Create sparse coefficient matrix. 
gridDimension = 1024;
A = delsq(numgrid("A",gridDimension));

% Create right-hand side vector.
b = zeros(size(A,1),1);
b(1000) = 1;
    
% Use incomplete LU factorization to create lower and upper triangular
% preconditioner matrices.
options.type = "ilutp";
options.droptol = 1e-6;
options.thresh = 0;
[L,U] = ilu(A,options);

% Solve system of linear equations with preconditioner matrices.
A = gpuArray(A);
tol = 1e-12;
maxit = 20;
x = gmres(A,b,[],tol,maxit,L,U);

GPU arrayfun: Support for P-code files and functions defined in class definition files

You can now use P-code files with arrayfun. You can:

  • Use arrayfun to evaluate a function obfuscated as a P-code file.

  • Call arrayfun inside a P-code file.

  • Use arrayfun when the function it applies contains a call to a function obfuscated as a P-code file.

For more information about P-code files, see Create a Content-Obscured File with P-Code.

You can also call arrayfun in a class method to evaluate functions defined in the class definition file (a file with a .m extension that contains the classdef keyword).

For example, this class contains a method, output, that uses arrayfun to evaluate a local function, localFun.

classdef TestClass
    methods
        function output = func(obj,x)
            output = arrayfun(@localFun,x);
        end
    end
end

function output = localFun(x)
output = x.*x;
end

For more information about defining classes in MATLAB, see Creating a Simple Class.

GPU mldivide: Improved performance for overdetermined systems

The mldivide function (\) shows improved performance when solving linear systems on a GPU when input matrix A has more rows than columns (the system is overdetermined) and has at least 5000 elements. For example, in the following code, solving a sparse linear system with an 80-by-81 input matrix is about 1.8x faster than in the previous release:

function timeMldivide

% Create rectangular matrix and column vector.
n = 80;
A = rand(n+1,n,"gpuArray");
b = rand(n+1,1,"gpuArray");

% Time mldivide.
gputimeit(@() A\b)

end

The approximate execution times are:

R2024a: 4.6 milliseconds

R2024b: 2.5 milliseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeMldivide function.

GPU Sparse mldivide: Improved performance for triangular matrices

The mldivide function (\) shows improved performance and reduced memory usage on a GPU when solving sparse linear systems with a triangular input matrix A. For example, in the following code, solving a sparse linear system with a 271201-by-271201 input matrix is about 1800x faster than in the previous release:

function timeSparseMldivide

% Create sparse triangular matrix.
n = 300; 
A = gallery("wathen",n,n);
A = tril(A);
A = gpuArray(A);

% Create dense vector.
b = ones(size(A,1),1,"gpuArray");

% Time mldivide.
gputimeit(@() A\b)

end

The approximate execution times are:

R2024a: 36.3 s

R2024b: 0.02 s

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeSparseMldivide function.

GPU Cholesky Factorization: Improved performance

The chol function shows improved performance on a GPU when factorizing matrices that are at least 600-by-600. For example, factorizing the matrix in the following code is about 1.4x faster than in the previous release:

function timeChol

% Create symmetric positive definite matrix.
n = 1000;
A = rand(n,"gpuArray");
SPD = A'*A;

% Time Cholesky factorization.
gputimeit(@() chol(SPD))

end

The approximate execution times are:

R2024a: 59 milliseconds

R2024b: 41 milliseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeChol function.

GPU LU Factorization: Improved performance of two-output syntax

The lu function shows improved performance on a GPU when returning two outputs and factorizing matrices that are at least 140-by-140. For example, factorizing the matrix in the following code is about 2.4x faster than in the previous release:

function timeLU

% Create symmetric positive definite matrix.
n = 300;
A = rand(n,"gpuArray");
SPD = A'*A;

% Time LU factorization.
gputimeit(@() lu(SPD),2)

end

The approximate execution times are:

R2024a: 31 milliseconds

R2024b: 13 milliseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeLU function.

GPU arrayfun: Improved performance when parent workspace variable changes size

The arrayfun function shows improved performance on a GPU when the function applied by arrayfun uses a variable in the workspace of a parent function and that variable changes size. If the number of dimensions of the variable in the workspace of the parent function changes, then there is no performance improvement. For example, calling arrayfun several times in the following code is about 11.2x faster than in the previous release:

function timeArrayfun

% Reset GPU to clear previously compiled functions.
gpuDevice([]);
gpu = gpuDevice;
wait(gpu)

% Create array of sizes.
size = 10:1:1000;

% Time arrayfun call while changing the size of array upLevel.
tic
for idx = 1:numel(size)
    upLevel = rand(size(idx),"gpuArray");
    randomIndex = randi(numel(upLevel),10,"gpuArray");

    out = arrayfun(@accessUplevel,randomIndex);
end
toc

    function out = accessUplevel(randomIndex)
        % Define function called by arrayfun that accesses variable in
        % parent workspace.
        out = upLevel(randomIndex);
    end

end

The approximate execution times are:

R2024a: 4.71 s

R2024b: 0.42 s

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeArrayfun function.

Job Properties: Determine the storage size (in bytes) of jobs

You can now determine the number of bytes the data for your job occupies in the job storage location. The data includes task input and output arguments, attached files, and entries in the job's FileStore and ValueStore objects.

To access the storage size (in bytes) of your job, query the StorageBytes property of the parallel.Job object.

job = batch(@rand,1,{5}); 
job.StorageBytes 
ans = 

9425

You can also view the storage size of your job in the Job Monitor.

The Job Manager displays the jobs that use the Processes (default) profile, with a box around the Storage Bytes column.

Cluster Administration: Export cluster monitoring metrics for integration with Prometheus and Grafana

Set up your MATLAB Job Scheduler to export cluster monitoring metrics such as cluster status, worker utilization, and licenses in use. Use these metrics to monitor the health of the MATLAB Job Scheduler cluster, diagnose issues, and optimize performance.

You can gather the exported metrics with a cluster monitoring system such as Prometheus® and visualize them in a preconfigured Grafana® dashboard. This setup allows for live cluster monitoring and alerts. For more information, see Configure Metrics for MATLAB Job Scheduler (MATLAB Parallel Server).

Gather In-Memory Arrays: Improved performance

The gather function shows improved performance when the input is not a gpuArray, distributed array, or tall array. For example, gathering the data in the following code is about 9x faster than in the previous release:

function timeGather

% Create in-memory data.
x = 1;

% Time gather operations.
n = 1e7;
tic
for idx = 1:n
    y = gather(x);
end
toc/n

end

The approximate execution times are:

R2024a: 0.51 microseconds

R2024b: 0.06 microseconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system by calling the timeGather function.

Distributed Arrays: Use functions with new distributed array support

These functions have new and enhanced distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

Tall Arrays: Use functions with new and enhanced tall array support

These functions have new and enhanced tall array support:

For more information, see Tall Arrays for Out-of-Memory Data.

Signal Processing: New examples

These new examples show how to accelerate signal feature extraction and classification:

Support for MPICH: Upgrade to MPICH 4.1.2

Parallel computing products now ship with MPICH 4.1.2 for use on Linux for third-party schedulers.

 Functionality being removed or changed

parfor no longer errors when you index sliced variables with colon-vector expressions

Behavior change

In R2024b, you can index sliced variables in parfor-loop statements with colon-vector indexing expressions. Before R2024b, indexing sliced variables with colon-vector expressions results in an error. If you have code that relies on this behavior, update your code to avoid compatibility issues.

Support for Pascal GPUs will be removed

Warns

Support for Pascal architecture GPUs with compute capability 6.0 to 6.2 will be removed in a future release. At that time, GPU computing in MATLAB will require a GPU device with compute capability 7.0 or greater.

In R2024b, Pascal architecture GPUs are still supported. MATLAB issues a warning the first time you use a Pascal GPU.

For more information about supported GPU devices, see GPU Computing Requirements.

R2024a

New Features, Bug Fixes, Compatibility Considerations

Improved Scalability: Support for parallel pools with up to 2000 workers

Parallel Computing Toolbox now supports parallel pools with up to 2000 workers. To learn about parallel pools, see Run Code on Parallel Pools.

 Thread-Based Parallel Pool: Specify maximum number of thread workers in parfor-loop

You can now specify the maximum number of workers when executing parfor-loops on a thread-based parallel pool.

For example, this code starts a pool with 6 thread workers and runs the body of a parfor-loop on a maximum of 2 thread workers.

parpool("Threads",6);
parfor (i=1:6,2)
disp(i)
end

 Compatibility Considerations

Starting in R2024a, specifying the maximum number of thread workers in a parfor-loop no longer throws an error. For more information, see parfor no longer errors when you specify maximum number of thread workers in parfor-loop.

Thread-Based Environment: Use new functionality on thread workers

These MATLAB functions now have thread-based support:

For more information, see Run MATLAB Functions in Thread-Based Environment.

GPU Functionality: Use functions with new and enhanced gpuArray support

MATLAB

These MATLAB functions have new and enhanced gpuArray support:

  • bitget — You can now use this function with unsigned integers and 64-bit integers.

  • bitset — You can now use this function with unsigned integers and 64-bit integers.

  • createArray

  • isuniform

  • paddata

  • resize

  • sound — This function now accepts GPU arrays but does not run on a GPU.

  • soundsc — This function now accepts GPU arrays but does not run on a GPU.

  • subsref — You can now reference entire rows or columns of sparse GPU arrays by index.

  • trimdata

For more information, see Run MATLAB Functions on a GPU.

Statistics and Machine Learning Toolbox

These Statistics and Machine Learning Toolbox functions have new gpuArray support:

  • The fitrlinear (Statistics and Machine Learning Toolbox) and fitclinear (Statistics and Machine Learning Toolbox) functions accept GPU array input arguments and can execute on a GPU.

  • Most object functions of the models RegressionLinear (Statistics and Machine Learning Toolbox), RegressionPartitionedLinear (Statistics and Machine Learning Toolbox), ClassificationLinear (Statistics and Machine Learning Toolbox), and ClassificationPartitionedLinear (Statistics and Machine Learning Toolbox) now support GPU array input arguments and can execute on a GPU. The following functions do not support GPU array input arguments:

    • RegressionLinear — incrementalLearner, lime, shapley, and update

    • ClassificationLinear — incrementalLearner, lime, shapley, and update

  • geomean (Statistics and Machine Learning Toolbox), harmmean (Statistics and Machine Learning Toolbox), trimmean (Statistics and Machine Learning Toolbox) — You can now use the all and vecdim input arguments.

For a full list of Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

Communications Toolbox

These Communications Toolbox system objects and functions have new gpuArray support:

For a full list of Communications Toolbox functions with GPU functionality, see Functions with gpuArray support (Communications Toolbox).

5G Toolbox

These 5G Toolbox functions have new gpuArray support:

For a full list of 5G Toolbox functions with GPU functionality, see Functions with gpuArray support (5G Toolbox).

Audio Toolbox

This Audio Toolbox™ object has new gpuArray support:

  • audioDatastore (Audio Toolbox) — You can now return data on the GPU by setting the OutputEnvironment property to "gpu".

For a full list of Audio Toolbox functions with GPU functionality, see Functions with gpuArray support (Audio Toolbox).

Wavelet Toolbox

This Wavelet Toolbox function has new gpuArray support:

  • wsst (Wavelet Toolbox)

For a full list of Wavelet Toolbox functions with GPU functionality, see Functions with gpuArray support (Wavelet Toolbox).

GPU arrayfun: Support for cell array case expressions in switch, case, otherwise

Use a cell array as the case expression to compare the switch expression against multiple values with switch, case, otherwise within the function you apply using arrayfun. For example, you can use case {x1,y1} to execute the corresponding code if the switch expression matches at least one of x1 and y1.

GPU Device: Identify GPUs using their UUIDs

Inspect the universally unique identifier (UUID) of your GPU using the UUID property of a GPUDevice object. You can use the UUID to distinguish otherwise identical GPUs.

For example, you can inspect the UUID of four GPUs using the gpuDeviceTable function.

gpuDeviceTable(["Index","Name","UUID"])
    Index           Name                              UUID                   
    _____    __________________    __________________________________________

      1      "NVIDIA RTX A5000"    "GPU-957b509e-322a-ae88-59c8-b7435d0f98f4"
      2      "NVIDIA RTX A5000"    "GPU-6f3ad2c0-5ea1-b1a2-1dca-cd756d10dbc0"
      3      "NVIDIA RTX A5000"    "GPU-41c24f34-c915-919b-0bb3-20d07117e0ec"
      4      "NVIDIA RTX A5000"    "GPU-23ab01ce-2f42-2d0c-0f6b-db7ac8c10867"

Alternatively, you can select a GPU using the gpuDevice function and query its UUID.

D = gpuDevice;
D.UUID
    'GPU-957b509e-28ca-ae88-59c8-b7435d0f98f4'

Cluster Administration: Access MATLAB Job Scheduler job history

You can now access job history information about MATLAB Parallel Server use on your MATLAB Job Scheduler. This information allows you to audit cluster usage and gain insight into cluster usage patterns based on job type, MATLAB version, and task duration.

For more information, see Manage and Access MATLAB Job Scheduler Cluster Job History (MATLAB Parallel Server).

MATLAB Parallel Server in Spark: Enhanced support for Spark clusters

This release introduces enhanced support for Spark™ based clusters integrated with MATLAB Parallel Server.

You can now create and validate cluster profiles for Spark based clusters integrated with MATLAB Parallel Server. Cluster profiles allow you to define properties for the Spark cluster object using the Cluster Profile Manager or from the command line. You can also export and share the cluster profile with other MATLAB users to connect to the cluster. For more details about creating Spark cluster profiles, see Configure for Spark Clusters (MATLAB Parallel Server).

parallel.cluster.Spark cluster objects now have support for several Common Job Scheduler properties and object functions. For more details, see parallel.cluster.Spark.

Distributed Arrays: Use functions with new distributed array support

These functions have new and enhanced distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

Workflow Examples: New and updated examples and topics

This new example shows how to use datastores to transform large raw data into a state ready for future analysis and save to Parquet files using parallel workers:

These new examples show how to accelerate your code using a GPU:

These new topics and examples show different approaches for scaling up parallel code to accommodate large-scale computations on HPC clusters:

This updated example shows how to solve a simple optimization problem using parfeval:

 GPU Support: Upgrade to CUDA 12.2

GPU computing in MATLAB now uses CUDA 12.2. You can use new CUDA features when you write custom CUDA code for running in MATLAB via mexcuda or parallel.gpu.CUDAKernel.

 Compatibility Considerations

Parallel Batch Jobs: Disable SPMD support for parallel pools in batch jobs

You can now configure the parallel pools in batch jobs offloaded to local or MATLAB Job Scheduler clusters to run without single-program multiple-data (SPMD) support. This settings allows parallel pools to keep running even if workers abort during parfor execution.

To disable SPMD support, set the SpmdEnabled name-value argument to false in the call to the batch or createCommunicatingJob function. For example, this code offloads a job that runs the myFunction function on a parallel pool without SPMD support.

j = batch(@myFunction,1,{x,y},Pool=4,SpmdEnabled=false)

Third-Party Cluster Schedulers: Built-in integrations now align with plugin scripts

When you integrate MATLAB with your third-party scheduler using a built-in cluster profile, MATLAB now derives the built-in cluster profile from the generic plugin scripts published on GitHub. This change allows you to access the range of features available with the plugin scripts and customize how MATLAB interfaces with your scheduler setup.

Third-Party Cluster Schedulers: Use integrations for AWS Batch, Grid Engine, and HTCondor schedulers

The Add Cluster Profile button in the Cluster Profile Manager now includes options to create cluster profiles for AWS Batch, Grid Engine, and HTCondor scheduler clusters. When you select one of these options, the software creates a scheduler specific default profile which you can modify as needed.

Sparse Matrix Products: Improved performance on GPU

Computing the product of some sparse matrices on the GPU shows improved performance. The amount of speedup depends on the sparsity pattern of the matrix. For example, in the following code, computing the product of two sparse matrices is about 6x faster than in the previous release:

function timeSparseMatrixProduct

% Create sparse matrix.
A = gpuArray(gallery("neumann",4000000));

% Time matrix product.
gputimeit(@() A*A)

end

The approximate execution times are:

R2023b: 0.06 s

R2024a: 0.01 s

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeSparseMatrixProduct function.

Preconditioned Sparse Solvers: Improved performance using preconditioned iterative solvers on GPU

Solving large sparse linear systems using preconditioned iterative solvers on the GPU shows improved performance. The amount of speedup depends on the sparsity pattern of the matrices. For example, in the following code, solving a sparse linear system with a 752001-by-752001 coefficient matrix using the bicg function is about 33x faster than in the previous release:

function timePreconditionedSolver

% Create large sparse matrix.
A = gpuArray(gallery("wathen",500,500));

% Create a random right-hand side matrix.
actualSolution = rand(size(A,1),1);
b = A*actualSolution;

% Create preconditioner matrix.
k = 3;
M = tril(triu(A,-k),k);

% Time solver.
maxNumberOfIterations = 100;
gputimeit(@() bicg(A,b,[],maxNumberOfIterations,M))

end

The approximate execution times are:

R2023b: 26.8 s

R2024a: 0.8 s

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timePreconditionedSolver function.

 Functionality being removed or changed

Support for Kepler architecture GPUs is removed

Errors

Starting in R2024a, Kepler architecture GPUs are no longer supported and the range of compute capabilities supported is 5.0 to 9.x.

For more information, see GPU Computing Requirements.

Change to supported CUDA Toolkit

Errors

Starting in R2024a, CUDA Toolkit version 11.8 is no longer supported. To create parallel.gpu.CUDAKernel objects using libraries that are not installed with MATLAB or if you have previously installed the CUDA Toolkit, install CUDA Toolkit version 12.2 instead.

For more information, see Install CUDA Toolkit (Optional).

parfor no longer errors when you specify maximum number of thread workers in parfor-loop

Behavior change

In R2024a, you can specify the maximum number of thread workers to execute your parfor-loop. Before R2024a, specifying the maximum number of thread workers to execute a parfor-loop results in an error. If you have code that relies on this behavior, update your code to avoid compatibility issues. If you do not specify the maximum number of workers for a parfor-loop, MATLAB uses as many workers as are available in your parallel pool, as in previous releases.

Output of empty distributed tables and timetables displays the variable names of the distributed table or timetable

Behavior change

Starting in R2024a, MATLAB displays the variable names of empty distributed tables and timetables. Before R2024a, MATLAB only displays the size of distributed tables and timetables.

For example, this code creates an empty distributed table. In R2023b, MATLAB displays only the table size. In R2024a, MATLAB displays the table size and the variable names in a table header.

dT = distributed(table(Size=[0,3], ...
    VariableType=["double","double","string"], ...
    VariableNames=["Temperature","WindSpeed","Station"]))

Output in R2023bOutput in R2024a
dT =
 0×3 empty distributed table
dT =
  0×3 empty distributed table
    Temperature    WindSpeed    Station
    ___________    _________    _______

Output of distributed cell array displays the contents of the distributed cell array

Behavior change

Before R2024a, when you create a distributed cell array, MATLAB only provides a summary of the distributed object. In R2024a, MATLAB displays the size and data type of the arrays contained in distributed cells.

For example, this code creates a distributed cell array that contains a double value and an array of 100 double values. In R2023b, MATLAB provides a summary of the distributed cell array. In R2024a, MATLAB displays the size and data type of the arrays contained in the distributed cell.

dD = distributed({3.14,[1:100]})

Output in R2023bOutput in R2024a
dD =

    distributed object of size [1×2] with underlying class: cell

dD =
  1×2 distributed cell array
    {[3.1400]}    {1×100 double}

Cholesky factorization function returns symmetric positive definite flag stored in host memory

Behavior change

Before R2024a, if you call the chol function on a gpuArray matrix and return two outputs, then the software returns the second output (flag) as a gpuArray. In R2024a, the software returns the flag output as a scalar of type double stored in host memory and not as a gpuArray.

For example, this code factorizes matrix A and returns the flag output indicating whether A is symmetric positive definite. In R2023b, the software returns flag as a gpuArray. In R2024a, the software returns flag as a scalar of type double stored in host memory.

A = gpuArray(gallery("lehmer",6));
[R,flag] = chol(A);
isgpuarray(flag)

Output in R2023bOutput in R2024a
ans =

    1

ans =

    0

R2023b

New Features, Compatibility Considerations

Thread-Based Parallel Pools: Use ValueStore and FileStore objects on thread workers

You can now use ValueStore and FileStore objects on Parallel Computing Toolbox ThreadPool workers. The software creates ValueStore and FileStore objects when you create a ThreadPool object on your local machine. To access the ValueStore and FileStore objects on ThreadPool workers, use the getCurrentValueStore and getCurrentFileStore functions, respectively.

Thread-Based Environment: Use new functionality on thread workers

These MATLAB functions now have thread-based support:

addCauseaddCorrectionaddedgeaddlisteneraddnodeadjacency
alphanumericBoundaryalphanumericsPatternappendappend (timeseries)append (unit testing)array2table
array2timetableasFewOfPatternasManyOfPatternbarycentricToCartesianbctreebfsearch
biconncompbvpgetbvpsetcaldayscalendarDurationcalmonths
calquarterscalweekscalyearscartesianToBarycentriccaseInsensitivePatterncategorical
cell2tablecentralitycharacterListPatternconncompconvertCharsToStringsconvertContainedStringsToChars
convertStringsToCharsconvertToconvertvarsconvexHullcopyfilecospi
datetimeddegetddesetdegreedelete (handle)delete
dfsearchdigitBoundarydigitsPatterndigraphdistancesdrawnow
durationedgeAttachmentsedgecountfaceNormalfeatureEdgesfileattrib
filemarkerfilepartsfindedgefindnodefindobjfindprop
findstrfreeBoundarygather (tall)getReportgraphhandle
incidenceindegreeinedgesinnerjoinintersectintersect (polyshape)
iscalendardurationiscategoricalisConnectedisdatetimeisdurationisInterior
isisomorphicisjavaisletterismultigraphisomorphismisspace
isstristableistimetableisvalidjsondecodejsonencode
KeyValueStorelaplacianlasterrlasterrorletterBoundarylettersPattern
lineBoundarylookAheadBoundarylookBehindBoundarymaskedPatternmaxflowmeta.class.fromName
meta.package.fromNameminspantreemovefilemustBeAnamedPatternnargchk
nearestnearestNeighbornearestNeighbor (alpha shape)newlinenotifynumedges
numnodesodegetoptionalPatternoutdegreeoutedgesouterjoin
parenAssignparenReferenceparsepathseppatternpdepe
pdevalpointLocationpolybufferpolyshapepossessivePatternprefdir
recycleregexpPatternreordernodesrmdirrmedgermnode
rowfunrows2varsshortestpathshortestpathtreesimplifysimplify (polyshape)
sinpistlreadstlwritestruct2tabletabletable (unit testing)
table2arraytable2celltable2structtable2timetabletempdirtextBoundary
throwthrowAsCallertimetimerangetimetabletimetable2table
unionunion (polyshape)unstackvarfunvartypever
verLessThanversionvertexAttachmentsvertexNormalvoronoiDiagramwebread
websavewebwritewhitespaceBoundarywhitespacePatternwildcardPatternwithtol

For more information, see Run MATLAB Functions in Thread-Based Environment.

 Parallel Pools: Efficient scheduling of parfor and parfeval computations on process-based parallel pools

MATLAB now efficiently schedules parfor and parfeval computations on process-based parallel pools as pool workers become available. You can now run parfeval and parfor computations on process-based pools on a local machine or a remote cluster concurrently.

For example, this code calls the parfevalWithparfor function, which runs a long-running parfeval computation in the background and then executes a parfor-loop. The code is about 1.5x faster than in the previous release because the software runs the parfeval and parfor computations simultaneously, removing 14 seconds of scheduling overhead.

function t = poolTimingTest

% Start parallel pool if none exists
pool = gcp("nocreate");

    function parfevalWithparfor
        future = parfeval(@pause,0,30);
        parfor ii = 1:10
            pause(ii);
        end
        wait(future);
    end

% Time executing parfor and parfeval concurrently
t = timeit(@() parfevalWithparfor);

end

The approximate execution times are:

R2023a: 44.03 seconds

R2023b: 30.03 seconds

 Compatibility Considerations

In previous releases, MATLAB schedules parfor or parfeval computations to run separately. If you have code that relies on this behavior, update your code to avoid compatibility issues.

Parallel Workflow Examples: New and updated examples and topics

This new topic contains information about accelerating code in MATLAB and using parallel computing capabilities to efficiently run code on multicore and multiprocessor computers:

This new topic helps you choose the right data management tools and workflows for your needs:

This new example shows how to send data to workers in a data queue:

This updated example shows to compare how fast functions run on the client and on a parallel pool:

Support for Apple silicon Macs

Parallel Computing Toolbox now supports Apple silicon Macs with this limitation:

  • Distributed and codistributed arrays are not supported for local process pools.

gpurng Function: Specify random number generator without specifying seed

You can now use the new syntax gpurng(generator) to specify the algorithm that the random number generator on the GPU uses. Use this syntax to set the random number algorithm without specifying the seed. The gpurng function uses a default seed of 0. This syntax is equivalent to gpurng(0,generator). For example, gpurng("philox") initializes the Philox 4x32 generator with a seed of 0. For more information, see gpurng.

GPU Support for switch, case, and otherwise inside arrayfun

You can now use switch conditional statements in functions you apply using arrayfun with gpuArray input. This functionality has these limitations:

  • Case expressions support only numeric and logical values.

  • Using a cell array as the case expression to compare the switch expression against multiple values, for example, case {x1,y1} is not supported.

GPU Functionality: Use functions with new and enhanced gpuArray support

These MATLAB functions have new and enhanced gpuArray support:

These Statistics and Machine Learning Toolbox functions have new gpuArray support:

  • pearscdf (Statistics and Machine Learning Toolbox)

  • pearspdf (Statistics and Machine Learning Toolbox)

  • pearsrnd (Statistics and Machine Learning Toolbox)

For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

These Signal Processing Toolbox functions have new gpuArray support:

For a list of all Signal Processing Toolbox functions with GPU functionality, see Functions with gpuArray support (Signal Processing Toolbox).

These Communications Toolbox functions have new gpuArray support:

For a list of all Communications Toolbox functions with GPU functionality, see Functions with gpuArray support (Communications Toolbox).

This Wavelet Toolbox function has new gpuArray support:

For a list of all Wavelet Toolbox functions with GPU functionality, see Functions with gpuArray support (Wavelet Toolbox).

GPU Arrays: Improved performance

Some workflows using gpuArray objects show improved performance. For example, simulating Conway's "Game of Life" on the GPU in this example is about 2x faster than in the previous release:

function timeGameOfLife

% Select GPU device.
gpu =  gpuDevice;
wait(gpu)

% Start timing.
tic

% Define simulation parameters.
gridSize = 1000;
numGenerations = 5000;
initialGrid = (rand(gridSize,gridSize) > .75);
grid = gpuArray(initialGrid);
p = [1 1:gridSize-1];
q = [2:gridSize gridSize];

% Loop over generations.
for generation = 1:numGenerations
    % Count number of neighbors.
    neighbours = grid(:,p) + grid(:,q) + grid(p,:) + grid(q,:) + ...
        grid(p,p) + grid(q,q) + grid(p,q) + grid(q,p);

    % Update the grid. A live cell with two live neighbors, or any cell with
    % three live neighbors, is alive at the next step.
    grid = (grid & (neighbours == 2)) | (neighbours == 3);
end

% Gather back to host memory.
gather(grid);
wait(gpu)

% Record the elapsed time.
t = toc;

end

The approximate execution times are:

R2023a: 2.8 s

R2023b: 1.2 s

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeGameOfLife function.

Improved Scalability: Use MATLAB Job Scheduler clusters with up to 10,000 workers

MATLAB Parallel Server with MATLAB Job Scheduler now supports clusters with up to 10,000 workers. Support for large parallel pools remains at 1024 workers.

When you scale above 1000 workers, you must increase the heap memory available to the job manager. For more information, see Customize Startup Parameters (MATLAB Parallel Server).

Cluster Scheduling: Specify load-balancing scheduling algorithm for MATLAB Job Scheduler

You can now select a scheduling algorithm for your MATLAB Job Scheduler that balances the workload more evenly across your cluster nodes. Specify the scheduling algorithm using the SCHEDULING_ALGORITHM parameter in the mjs_def file. For more details, see Define MATLAB Job Scheduler Startup Parameters (MATLAB Parallel Server).

 Big Data Workflows: Convert between tall arrays and distributed arrays

You can now convert a tall array to a distributed array to access MATLAB functions that have distributed array support. To convert tall arrays to distributed arrays, use the distributed function with a tall array.

You can also convert a distributed array to a tall array to access functions that have tall array support. To convert distributed arrays into tall arrays, use the tall function with a distributed array.

 Compatibility Considerations

Using a distributed array in the tall function or a tall array in the distributed function throws an error in earlier releases. If your code relies on the errors that earlier releases of MATLAB throw for those conversions, such as within a try/catch block, update your code so it does not rely on those errors.

Distributed Arrays: Faster distribution of local arrays to workers

Creating distributed arrays from large local arrays shows improved performance. For example this code calls the distributeLargeArray function which distributes a large array to the workers in a parallel pool. The code is about 4.8x faster than in the previous release.

function t = distributedTimingTest

% Start parallel pool if none exists
pool = gcp("nocreate");

% Prepare large array
W = triu(gallery("wathen", 1000, 1000));

f = @() distributed(W);

% Time distributing large array
t = timeit(f);

end

The approximate execution times are:

R2023a: 4.20 seconds

R2023b: 0.87 seconds

The code was timed on a Windows 10, Intel(R) Xeon(R) CPU E5-1650 v3 @ 3.50 GHz test system using the distributed function on a parallel pool with six workers.

Distributed Arrays: Use functions with new distributed array support

These functions have new distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

 Functionality being removed or changed

arrayfun with GPU Arrays: Passing arrays from parent workspace to nested function and indexing into array within nested function now errors

Behavior change

In this code, you create the parentWorkspaceVar variable in the parent workspace of the foo function. If you use foo in an arrayfun call with gpuArray input and if the foo function passes parentWorkspaceVar as an input to a nested function within foo, the code errors.

Instead of passing the parent workspace variable (parentWorkspaceVar) to the nested function (bar), use the parent workspace variable directly. This variable is already in the scope of the nested function.

ErrorsAlternative
function y = exampleFunction
    parentWorkspaceVar = 1:9;
    x = ones(2,"gpuArray");
    y = arrayfun(@foo,x);

    function y = foo(x)
        y = bar(parentWorkspaceVar);

        function y = bar(z) % Errors
            y = z(1);
        end

    end

end
function y = exampleFunction
    parentWorkspaceVar = 1:9;
    x = ones(2,"gpuArray");
    y = arrayfun(@foo,x);

    function y = foo(x)
        y = bar;

        function y = bar
            y = parentWorkspaceVar(1); % Use parent workspace variable directly.
        end

    end

end

arrayfun with GPU Arrays: Functions writing into variables created in parent function now error

Behavior change

In this code, you create the workspaceVar variable in the workspace of the bar function. If you use bar in an arrayfun call with gpuArray input and if a nested function foo writes into workspaceVar, the code errors.

Instead of writing to the variable (workspaceVar) within the nested function (foo), add another output to the nested function and use the output to write to the variable.

ErrorsWorkaround
function x = exampleFunction
    z = ones(2,"gpuArray");
    x = arrayfun(@bar,z);

    function y = bar(z)
        workspaceVar = 2;
        foo(z);

        function x = foo(z)
            workspaceVar = 10; % Errors
            x = z;
        end

        y = workspaceVar;
    end

end
function x = exampleFunction
    z = ones(2,"gpuArray");
    x = arrayfun(@bar,z);

    function y = bar(z)
        workspaceVar = 2;
        [out,workspaceVar] = foo(z); % Write to workspaceVar outside foo.

        function [x,y] = foo(z) % Add another output y to the nested function.
            y = 10; 
            x = z;
        end

        y = workspaceVar;
    end

end

parallel.pool.Constant with no arguments now returns invalid Constant object

Behavior change

When you call the parallel.pool.Constant function without input arguments, it initializes a Constant object in an invalid state. In previous releases, calling the parallel.pool.Constant function without input arguments errors.

You can use parallel.pool.Constant with no arguments to assign invalid Constant objects to array elements. When you create or grow an array of Constant objects without assigning values to each element, any new elements of the array contain invalid Constant elements.

R2023a

New Features, Compatibility Considerations

 PreferredPoolNumWorkers: Specify limited number of workers per cluster profile

The global preference for the default number of workers in a pool is replaced by a per-profile property. You can configure the new PreferredPoolNumWorkers cluster object property in the cluster profile manager or at the command line. The PreferredPoolNumWorkers property default is the number of workers available to the cluster (NumWorkers) for personal cloud cluster profiles or the minimum of NumWorkers and 32 for other nonlocal profiles.

Use this new property to create pools with a reasonable number of workers when you call parpool without requesting a specific number of workers. You can also use this property to control how many workers you start on clusters to avoid taking up too many resources.

For local profiles, the default parallel pool size is NumWorkers. For all profiles, you can continue to request a specific number of workers when you call parpool to start a parallel pool outside the main body of code. If you specify a pool size at the command line, you override your preferences. For more details, see Pool Size and Cluster Selection.

 Compatibility Considerations

The Preferred number of workers in a parallel pool option in Parallel Preferences has been removed. Use the PreferredPoolNumWorkers or NumWorkers cluster object properties instead. For more information, see Functionality being removed or changed.

Parallel Menu: Select a GPU using the Parallel menu

You can now select which GPU device to use for computation from the MATLAB desktop. On the Home tab, in the Environment area, select Parallel > Select GPU Environment. For each available GPU, the menu displays the total memory, the multiprocessor count, and an indication of when the device was last used. To inspect more properties of your GPU devices, use the gpuDevice function.

The Parallel menu, including the Select GPU Environment pane showing two GPU devices. A tick next to the first device indicates that it is the selected device.

GPUDevice Object: Changes to device properties

gpuDevice objects that the gpuDevice function returns now have these additional properties:

  • GraphicsDriverVersion — Graphics driver version currently in use.

  • DriverModel — Operating model of the graphics driver.

    On Windows operating systems, the model options are:

    • 'WDDM' — Display model.

    • 'TCC' — Compute model.

    On other operating systems, DriverModel is 'N/A'.

  • CachePolicy — Policy that determines how much GPU memory the GPU can cache to accelerate computation. You can set the CachePolicy property to "balanced", "minimum", or "maximum".

For example, select a GPU using the gpuDevice function to inspect the new properties.

D = gpuDevice
D = 

  CUDADevice with properties:

                      Name: 'NVIDIA RTX A5000'
                     Index: 1
         ComputeCapability: '8.6'
            SupportsDouble: 1
     GraphicsDriverVersion: '511.79'
               DriverModel: 'TCC'
            ToolkitVersion: 11.2000
        MaxThreadsPerBlock: 1024
          MaxShmemPerBlock: 49152 (49.15 KB)
        MaxThreadBlockSize: [1024 1024 64]
               MaxGridSize: [2.1475e+09 65535 65535]
                 SIMDWidth: 32
               TotalMemory: 25553076224 (25.55 GB)
           AvailableMemory: 25145376768 (25.15 GB)
               CachePolicy: 'balanced'
       MultiprocessorCount: 64
              ClockRateKHz: 1695000
               ComputeMode: 'Default'
      GPUOverlapsTransfers: 1
    KernelExecutionTimeout: 0
          CanMapHostMemory: 1
           DeviceSupported: 1
           DeviceAvailable: 1
            DeviceSelected: 1

Change the caching policy to allow the GPU to cache the maximum amount of memory for accelerating computation.

D.CachePolicy = "maximum";

The DriverVersion property no longer displays by default, but you can still query the property using dot notation. There are no plans to remove the DriverVersion property.

GPU Functionality: Use functions with new and enhanced gpuArray support

These MATLAB functions have new and enhanced gpuArray support:

For more information, see Run MATLAB Functions on a GPU.

These Statistics and Machine Learning Toolbox functions have new and enhanced gpuArray support:

  • regress (Statistics and Machine Learning Toolbox)

  • fitrsvm (Statistics and Machine Learning Toolbox) — You can now specify gpuArray inputs.

    • Most object functions of RegressionSVM (Statistics and Machine Learning Toolbox), CompactRegressionSVM (Statistics and Machine Learning Toolbox), and RegressionPartitionedSVM (Statistics and Machine Learning Toolbox) run on a GPU when you provide GPU array input arguments. These object functions do not support GPU array input arguments.

      Object FunctionModel Type
      incrementalLearner (Statistics and Machine Learning Toolbox)RegressionSVM (Statistics and Machine Learning Toolbox) and CompactRegressionSVM (Statistics and Machine Learning Toolbox)
      lime (Statistics and Machine Learning Toolbox)RegressionSVM (Statistics and Machine Learning Toolbox) and CompactRegressionSVM (Statistics and Machine Learning Toolbox)
      shapley (Statistics and Machine Learning Toolbox)RegressionSVM (Statistics and Machine Learning Toolbox) and CompactRegressionSVM (Statistics and Machine Learning Toolbox)
      update (Statistics and Machine Learning Toolbox)CompactRegressionSVM (Statistics and Machine Learning Toolbox)
    • You can now use the gather (Statistics and Machine Learning Toolbox) function to create a RegressionSVM (Statistics and Machine Learning Toolbox), CompactRegressionSVM (Statistics and Machine Learning Toolbox), or RegressionPartitionedSVM (Statistics and Machine Learning Toolbox) object from an equivalent object fitted with GPU array data. The properties of the new object are stored in host memory.

For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

These Signal Processing Toolbox functions have new and enhanced gpuArray support:

For a list of all Signal Processing Toolbox functions with GPU functionality, see Functions with gpuArray support (Signal Processing Toolbox).

These Wavelet Toolbox functions have new and enhanced gpuArray support:

For a list of all Wavelet Toolbox functions with GPU functionality, see Functions with gpuArray support (Wavelet Toolbox).

These Communications Toolbox functions have new gpuArray support:

For a list of all Communications Toolbox functions with GPU functionality, see Functions with gpuArray support (Communications Toolbox).

PTX Build Option for mexcuda: Compile PTX files using mexcuda

You can now compile parallel thread execution (PTX) files using the mexcuda function by specifying the -ptx build option. The function compiles the PTX file and gives it the same name as the input CUDA C++ (CU) source file.

You no longer need to install the CUDA Toolkit to compile a PTX file from a CU file.

For an updated overview of the workflow for creating and executing your own custom CUDAKernel object, see Run CUDA or PTX Code on GPU.

Support for New GPU Architectures: Update to NVIDIA CUDA 11.8

Parallel Computing Toolbox software now uses CUDA version 11.8, which supports NVIDIA GPUs with compute capability up to 9.x, including Hopper architecture GPUs. For more information, see GPU Computing Requirements.

Workflow Examples: New and updated examples and topics

These new and updated examples show how to accelerate your code using a GPU:

This updated example shows how to monitor the training of deep neural networks in a batch job using a ValueStore object:

This new topic helps you choose the right parallel language and workflow for your needs:

Thread-Based Environment: Use new functionality on thread workers

These MATLAB functions and classes now have thread-based support:

For more information, see Run MATLAB Functions in Thread-Based Environment.

Tall Arrays: Use functions with new and enhanced tall array support

These functions have new and enhanced tall array support:

For more information, see Tall Arrays.

Distributed and Tall tables and timetables: Perform calculations directly on distributed and tall tables and timetables without extracting their data

You can now perform calculations directly on distributed and tall tables and timetables without extracting their data. In previous releases, all calculations require you to extract data from your distributed and tall tables and timetables by indexing into them.

For example, to scale a distributed table that contains numeric data, multiply it by a scale factor.

table = array2table(rand(2),VariableNames={"V1","V2"});
distTable = distributed(table)
distTable =
  2×2 distributed table
      V1          V2   
    _______    ________

    0.69483     0.95022
     0.3171    0.034446
distTable = distTable .* 10
distTable =
  2×2 distributed table
      V1          V2   
    _______    ________

    6.9483     9.5022
     3.171    0.34446

These MATLAB functions have new support for performing calculations directly on distributed and tall tables and timetables.

For more information about performing calculations directly on tables and timetables, see Direct Calculations on Tables and Timetables.

 Composite arrays: Use gather to convert a Composite array into a local cell array.

You can now use the gather function to convert a Composite array containing data on parallel workers into a cell array in the local workspace.

For example, gather data after the execution of an spmd statement.

parpool("Threads",3)
spmd, c = ones(spmdIndex); end
data = gather(c)
data =

  1×3 cell array

    {[1]}    {2×2 double}    {3×3 double}

 Compatibility Considerations

In previous releases, when you use gather on a Composite array, MATLAB returns the same Composite array as the output. If you have code that relies on this behavior, update your code to avoid compatibility issues.

Distributed Arrays: Use functions with new and enhanced distributed array support

Discover Clusters Supports Third-Party Schedulers

The Discover Clusters dialog can now locate third-party scheduler clusters for different releases of parallel computing products. For information about cluster discovery, see Discover Clusters.

Third-Party Cluster Schedulers: Access simplified plugin scripts for integration on GitHub

When you integrate MATLAB with your scheduler using MATLAB Parallel Server, you can now use a single set of scripts and specify the values of one of these AdditionalProperties properties to describe your network configuration.

Network ConfigurationPrevious Submission ModeNew Properties to Set

The MATLAB client does not share the file system with the cluster nodes.

Shared

  • HasSharedFilesystem = true

The MATLAB client is unable to directly submit jobs to the third-party scheduler.

Remote

  • HasSharedFilesystem = true

  • ClusterHost = HeadnodeHostname

The MATLAB client does not share the file system with the cluster nodes and is unable to directly submit jobs to the third-party scheduler.

Nonshared

  • HasSharedFilesystem = false

  • ClusterHost = HeadnodeHostname

In previous releases, you have to choose a set of scripts for shared, nonshared, and remote submission modes

The simplified plugin scripts for these third-party schedulers are now available in these GitHub repositories:

You can still download the plugin scripts from MATLAB Central™ File Exchange.

Third-Party Cluster Schedulers: Open ports on workers to listen for connections from client

You can now open listening ports on MATLAB Parallel Server workers running on third-party scheduler clusters to enable your MATLAB client to connect to the workers. Use this functionality to run interactive parallel pools on third-party scheduler clusters without the need to modify your network systems such as firewalls. For more details, see pctconfig.

LDAP Support for MATLAB Job Scheduler Clusters: Authenticate cluster logins against LDAP credentials

Validate and control user access to a MATLAB Job Scheduler cluster by using a Lightweight Directory Access Protocol (LDAP) server to authenticate user logins. For more information, see Configure LDAP Server Authentication for MATLAB Job Scheduler (MATLAB Parallel Server).

Cluster Security: Restrict use of MATLAB Job Scheduler cluster commands

You can now restrict the use of cluster changing commands to specific users by verifying commands sent to MATLAB Job Scheduler clusters. In previous releases, any user can execute commands that can change the state of cluster. For more information, see Set Cluster Command Verification (MATLAB Parallel Server).

Parallel Computing in MATLAB Online Server Environment: Speed up MATLAB Online workflows

You can now integrate Parallel Computing Toolbox software and MATLAB Parallel Server with your MATLAB Online Server™ instance and use parallel computing to speed up MATLAB Online™ workflows.

MATLAB Parallel Server in Kubernetes: Configure and run MATLAB Parallel Server workers on Kubernetes clusters

You can now integrate MATLAB Parallel Server with your Kubernetes® cluster and interface Parallel Computing Toolbox software on your computer to the Kubernetes cluster. For more details, see the Parallel Computing Toolbox for MATLAB Parallel Server with Kubernetes plugin script on GitHub or MATLAB Central File Exchange.

 Functionality being removed or changed

Preferred number of workers in a parallel pool option has been removed

Behavior change

The Preferred number of workers in a parallel pool option in Parallel Preferences has been removed. This change gives you more control over how many workers to start on nonlocal clusters.

For MATLAB Job Scheduler, third party schedulers, and cloud clusters, you can use the PreferredPoolNumWorkers property of your cluster profiles to specify the preferred number of workers to start in a parallel pool. For the local Processes and Threads profiles, use the NumWorkers property to specify the default number of workers to start in a parallel pool.

rng("default") on MATLAB parallel workers now sets random number generator settings to worker default

Behavior change

When you call the rng function with the "default" argument on MATLAB parallel workers, MATLAB resets the random number generator settings to the worker default values. The default corresponds to the Threefry generator with 20 rounds and a seed of 0

In previous releases, when you run rng("default") on parallel workers, MATLAB changes the worker random number generator settings to the client default values. The default corresponds to the Mersenne Twister generator with a seed of 0.

parallel.pool.Constant objects no longer automatically transferred to workers

Behavior change

MATLAB no longer automatically transfers the parallel.pool.Constant object from your current MATLAB session to workers in a parallel pool. MATLAB sends the Constant object to workers only if the object is required to execute your code.

Thread-Based Parallel Pool: startup or matlabrc no longer runs when workers start

Behavior change

When you start a thread-based parallel pool, the thread-based workers no longer run the startup or matlabrc file. All startup options on the client are automatically mirrored to the thread-based workers.

pmode has been removed

Errors

The pmode function has been removed. To execute commands interactively on multiple workers, use spmd instead.

pload and psave have been removed

Errors

The pload and psave functions have been removed. To save and load data on workers, in the form of composite arrays or distributed arrays, use dsave and dload instead.

R2022b

New Features, Compatibility Considerations

Thread-Based Parallel Pool: Specify number of thread workers using parpool

You can now specify the pool size of a thread-based parallel pool by using parpool.

For example, start a pool of four thread workers.

parpool("Threads",4);

You can also set Threads as the default parallel environment and start a parallel pool with the largest possible number of workers.

parallel.defaultProfile("Threads");
parpool([1 50]);

 Parallel menu: Select Threads using the Parallel menu

You can now select the Threads parallel environment option from the MATLAB desktop Parallel menu.

Local machine Processes and Threads environments available to select.

  • You can set the default parallel environment on your local machine from the MATLAB desktop Home tab, in the Environment area, by selecting one of these options:

    • Parallel > Select Parallel Environment > Processes

    • Parallel > Select Parallel Environment > Threads

You can also change the default profile in Parallel Preferences. For more information, see Specify Your Parallel Preferences.

 Compatibility Considerations

The local profile is no longer recommended. For more information, see Functionality being removed or changed.

gpuDevice Object: Display values for byte-based memory properties with appropriate units

The gpuDevice object now displays the TotalMemory, AvailableMemory, and MaxShmemPerBlock properties with appropriately scaled units. The object displays these properties as bytes (B), kilobytes (KB), megabytes (MB), or gigabytes (GB). Accessing any of these properties using dot notation still returns a value in bytes.

For example, create a gpuDevice object and access its TotalMemory property.

g = gpuDevice
g = 

  CUDADevice with properties:

                      Name: 'NVIDIA RTX A5000'
                     Index: 1
         ComputeCapability: '8.6'
            SupportsDouble: 1
             DriverVersion: 11.6000
            ToolkitVersion: 11.2000
        MaxThreadsPerBlock: 1024
          MaxShmemPerBlock: 49152 (49.15 KB)
        MaxThreadBlockSize: [1024 1024 64]
               MaxGridSize: [2.1475e+09 65535 65535]
                 SIMDWidth: 32
               TotalMemory: 25553076224 (25.55 GB)
           AvailableMemory: 25153765376 (25.15 GB)
       MultiprocessorCount: 64
              ClockRateKHz: 1695000
               ComputeMode: 'Default'
      GPUOverlapsTransfers: 1
    KernelExecutionTimeout: 0
          CanMapHostMemory: 1
           DeviceSupported: 1
           DeviceAvailable: 1
            DeviceSelected: 1
g.TotalMemory
ans =

   2.5553e+10

Cloud GPU Computing: Access Cloud GPUs and Clusters in AWS Using Cloud Center

If you do not have a GPU available, you can speed up your MATLAB code with one or more high-performance NVIDIA GPUs in the cloud. Working in the cloud requires some initial setup, but using cloud resources can significantly accelerate your code without requiring you to buy and set up your own local GPUs. The simplest way to access a high-performance cloud GPU is to use Cloud Center.

From April 2022, you can use Cloud Center to start MATLAB in Amazon Web Services (AWS). You can start a single machine, with MATLAB installed, that you can access from a web browser or a remote desktop application. To get started, see Get Started with Cloud Center and Start MATLAB on Amazon Web Services (AWS) Using Cloud Center. To learn more about other cloud options for GPU access, see Run MATLAB using GPUs in the Cloud.

You can continue to use Cloud Center to start and manage MATLAB Parallel Server clusters that you can access from any MATLAB client. MATLAB Parallel Server clusters provide more hardware, including GPU and CPU clusters.

New Examples: Use ValueStore to monitor batch jobs

These new examples show how to retrieve data and monitor batch jobs with a ValueStore object:

Parallel workflow examples: Use clusters for 5G Communications simulations

This new Communications Toolbox example shows how to accelerate simulations using clusters for 5G LDPC Block Error Rate:

Data Analysis: Use mapreduce on Spark clusters

Parallel Computing Toolbox and MATLAB Parallel Server support the use of Spark clusters for the execution environment of mapreduce applications. For more information, see:

Parallel Computing page: Learn about MathWorks solutions that help you take advantage of more hardware resources

Use the new Parallel Computing page to discover solutions that use parallel, GPU, and cloud computing.

Multistep Authentication for Third-Party Schedulers: Connect to remote clients with multiple authentication options

You can now use a RemoteClusterAccess object to perform multistep authentication with any number of authentication options. To connect to a cluster with multiple authentication requirements, specify the AuthenticationMode property of the RemoteClusterAccess object as a string array or cell array containing a combination of "Agent", "IdentityFile", "Multifactor", and "Password".

For example, create a RemoteClusterAccess object to connect to a cluster that requires a password and identity file.

RemoteClusterAccess(AuthenticationMode={"IdentityFile","Password"});

The sample plugin scripts for third-party schedulers now accept the cell array {"IdentityFile","Password"} as the value of the AuthenticationMode property of an AdditionalProperties object, which sets the same AuthenticationMode property value in the RemoteClusterAccess object.

Thread-Based Environment: Use new functionality on thread workers

These Parallel Computing Toolbox functions now have new thread-based support:

These Image Processing Toolbox functions now have thread-based support.

adapthisteq (Image Processing Toolbox)bfscore (Image Processing Toolbox)bwconncomp (Image Processing Toolbox)bwdist (Image Processing Toolbox)
bwferet (Image Processing Toolbox)bwlabel (Image Processing Toolbox)bwlabeln (Image Processing Toolbox)bwlookup (Image Processing Toolbox)
bwmorph3 (Image Processing Toolbox)bwpack (Image Processing Toolbox)bwskel (Image Processing Toolbox)bwtraceboundary (Image Processing Toolbox)
bwunpack (Image Processing Toolbox)demosaic (Image Processing Toolbox)edge (Image Processing Toolbox)edge3 (Image Processing Toolbox)
entropyfilt (Image Processing Toolbox)exrHalfAsSingle (Image Processing Toolbox)exrinfo (Image Processing Toolbox)exrread (Image Processing Toolbox)
exrwrite (Image Processing Toolbox)fibermetric (Image Processing Toolbox)grabcut (Image Processing Toolbox)histeq (Image Processing Toolbox)
hough (Image Processing Toolbox)imabsdiff (Image Processing Toolbox)imapplymatrix (Image Processing Toolbox)imbothat (Image Processing Toolbox)
imboxfilt (Image Processing Toolbox)imboxfilt3 (Image Processing Toolbox)imclose (Image Processing Toolbox)imdiffusefilt (Image Processing Toolbox)
imdilate (Image Processing Toolbox)imerode (Image Processing Toolbox)imfill (Image Processing Toolbox)imfilter (Image Processing Toolbox)
imgaussfilt (Image Processing Toolbox)imgaussfilt3 (Image Processing Toolbox)imgradient (Image Processing Toolbox)imgradient3 (Image Processing Toolbox)
imhist (Image Processing Toolbox)imlincomb (Image Processing Toolbox)imnlmfilt (Image Processing Toolbox)imopen (Image Processing Toolbox)
imreconstruct (Image Processing Toolbox)imregionalmax (Image Processing Toolbox)imregionalmin (Image Processing Toolbox)imsegfmm (Image Processing Toolbox)
imtophat (Image Processing Toolbox)inpaintCoherent (Image Processing Toolbox)inpaintExemplar (Image Processing Toolbox)integralImage (Image Processing Toolbox)
integralImage3 (Image Processing Toolbox)intlut (Image Processing Toolbox)iradon (Image Processing Toolbox)isexr (Image Processing Toolbox)
label2idx (Image Processing Toolbox)medfilt2 (Image Processing Toolbox)medfilt3 (Image Processing Toolbox)modefilt (Image Processing Toolbox)
multissim (Image Processing Toolbox)multissim3 (Image Processing Toolbox)nitfread (Image Processing Toolbox)obliqueslice (Image Processing Toolbox)
ordfilt2 (Image Processing Toolbox)planar2raw (Image Processing Toolbox)poly2mask (Image Processing Toolbox)qtdecomp (Image Processing Toolbox)
radon (Image Processing Toolbox)raw2planar (Image Processing Toolbox)raw2rgb (Image Processing Toolbox)rawread (Image Processing Toolbox)
regionprops (Image Processing Toolbox)regionprops3 (Image Processing Toolbox)rgb2lab (Image Processing Toolbox)rgb2lightness (Image Processing Toolbox)
superpixels (Image Processing Toolbox)superpixels3 (Image Processing Toolbox)watershed (Image Processing Toolbox) 

For more information, see Run MATLAB Functions in Thread-Based Environment.

GPU Functionality: Use functions with new and enhanced gpuArray support

These MATLAB functions have new and enhanced gpuArray support:

For more information, see Run MATLAB Functions on a GPU.

These Statistics and Machine Learning Toolbox functions have enhanced gpuArray support:

  • fitcecoc (Statistics and Machine Learning Toolbox) - Support for categorical predictors for classification tree learners

  • fitcensemble (Statistics and Machine Learning Toolbox) - Support for categorical predictors for classification tree learners

  • fitcsvm (Statistics and Machine Learning Toolbox)

    • Support for hyperparameter optimization

    • When fitcsvm fits the SVM model on a GPU, some of the properties of the ClassificationSVM (Statistics and Machine Learning Toolbox) model are stored in GPU memory

    • Use gather to create a ClassificationSVM or CompactClassificationSVM object with properties stored in the local workspace from an equivalent object with properties stored in GPU memory

  • fitctree (Statistics and Machine Learning Toolbox) - Support for categorical predictors

  • fitensemble (Statistics and Machine Learning Toolbox) - Support for categorical predictors

  • fitrensemble (Statistics and Machine Learning Toolbox) - Support for categorical predictors

  • fitrtree (Statistics and Machine Learning Toolbox) - Support for categorical predictors

  • pca (Statistics and Machine Learning Toolbox) - Improved performance when you use the SVD algorithm

  • Object functions of the CompactClassificationSVM (Statistics and Machine Learning Toolbox) model execute on a GPU if the model is fitted with GPU arrays

  • Object functions of the ClassificationSVM (Statistics and Machine Learning Toolbox) model execute on a GPU if the model is fitted with GPU arrays

For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

These Signal Processing Toolbox functions have new and enhanced gpuArray support:

  • instbw (Signal Processing Toolbox)

  • instfreq (Signal Processing Toolbox)

  • meanfreq (Signal Processing Toolbox)

  • poctave (Signal Processing Toolbox)

  • sosfilt (Signal Processing Toolbox) - Support for FIR second-order subsections of the input filter

For a list of all Signal Processing Toolbox functions with GPU functionality, see Functions with gpuArray support (Signal Processing Toolbox).

This Audio Toolbox function has enhanced gpuArray support :

For a list of all Audio Toolbox functions with GPU functionality, see Functions with gpuArray support (Audio Toolbox).

This Image Processing Toolbox function has new gpuArray support:

For a list of all Image Processing Toolbox functions with GPU functionality, see Functions with gpuArray support (Image Processing Toolbox).

Tall Arrays: Use functions with new and enhanced tall array support

These functions have new and enhanced tall array support:

  • allfinite

  • anymissing

  • anynan

  • rmoutliers - Support for defining outlier locations and returning the outlier indicator, thresholds, and center value outputs

  • round - Support for specifying a direction for rounding ties

For more information, see Tall Arrays.

Distributed Arrays: Use functions with new and enhanced distributed array support

These functions have new and enhanced distributed array support:

For more information, see Run MATLAB Functions with Distributed Arrays.

MATLAB Job Scheduler: Support for your own Java installation

You can now use your own Java 8 installation with MATLAB Job Scheduler (MJS). To use your own Java 8 installation with MJS, specify a path to a Java Runtime Environment (JRE) installation for the MJS_Java variable in your mjs_def file. For more information on where to find the mjs_def file, see Customize Startup Parameters (MATLAB Parallel Server).

Element-Wise Operations: Improved performance on GPU

Element-wise operations on large numbers of gpuArray objects show improved performance. Improvements are greater when you operate on large numbers of gpuArray objects. For example, adding one to every element of a cell array of the gpuArray data in this function is about 43.8x faster than in the previous release:

function timeElementWiseOps

% Prepare a cell array of gpuArray data
x = cell(5000,1);
x = cellfun(@(z) rand(10,"gpuArray"),x,"UniformOutput",false);

% Make all arrays different to one another
x = cellfun(@(z) z+rand(1,"gpuArray"),x,"UniformOutput",false);

% Time adding one to each element
gputimeit(@() cellfun(@(z) z+1,x,"UniformOutput",false))

end

The approximate execution times are:

R2022a: 5.70 seconds

R2022b: 0.13 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling the timeElementWiseOps function.

svd Function: Improved performance on GPU

The svd function shows improved performance when called with a gpuArray input matrix that is short and wide (n > 3m) or tall and narrow (m > 6n). For example, performing a singular value decomposition of a matrix in this function is about 1.4x faster than in the previous release:

function timeSVD

% Prepare input matrix
A = rand(1000,10000,"gpuArray");

% Time SVD
gputimeit(@() svd(A),3)

end

The approximate execution times are:

R2022a: 2.64 seconds

R2022b: 1.88 seconds

The code was timed on a Windows 10, Intel Xeon W-2133 @ 3.60 GHz test system with an NVIDIA RTX A5000 GPU by calling thetimeSVD function.

 Functionality being removed or changed

labXxx functions are not recommended and are renamed to spmdXxx

Still runs

To align these function names with their intended use within spmd blocks, these spmd block code execution and communication functions have been renamed:

These functions are no longer recommended, but they will continue to work. This table shows the recommended replacement functions.

FunctionalityRecommended ReplacementCompatibility Considerations
labBarrierspmdBarrierReplace all instances of labBarrier with spmdBarrier
labBroadcastspmdBroadcastReplace all instances of labBroadcast with spmdBroadcast
labindexspmdIndexReplace all instances of labindex with spmdIndex
labProbespmdProbeReplace all instances of labProbe with spmdProbe
labReceivespmdReceiveReplace all instances of labReceive with spmdReceive
labSendspmdSendReplace all instances of labSend with spmdSend
labSendReceivespmdSendReceiveReplace all instances of labSendReceive with spmdSendReceive
numlabsspmdSizeReplace all instances of numlabs with spmdSize

The labXxx functions will not be removed.

gcat, gop, and gplus are not recommended and are renamed to spmdCat, spmdReduce, and spmdPlus

Still runs

For performing SPMD operations on all the workers in an spmd block, gcat, gop, and gplus are no longer recommended. Use spmdCat, spmdReduce, and spmdPlus as direct replacements. The gcat, gop, and gplus functions will not be removed.

local profile has been renamed to Processes on the Parallel menu

Still runs

For a process-based parallel environment on a local machine, the local profile is no longer recommended. Use Processes instead.

  • To start a parallel pool of process workers, use this code.

    parpool("Processes")
    In previous releases, you used this code which is no longer recommended.
    parpool("local")

The local profile option has been removed from the Parallel menu but will continue to work when you use it programmatically.

parallel.defaultClusterProfile and parallel.clusterProfiles are renamed to parallel.defaultProfile and parallel.listProfiles

Still runs

The parallel.defaultClusterProfile and parallel.clusterProfiles functions are renamed. parallel.defaultClusterProfile is renamed to parallel.defaultProfile. parallel.clusterProfiles is renamed to parallel.listProfiles. The parallel.defaultClusterProfile and parallel.clusterProfiles will not be removed.

To update your code, replace all instances of parallel.defaultClusterProfile with parallel.defaultProfile and parallel.clusterProfiles with parallel.listProfiles.

remotecopy has been removed

Errors

Starting in R2022b, remotecopy (MATLAB Parallel Server) has been removed. To copy files to and from a remote host, use scp or sftp instead.

  • Previously you used remotecopy and -protocol scp to copy files to and from a remote host using the secure copy protocol (SCP).

    This table shows how to use scp instead.

    ErrorsRecommended
    remotecopy -remotehost host1 -local /my/file/path -to -remote /remote/file/path -protocol scp
    scp /my/file/path host1:/remote/file/path
    remotecopy -remotehost host1,host2 -local /my/file/path -to -remote /remote/file/path -protocol scp
    scp /my/file/path host1:/remote/file/path
    scp /my/file/path host2:/remote/file/path
    remotecopy -remotehost host1 -local /my/file/path -from -remote /remote/file/path -protocol scp
    scp host1:/remote/file/path /my/file/path
  • Previously you used remotecopy and -protocol sftp to copy files to and from a remote host using the secure file transfer protocol (SFTP).

    This table shows how to use sftp instead.

    ErrorsRecommended
    remotecopy -remotehost host1 -local /my/file/path -to -remote /remote/file/path -protocol sftp
    sftp /my/file/path host1:/remote/file/path
    
    remotecopy -remotehost host1,host2 -local /my/file/path -to -remote /remote/file/path -protocol sftp
    sftp /my/file/path host1:/remote/file/path
    sftp /my/file/path host2:/remote/file/path
    remotecopy -remotehost host1 -local /my/file/path -from -remote /remote/file/path -protocol sftp
    sftp host1:/remote/file/path /my/file/path

remotemjs has been removed

Errors

Starting in R2022b, remotemjs (MATLAB Parallel Server) has been removed. Use ssh instead.

Previously you used remotemjs to run MATLAB Job Scheduler commands on a remote host using -protocol ssh or -protocol winsc.

This table shows how to use ssh instead.

ErrorsRecommended
remotemjs <mjs options> -matlabroot <installfoldername> -remotehost host1
ssh host1 <installfoldername>/toolbox/parallel/bin/mjs <mjs options>
remotemjs <mjs options> -matlabroot <installfoldername> -remotehost host1,host2
ssh host1 <installfoldername>/toolbox/parallel/bin/mjs <mjs options>
ssh host2 <installfoldername>/toolbox/parallel/bin/mjs <mjs options>

R2022a

New Features, Compatibility Considerations

ValueStore and FileStore objects: Retrieve data and files on MATLAB clients during job execution

ValueStore and FileStore objects now allow you to store data and files from MATLAB workers that can be retrieved by MATLAB clients during the execution of a job (even while the job is still running). These objects are not held in system memory, so they can be used to store large results. A FileStore object provides a universal location for workers to copy files. MATLAB clients can then access these files regardless of the specific cluster environment where the worker is running.

The ValueStore and FileStore objects are automatically created when you create a job on a cluster, a parallel pool of process workers on your local machine, or a parallel pool of workers on a cluster of machines. To access the ValueStore and FileStore objects on a worker, use the getCurrentValueStore and getCurrentFileStore functions, respectively.

Multifactor Authentication for the Generic Scheduler Interface: Connect to remote clients with RemoteClusterAccess using two or more authentication factors

RemoteClusterAccess objects now allow you to perform multifactor authentication, including two-factor authentication. To do so, specify 'Multifactor' for the AuthenticationMode name-value argument. For details, see the RemoteClusterAccess reference page. The sample plugin scripts for third-party schedulers now accept 'Multifactor' as the value for the AuthenticationMode property of AdditionalProperties, which sets that value on RemoteClusterAccess.

Cluster resizing: Customize your MATLAB Job Scheduler cluster to resize automatically

You can customize your MATLAB Job Scheduler (MJS) cluster to resize automatically. After you set up auto-resizing, also called auto-scaling, the cluster can automatically change the number of workers with the amount of work submitted. The cluster grows (scales up) when there is more work to do and shrinks (scales down) when there is less work to do. This allows you to use your compute resources more efficiently and can result in cost savings. To learn more, see Set up MATLAB Job Scheduler Cluster for Auto-Resizing (MATLAB Parallel Server).

Parallel Pools: Check status of pools

Starting in R2022a, you can query if any type of pool is currently running by using the Busy property of the pool. This property indicates whether the parallel pool is busy, specified as true or false. The pool is busy if there is outstanding work for the pool to complete.

Background Pool: Number of thread workers no longer limited to 8 workers

Before R2022a, the NumWorkers property was capped at a maximum of 8 workers when Parallel Computing Toolbox was installed, and 1 worker when not installed. This limit is no longer in place when the Parallel Computing Toolbox is installed. When Parallel Computing Toolbox is not installed, the background pool remains limited to 1 worker.

Thread-Based Parallel Pool: See futures in active pool

Starting in R2022a, you can query all queued and running futures on a parallel pool by using the FevalQueue property of the pool. To create futures, use parfeval and parfevalOnAll. For more information on futures, see Future.

cancelAll Method: Cancel currently queued and running futures in a parallel pool

cancelAll cancels all futures currently queued or running in a parallel pool. Queued or running futures are listed in the FevalQueue property.

For example, you can use cancelAll to stop all Futures in an FevalQueue.

pool = parpool;
cancelAll(pool.FevalQueue);

GPU Functionality: Use new and enhanced gpuArray functions

New GPU support in MATLAB:

For more information, see Run MATLAB Functions on a GPU.

New GPU support in Statistics and Machine Learning Toolbox:

  • fitensemble (Statistics and Machine Learning Toolbox)

  • fitrensemble (Statistics and Machine Learning Toolbox)

  • New support for probability functions: gev*, gp*, nbin*

For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

The following functions have new and enhanced gpuArray support in Signal Processing Toolbox:

For a list of all Signal Processing Toolbox functions with GPU functionality, see Functions with gpuArray support (Signal Processing Toolbox).

The following functions have new gpuArray support in Audio Toolbox:

For a list of all Audio Toolbox functions with GPU functionality, see Functions with gpuArray support (Audio Toolbox).

The following functions have new gpuArray and dlArray support in Wavelet Toolbox:

For a list of all Wavelet Toolbox functions with GPU functionality, see Functions with gpuArray support (Wavelet Toolbox).

Support for NVIDIA CUDA 11.2: Update to CUDA Toolkit 11.2

The parallel computing products now use CUDA toolkit version 11.2. To generate CUDA kernel objects from CU code or compile CUDA compatible source code, libraries, and executables using GPU Coder™, you must use toolkit version 11.2. For more information, see Run CUDA or PTX Code on GPU.

Tall Arrays: Use new and enhanced tall array functionality

For more information, see Tall Arrays.

Distributed Arrays: Use new and enhanced distributed array functionality

For more information, see Run MATLAB Functions with Distributed Arrays.

Parallel Features in Other Products

Parallel features added in other products:

  • Experiment Manager: Offload deep learning experiments as batch jobs in a cluster

    Starting in R2022a, Experiment Manager (Deep Learning Toolbox) supports offloading experiments as batch jobs in a cluster. You can configure the cluster to run multiple trials at the same time or to run a single trial at a time on multiple parallel workers. While the experiment is running in the cluster, you can run other experiments, close the app and continue using MATLAB, or close your MATLAB session. For more information, see Offload Experiments as Batch Jobs to Cluster (Deep Learning Toolbox).

  • Parallel Simulations: Perform parameter sweeps using Parameter Combinations

    In R2022a, you can use Parameter Combinations in the Multiple Simulations panel of the Simulink® Editor for workflows with multiple simulations, such as Monte-Carlo simulations and parameter sweeps. Parameter Combinations allows you to create sequential and exhaustive combinations of parameters, specify value ranges, and run simulations with these combinations. The Multiple Simulations panel was introduced in R2021b. For more information, see Multiple Simulations Panel: Simulate for Different Values of Stiffness for a Vehicle Dynamics System (Simulink).

  • Machine Learning Apps: Train draft models in parallel or train in the background to keep the app responsive

    In Classification Learner (Statistics and Machine Learning Toolbox) and Regression Learner (Statistics and Machine Learning Toolbox), you can train multiple draft models in parallel. You can use the Use Parallel or Use Background buttons while training models.

Solving Linear System: Improved performance when solving linear systems A*X = B with gpuArray for symmetric positive definite matrices A

Solving a linear system of the form A*X = B with gpuArray by executing A\B shows improved performance when A is a symmetric positive definite matrix.

For example, this code solves A*X = B for a 10,000-by-10,000 symmetric positive definite matrix A and a 10,000-by-1 column vector B. The code is about 2.4x faster than in the previous release.

function timingTest
rng default
R = rand(10000,"gpuArray");
A = R'*R;
B = ones(10000,1,"gpuArray");
X = A\B;
end

The approximate execution times are:

R2021b: 1.03 s

R2022a: 0.43 s

The code was timed on a Windows 10, Intel Xeon CPU E5-1640 v3 @ 3.50 GHz with an NVIDIA Titan V GPU test system using the gputimeit function:

gputimeit(@timingTest)

 Functionality being removed or changed

distcomp folder removed, now named parallel

The distcomp folder has been removed.

In R2019b, the Parallel Computing Toolbox folder was renamed. Since then, the name of the toolbox folder has been parallel. If you need to reference the location of the toolbox, update your references to use toolbox/parallel instead of toolbox/distcomp.

R2021b

New Features, Compatibility Considerations

Parallel Language in MATLAB: Share parallel code with any MATLAB user

From R2021b, you can use more parallel language features in serial without Parallel Computing Toolbox. To scale up and speed up computations that use these features, use parallel pools and Parallel Computing Toolbox.

Additionally, you can now share parallel code with other MATLAB users who do not have Parallel Computing Toolbox.

The following features are now available in MATLAB:

  • parfeval and related functionality such as afterEach and afterAll

  • parallel.pool.DataQueue, parallel.pool.PollableDataQueue, and related functionality such as afterEach

For more information, see Background Processing and Write Portable Parallel Code.

GPU Functionality: Use new and enhanced gpuArray functions

For more information, see Run MATLAB Functions on a GPU.

GPU Functionality: Use new and enhanced gpuArray functions in Statistics and Machine Learning Toolbox

  • betafit (Statistics and Machine Learning Toolbox)

  • fitcecoc (Statistics and Machine Learning Toolbox)

  • fitcensemble (Statistics and Machine Learning Toolbox)

  • fitctree (Statistics and Machine Learning Toolbox)

  • fitdist (Statistics and Machine Learning Toolbox)

  • fitrtree (Statistics and Machine Learning Toolbox)

  • gevfit (Statistics and Machine Learning Toolbox)

  • gpfit (Statistics and Machine Learning Toolbox)

  • ksdensity (Statistics and Machine Learning Toolbox)

  • mle (Statistics and Machine Learning Toolbox)

  • mvksdensity (Statistics and Machine Learning Toolbox)

  • nbinfit (Statistics and Machine Learning Toolbox)

For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray support (Statistics and Machine Learning Toolbox).

GPU Functionality: Use new and enhanced gpuArray functions for working with signals and audio

Memory Usage: Use whos to check memory used by gpuArray and distributed variables

You can now use the whos function to check the amount of memory used by gpuArray and distributed variables.

Previously, the Bytes variable of the output of whos showed the number of bytes of the pointer to the variable in the local memory of the host machine. Now, the Bytes variable displays the amount of GPU or distributed memory allocated to that variable. For gpuArray variables, whos displays the amount of GPU memory used by that variable. For distributed variables, whos displays the total memory used by that variable across all workers in the pool.

Distributed Arrays: Use new and enhanced distributed array functionality

For more information, see Run MATLAB Functions with Distributed Arrays.

Thread-Based Environment: Use new and enhanced functionality on threads for working with audio, video, and images

For more information, see Run MATLAB Functions in Thread-Based Environment.

 Functionality Being Removed or Changed

parfeval and parfevalOnAll can now run in serial with no pool

Behavior change

Starting in R2021b, you can now run parfeval and parfevalOnAll in serial with no pool. This behavior allows you to share parallel code that you write with users who do not have Parallel Computing Toolbox.

When you use the syntaxes parfeval(fcn,n,X1,...,Xm) or parfevalOnAll(fcn,n,X1,...,Xm), MATLAB tries to use an open parallel pool if you have Parallel Computing Toolbox. If a parallel pool is not open, MATLAB will create one if automatic pool creation is enabled.

If parallel pool creation is disabled or if you do not have Parallel Computing Toolbox, the function is evaluated in serial. In previous releases, MATLAB threw an error instead.

mldivide and decomposition now produce the same results for distributed arrays

Behavior change

Starting in R2021b, mldivide and decomposition now produce the same results when you run the following code using distributed arrays A and b.

X = A \ b;
X = decomposition(A) \ b;
Previously, you sometimes saw different results when you used decomposition before mldivide for a non-square distributed matrix A.

Default behavior for decomposition has changed for distributed arrays

Behavior change

Starting in R2021b, the algorithm for decomposition with type set to 'auto' (default) has changed for distributed array input. You see this behavior change when you use decomposition to decompose a distributed matrix that is Hermitian, banded, or permuted triangular.

When you use decomposition(A) or decomposition(A,'auto') with a distributed matrix A, the type of decomposition selected by MATLAB is limited to the supported decomposition types for distributed arrays.

  • For a distributed dense matrix A, in the syntax decomposition(A,type) the decomposition types 'ldl', 'cod', and 'hessenberg' are not supported.

  • For a distributed sparse matrix A, in the syntax decomposition(A,type) the decomposition types 'chol', 'cod', and 'hessenberg' are not supported.