Thread-Based Parallel Pools: Use ValueStore and
FileStore objects on thread workers
You can now use ValueStore and FileStore objects on Parallel Computing Toolbox™
ThreadPool workers. The software creates
ValueStore and FileStore objects when you
create a ThreadPool object on your local machine. To access the
ValueStore and FileStore objects on
ThreadPool workers, use the getCurrentValueStore and getCurrentFileStore functions, respectively.
Thread-Based Environment: Use new functionality on thread workers
These MATLAB® functions now have thread-based support:
For more information, see Run MATLAB Functions in Thread-Based Environment.
Parallel Pools: Efficient scheduling of parfor and
parfeval computations on process-based parallel
pools
MATLAB now efficiently schedules parfor and parfeval computations on process-based parallel pools as pool
workers become available. You can now run parfeval and
parfor computations on process-based pools on a local
machine or a remote cluster concurrently.
For example, this code calls the parfevalWithparfor function,
which runs a long-running parfeval computation in the
background and then executes a parfor-loop. The code is about
1.5x faster than in the previous release because the software runs the
parfeval and parfor computations
simultaneously, removing 14 seconds of scheduling
overhead.
function t = poolTimingTest % Start parallel pool if none exists pool = gcp("nocreate"); function parfevalWithparfor future = parfeval(@pause,0,30); parfor ii = 1:10 pause(ii); end wait(future); end % Time executing parfor and parfeval concurrently t = timeit(@() parfevalWithparfor); end
The approximate execution times are:
R2023a: 44.03 seconds
R2023b: 30.03 seconds
In previous releases, MATLAB schedules parfor or
parfeval computations to run separately. If you have
code that relies on this behavior, update your code to avoid compatibility
issues.
Parallel Workflow Examples: New and updated examples and topics
This new topic contains information about accelerating code in MATLAB and using parallel computing capabilities to efficiently run code on multicore and multiprocessor computers:
This new topic helps you choose the right data management tools and workflows for your needs:
This new example shows how to send data to workers in a data queue:
This updated example shows to compare how fast functions run on the client and on a parallel pool:
Support for Apple silicon Macs
Parallel Computing Toolbox now supports Apple silicon Macs with this limitation:
Distributed and codistributed arrays are not supported for local process pools.
gpurng Function: Specify random number generator without
specifying seed
You can now use the new syntax gpurng(generator) to specify the
algorithm that the random number generator on the GPU uses. Use this syntax to set
the random number algorithm without specifying the seed. The
gpurng function uses a default seed of 0. This syntax is
equivalent to gpurng(0,generator). For example,
gpurng("philox") initializes the Philox 4x32 generator with a
seed of 0. For more information, see gpurng.
GPU Support for switch, case, and
otherwise inside arrayfun
You can now use switch conditional statements in functions you apply using arrayfun with gpuArray input. This functionality
has these limitations:
Case expressions support only numeric and logical values.
Using a cell array as the case expression to compare the switch expression
against multiple values, for example, case {x1,y1} is not
supported.
GPU Functionality: Use functions with new and enhanced
gpuArray support
These MATLAB functions have new and enhanced gpuArray support:
These Statistics and Machine Learning Toolbox™ functions have new gpuArray support:
For a list of all Statistics and Machine Learning Toolbox functions with GPU functionality, see Functions with gpuArray
support (Statistics and Machine Learning Toolbox).
These Signal Processing Toolbox™ functions have new gpuArray support:
ifsst (Signal Processing Toolbox)
rpmfreqmap (Signal Processing Toolbox)
rpmordermap (Signal Processing Toolbox)
tfestimate (Signal Processing Toolbox)
For a list of all Signal Processing Toolbox functions with GPU functionality, see Functions with gpuArray
support (Signal Processing Toolbox).
These Communications Toolbox™ functions have new gpuArray support:
awgn (Communications Toolbox)
ldpcDecode (Communications Toolbox)
ofdmEqualize (Communications Toolbox)
qamdemod (Communications Toolbox)
qammod (Communications Toolbox)
For a list of all Communications Toolbox functions with GPU functionality, see Functions with gpuArray
support (Communications Toolbox).
This Wavelet Toolbox™ function has new gpuArray support:
dwtleader (Wavelet Toolbox)
For a list of all Wavelet Toolbox functions with GPU functionality, see Functions with gpuArray
support (Wavelet Toolbox).
GPU Arrays: Improved performance
Some workflows using gpuArray objects show improved performance. For example, simulating
Conway's "Game of Life" on the GPU in this example is about 2x faster than in the
previous release:
function timeGameOfLife % Select GPU device. gpu = gpuDevice; wait(gpu) % Start timing. tic % Define simulation parameters. gridSize = 1000; numGenerations = 5000; initialGrid = (rand(gridSize,gridSize) > .75); grid = gpuArray(initialGrid); p = [1 1:gridSize-1]; q = [2:gridSize gridSize]; % Loop over generations. for generation = 1:numGenerations % Count number of neighbors. neighbours = grid(:,p) + grid(:,q) + grid(p,:) + grid(q,:) + ... grid(p,p) + grid(q,q) + grid(p,q) + grid(q,p); % Update the grid. A live cell with two live neighbors, or any cell with % three live neighbors, is alive at the next step. grid = (grid & (neighbours == 2)) | (neighbours == 3); end % Gather back to host memory. gather(grid); wait(gpu) % Record the elapsed time. t = toc; end
The approximate execution times are:
R2023a: 2.8 s
R2023b: 1.2 s
The code was timed on a Windows® 10, Intel®
Xeon® W-2133 @ 3.60 GHz test system with an NVIDIA® RTX A5000 GPU by calling the timeGameOfLife
function.
Improved Scalability: Use MATLAB Job Scheduler clusters with up to 10,000 workers
MATLAB Parallel Server™ with MATLAB Job Scheduler now supports clusters with up to 10,000 workers. Support for large parallel pools remains at 1024 workers.
When you scale above 1000 workers, you must increase the heap memory available to the job manager. For more information, see Customize Startup Parameters (MATLAB Parallel Server).
Cluster Scheduling: Specify load-balancing scheduling algorithm for MATLAB Job Scheduler
You can now select a scheduling algorithm for your MATLAB Job Scheduler that balances the workload more evenly across your
cluster nodes. Specify the scheduling algorithm using the
SCHEDULING_ALGORITHM parameter in the
mjs_def file. For more details, see Define MATLAB Job Scheduler Startup Parameters (MATLAB Parallel
Server).
Big Data Workflows: Convert between tall arrays and distributed arrays
You can now convert a tall array to a distributed array to access MATLAB functions that have distributed array support. To convert tall arrays
to distributed arrays, use the distributed function with a tall array.
You can also convert a distributed array to a tall array to access functions that
have tall array support. To convert distributed arrays into tall arrays, use the
tall function with a distributed array.
Using a distributed array in the tall function or a tall
array in the distributed function throws an error in
earlier releases. If your code relies on the errors that earlier releases of
MATLAB throw for those conversions, such as within
a try/catch block, update your
code so it does not rely on those errors.
Distributed Arrays: Faster distribution of local arrays to workers
Creating distributed arrays from large local arrays shows improved performance.
For example this code calls the distributeLargeArray function
which distributes a large array to the workers in a parallel pool. The code is about
4.8x faster than in the previous
release.
function t = distributedTimingTest
% Start parallel pool if none exists
pool = gcp("nocreate");
% Prepare large array
W = triu(gallery("wathen", 1000, 1000));
f = @() distributed(W);
% Time distributing large array
t = timeit(f);
endThe approximate execution times are:
R2023a: 4.20 seconds
R2023b: 0.87 seconds
The code was timed on a Windows 10, Intel(R) Xeon(R) CPU E5-1650 v3 @ 3.50 GHz
test system using the distributed function on a parallel pool
with six workers.
Distributed Arrays: Use functions with new distributed array support
These functions have new distributed array support:
For more information, see Run MATLAB Functions with Distributed Arrays.
Functionality being removed or changed
arrayfun with GPU Arrays: Passing arrays from parent
workspace to nested function and indexing into array within nested function now
errors
Behavior change
In this code, you create the parentWorkspaceVar variable in
the parent workspace of the foo function. If you use
foo in an arrayfun call with gpuArray input and if the foo function passes
parentWorkspaceVar as an input to a nested function
within foo, the code errors.
Instead of passing the parent workspace variable
(parentWorkspaceVar) to the nested function
(bar), use the parent workspace variable directly. This
variable is already in the scope of the nested function.
| Errors | Alternative |
|---|---|
function y = exampleFunction parentWorkspaceVar = 1:9; x = ones(2,"gpuArray"); y = arrayfun(@foo,x); function y = foo(x) y = bar(parentWorkspaceVar); function y = bar(z) % Errors y = z(1); end end end |
function y = exampleFunction parentWorkspaceVar = 1:9; x = ones(2,"gpuArray"); y = arrayfun(@foo,x); function y = foo(x) y = bar; function y = bar y = parentWorkspaceVar(1); % Use parent workspace variable directly. end end end |
arrayfun with GPU Arrays: Functions writing into variables
created in parent function now error
Behavior change
In this code, you create the workspaceVar variable in the
workspace of the bar function. If you use
bar in an arrayfun call with gpuArray input and if a nested function foo
writes into workspaceVar, the code errors.
Instead of writing to the variable (workspaceVar) within
the nested function (foo), add another output to the nested
function and use the output to write to the variable.
| Errors | Workaround |
|---|---|
function x = exampleFunction z = ones(2,"gpuArray"); x = arrayfun(@bar,z); function y = bar(z) workspaceVar = 2; foo(z); function x = foo(z) workspaceVar = 10; % Errors x = z; end y = workspaceVar; end end |
function x = exampleFunction z = ones(2,"gpuArray"); x = arrayfun(@bar,z); function y = bar(z) workspaceVar = 2; [out,workspaceVar] = foo(z); % Write to workspaceVar outside foo. function [x,y] = foo(z) % Add another output y to the nested function. y = 10; x = z; end y = workspaceVar; end end |
parallel.pool.Constant with no arguments now returns
invalid Constant object
Behavior change
When you call the parallel.pool.Constant function without input arguments, it
initializes a Constant object in an invalid state. In previous
releases, calling the parallel.pool.Constant function
without input arguments errors.
You can use parallel.pool.Constant with no arguments to
assign invalid Constant objects to array elements. When you
create or grow an array of Constant objects without assigning
values to each element, any new elements of the array contain invalid
Constant elements.