Text Chunking: Split large documents into chunks for retrieval-augmented generation (RAG)
You can improve the quality of AI generated text by providing a large language model (LLM) with additional information from a set of documents. This process is known as retrieval-augmented generation (RAG). Use text chunking to split a large document into smaller documents, then choose only the relevant document chunks to give to the LLM.
To get started, use the splitTextChunks function to recursively split the text into
paragraphs, sentences, tokens, and characters until each chunk is smaller than the
target length. For more information on text chunking, see Text Chunks. For more information on RAG, see Retrieval-Augmented Generation (RAG).
For a more customized workflow, use any of these additional functions.
Split your document into sections and preserve the section metadata using one of these functions:
splitHTMLSections | Split an HTML-formatted document into HTML
sections according to the section tags
|
splitMarkdownSections | Split a Markdown-formatted document into Markdown
sections, for example according to ATX section tags
# , ## , …,
###### . |
splitCustomSections | Split a document into custom sections according to custom section delimiters. |
Split your documents or your chunks recursively into paragraphs,
sentences, and tokens using the splitTextChunks function.
To avoid redundancy, join similar adjacent chunks using the joinSimilarTextChunks function.
Add overlap between adjacent text chunks using the addTextChunkOverlap function.
After you have identified relevant text chunks, use the findTextChunkContext function to identify other text
chunks in the same section, subsection, etc., that you can add to the
relevant text chunk to provide additional context.
Finally, use the formatTextChunks function to generate a
Markdown-formatted prompt to provide to an LLM.
For an example showing this workflow, see Split Document Into Semantically Meaningful Text Chunks.
Use an LLM to generate text based on the retrieved text chunks. To connect to LLM APIs using MATLAB®, use the Large Language Models (LLMs) with MATLAB add-on. You can download Large Language Models (LLMs) with MATLAB from the Add-On Explorer.
Text Preprocessing: eraseURLs supports
tokenizedDocument
The eraseURLs function now supports tokenizedDocument inputs.
Tokenizers: Access vocabulary size
You can now access the vocabulary size of bpeTokenizer and bertTokenizer objects using the new NumTokens
property.
For example, you can use the vocabulary size to specify the number of words when
you create a wordEmbeddingLayer object.