removeShortWords — Remove short words from tokenized documents or bag-of-words models.
removeShortWords(documents, len) removes tokens whose character length is less than or equal to len from RunMat tokenizedDocument compatibility objects. removeShortWords(bag, len) removes matching vocabulary columns from RunMat bagOfWords objects.
Syntax
newDocumentsOrBag = removeShortWords(documentsOrBag, len)Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
documentsOrBag | Any | Yes | — | tokenizedDocument or bagOfWords object. |
len | NumericScalar | Yes | — | Maximum word length to remove. |
Returns
| Name | Type | Description |
|---|---|---|
newDocumentsOrBag | Any | Filtered tokenizedDocument or bagOfWords object. |
Errors
| Identifier | When | Message |
|---|---|---|
RunMat:textAnalyticsDocuments:InvalidInput | Inputs do not match a supported Text Analytics document or model helper form. | Text Analytics document helper received invalid input |
How removeShortWords works
documentsmust be a RunMattokenizedDocumentobject created bytokenizedDocument.bagmust be a RunMatbagOfWordsobject created bybagOfWords.lenmust be a positive integer scalar.- For tokenized documents, the result preserves
Shape,TokenizeMethod, andLanguagemetadata while recomputing document lengths and vocabulary. - For bag-of-words models, the result preserves document row count and removes count columns for words with length less than or equal to
len. removeStopWords,removeWords, andremoveLongWordsare implemented separately for other tokenized-document filtering workflows. Table-backed custom token filters remain tracked by the Text Analytics umbrella.
GPU memory and residency
removeShortWords filters host text/model objects and has no provider kernel.
Examples
Filter Tokenized Documents
documents = tokenizedDocument("a short document");
newDocuments = removeShortWords(documents, 1)Expected output:
`newDocuments` contains `short` and `document`.Filter A Bag Of Words
documents = tokenizedDocument(["a beta"; "an gamma"]);
bag = bagOfWords(documents);
newBag = removeShortWords(bag, 2)Expected output:
`newBag.Vocabulary` contains `beta` and `gamma`.Using removeShortWords with coding agents
Open a RunMat example with live inputs, then ask the agent to explain how removeShortWords changes the result.
Run a small removeShortWords example, explain the result, then change one input and compare the output.
FAQ
Does removeShortWords remove words shorter than len or shorter than or equal to len?⌄
It removes words with character length less than or equal to len, matching the MATLAB Text Analytics contract.
Does removeShortWords support raw strings?⌄
No. Convert raw text with tokenizedDocument first, or build a bagOfWords model and filter that object.
Does removeShortWords execute on the GPU?⌄
No. It filters host text/model metadata and has no runmat-accelerate provider path.
Related Strings functions
Text Analytics
addDependencyDetails · addEntityDetails · addLemmaDetails · addPartOfSpeechDetails · addSentenceDetails · addTypeDetails · bagOfNgrams · bagOfWords · cosineSimilarity · doc2sequence · encode · extractFileText · extractHTMLText · fastTextWordEmbedding · findElement · getAttribute · htmlTree · ind2word · isVocabularyWord · normalizeWords · readWordEmbedding · removeLongWords · removeStopWords · removeWords · stopWords · tokenDetails · tokenizedDocument · trainWordEmbedding · vaderSentimentScores · vec2word · word2ind · word2vec · wordEncoding · writeWordEmbedding
Transform
append · deblank · erase · eraseBetween · erasePunctuation · eraseURLs · extractAfter · extractBefore · extractBetween · insertAfter · insertBefore · join · lower · pad · replace · replaceBetween · reverse · split · splitlines · strcat · strip · strjoin · strjust · strrep · strsplit · strtrim · upper
Core
blanks · char · compose · convertCharsToStrings · convertContainedStringsToChars · convertStringsToChars · genvarname · int2str · isletter · isspace · isStringScalar · isstrprop · mat2str · native2unicode · newline · num2str · sprintf · sscanf · str2double · str2num · strcmp · strcmpi · string · string.empty · strings · strlength · strncmp · strncmpi · strtok · unicode2native
Search
contains · endsWith · matches · startsWith · strfind
Pattern
digitsPattern · lettersPattern · pattern · regexpPattern · textBoundary · wildcardPattern
Open-source implementation
Unlike proprietary runtimes, every RunMat function is open-source. Read exactly how removeShortWords is executed, line by line, in Rust.
- View the source for removeShortWords in Rust on GitHub
- Learn how the RunMat runtime works
- Found a bug? Open an issue with a minimal reproduction.
About RunMat
RunMat is an open-source runtime that executes MATLAB-syntax code blazing on any GPU. It is licensed under the Apache 2.0 license.
- RunMat automatically optimizes your math for GPU execution on Apple, Nvidia, and AMD hardware. No code changes needed. Simulations that took hours now take minutes.
- Start running code in seconds. RunMat runs in the browser, on the desktop, or from the CLI. No license server, no IT ticket.