removeStopWords — Remove stop words from tokenized documents.
removeStopWords(documents) removes stop words from RunMat tokenizedDocument compatibility objects using the object's language metadata and the same compatibility lists returned by stopWords.
Syntax
newDocuments = removeStopWords(documents)
newDocuments = removeStopWords(documents, Name, Value, ...)Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
documents | Any | Yes | — | tokenizedDocument object. |
NameValue | Any | Variadic | — | Name-value options: IgnoreCase. |
Returns
| Name | Type | Description |
|---|---|---|
newDocuments | Any | Filtered tokenizedDocument object. |
Errors
| Identifier | When | Message |
|---|---|---|
RunMat:removeStopWords:InvalidInput | Inputs do not match a supported removeStopWords form. | removeStopWords: invalid input |
How removeStopWords works
documentsmust be a RunMattokenizedDocumentobject created bytokenizedDocument.removeStopWords(documents)is case-insensitive by default.removeStopWords(documents,'IgnoreCase',false)removes only tokens that match the stop-word list case exactly.- The result preserves
Shape,TokenizeMethod, andLanguagemetadata while recomputing document lengths and vocabulary. - RunMat removes stop words from derived
lettersandothertoken types. Punctuation, URLs, hashtags, mentions, and emoticons are preserved by this helper. - English and German tokenized documents are supported by RunMat's current tokenizer. Japanese and Korean stop-word lists exist for
stopWords, but MeCab-backed Japanese/Korean tokenized-document support remains tracked by the broader Text Analytics compatibility issue.
GPU memory and residency
removeStopWords filters host text metadata and has no provider kernel.
Examples
Remove English Stop Words
documents = tokenizedDocument(["an example of a short sentence"; "a second short sentence"]);
newDocuments = removeStopWords(documents)Expected output:
`newDocuments` contains `example short sentence` and `second short sentence`.Remove Stop Words Case Sensitively
documents = tokenizedDocument(["The" "the" "word"], "TokenizeMethod", "none");
newDocuments = removeStopWords(documents, "IgnoreCase", false)Expected output:
`newDocuments` contains `The word`.Using removeStopWords with coding agents
Open a RunMat example with live inputs, then ask the agent to explain how removeStopWords changes the result.
Run a small removeStopWords example, explain the result, then change one input and compare the output.
FAQ
Does removeStopWords support raw strings?⌄
No. Convert raw text with tokenizedDocument first.
Can I provide a custom word list?⌄
No. Use of custom lists belongs to removeWords, which remains tracked by the broader Text Analytics compatibility issue.
Does removeStopWords execute on the GPU?⌄
No. It filters host text metadata and has no runmat-accelerate provider path.
Related Strings functions
Text Analytics
addDependencyDetails · addEntityDetails · addLemmaDetails · addPartOfSpeechDetails · addSentenceDetails · addTypeDetails · bagOfNgrams · bagOfWords · cosineSimilarity · doc2sequence · encode · extractFileText · extractHTMLText · fastTextWordEmbedding · findElement · getAttribute · htmlTree · ind2word · isVocabularyWord · normalizeWords · readWordEmbedding · removeLongWords · removeShortWords · removeWords · stopWords · tokenDetails · tokenizedDocument · trainWordEmbedding · vaderSentimentScores · vec2word · word2ind · word2vec · wordEncoding · writeWordEmbedding
Transform
append · deblank · erase · eraseBetween · erasePunctuation · eraseURLs · extractAfter · extractBefore · extractBetween · insertAfter · insertBefore · join · lower · pad · replace · replaceBetween · reverse · split · splitlines · strcat · strip · strjoin · strjust · strrep · strsplit · strtrim · upper
Core
blanks · char · compose · convertCharsToStrings · convertContainedStringsToChars · convertStringsToChars · genvarname · int2str · isletter · isspace · isStringScalar · isstrprop · mat2str · native2unicode · newline · num2str · sprintf · sscanf · str2double · str2num · strcmp · strcmpi · string · string.empty · strings · strlength · strncmp · strncmpi · strtok · unicode2native
Search
contains · endsWith · matches · startsWith · strfind
Pattern
digitsPattern · lettersPattern · pattern · regexpPattern · textBoundary · wildcardPattern
Open-source implementation
Unlike proprietary runtimes, every RunMat function is open-source. Read exactly how removeStopWords is executed, line by line, in Rust.
- View the source for removeStopWords in Rust on GitHub
- Learn how the RunMat runtime works
- Found a bug? Open an issue with a minimal reproduction.
About RunMat
RunMat is an open-source runtime that executes MATLAB-syntax code blazing on any GPU. It is licensed under the Apache 2.0 license.
- RunMat automatically optimizes your math for GPU execution on Apple, Nvidia, and AMD hardware. No code changes needed. Simulations that took hours now take minutes.
- Start running code in seconds. RunMat runs in the browser, on the desktop, or from the CLI. No license server, no IT ticket.