addPartOfSpeechDetails — Add part-of-speech tags to tokenized documents.
addPartOfSpeechDetails(documents) adds part-of-speech metadata to a RunMat tokenizedDocument compatibility object. Use tokenDetails to view the resulting PartOfSpeech column.
Syntax
updatedDocuments = addPartOfSpeechDetails(documents)
updatedDocuments = addPartOfSpeechDetails(documents,Name,Value)Inputs
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
documents | Any | Yes | — | tokenizedDocument object. |
NameValue | Any | Variadic | — | Name-value options: RetokenizeMethod, Abbreviations, DiscardKnownValues. |
Returns
| Name | Type | Description |
|---|---|---|
updatedDocuments | Any | Updated tokenized document object. |
Errors
| Identifier | When | Message |
|---|---|---|
RunMat:addPartOfSpeechDetails:InvalidInput | Input is not a supported tokenizedDocument object or option form. | addPartOfSpeechDetails: invalid input |
How addPartOfSpeechDetails works
documentsmust be a RunMattokenizedDocumentobject.addPartOfSpeechDetails(documents, Name, Value)supportsRetokenizeMethod,Abbreviations, andDiscardKnownValuesname-value pairs.RetokenizeMethoddefaults to"part-of-speech", which applies common part-of-speech retokenization such as English contractions,wanna/gonna, ellipses, and dotted initialisms. Use"none"to keep existing tokens.Abbreviationsis passed through to the sentence-detail prepass when sentence details are absent.DiscardKnownValuesdefaults tofalse, so existing known part-of-speech details are preserved while missing or empty entries are filled.DiscardKnownValues,truerecomputes every stored tag.- RunMat tags English and German tokens with deterministic compatibility taggers covering common pronouns, determiners, coordinate and subordinate conjunctions, adpositions, particles, auxiliary verbs, verbs, adverbs, adjectives, nouns, numerals, punctuation, and other token classes.
- If an existing tokenizedDocument object already carries Japanese or Korean language metadata, RunMat currently uses token-class-aware fallback tags and a conservative noun fallback for letter tokens. Public Japanese/Korean tokenizedDocument construction and MeCab-backed part-of-speech models remain tracked.
- Exact MathWorks statistical model parity remains tracked as broader Text Analytics work.
GPU memory and residency
addPartOfSpeechDetails updates host tokenizedDocument metadata and has no provider kernel.
Examples
Add Part-Of-Speech Details
documents = tokenizedDocument("The dogs are running.");
documents = addPartOfSpeechDetails(documents);
tdetails = tokenDetails(documents)Expected output:
`tdetails.PartOfSpeech` contains tags such as `"determiner"`, `"noun"`, `"auxiliary-verb"`, `"verb"`, and `"punctuation"`.Keep Existing Tokens
documents = addPartOfSpeechDetails(documents, "RetokenizeMethod", "none")Expected output:
`documents` receives part-of-speech details without the part-of-speech retokenization pass.Recompute Part-Of-Speech Details
documents = addPartOfSpeechDetails(documents, "DiscardKnownValues", true)Expected output:
`documents` has refreshed part-of-speech details.Using addPartOfSpeechDetails with coding agents
Open a RunMat example with live inputs, then ask the agent to explain how addPartOfSpeechDetails changes the result.
Run a small addPartOfSpeechDetails example, explain the result, then change one input and compare the output.
FAQ
Does addPartOfSpeechDetails use a full NLP model?⌄
No. This compatibility slice uses deterministic English rules and records exact MathWorks model parity as remaining Text Analytics work.
Does addPartOfSpeechDetails execute on the GPU?⌄
No. It updates host tokenizedDocument metadata.
Related Strings functions
Text Analytics
addDependencyDetails · addEntityDetails · addLemmaDetails · addSentenceDetails · addTypeDetails · bagOfNgrams · bagOfWords · cosineSimilarity · doc2sequence · encode · extractFileText · extractHTMLText · fastTextWordEmbedding · findElement · getAttribute · htmlTree · ind2word · isVocabularyWord · normalizeWords · readWordEmbedding · removeLongWords · removeShortWords · removeStopWords · removeWords · stopWords · tokenDetails · tokenizedDocument · trainWordEmbedding · vaderSentimentScores · vec2word · word2ind · word2vec · wordEncoding · writeWordEmbedding
Transform
append · deblank · erase · eraseBetween · erasePunctuation · eraseURLs · extractAfter · extractBefore · extractBetween · insertAfter · insertBefore · join · lower · pad · replace · replaceBetween · reverse · split · splitlines · strcat · strip · strjoin · strjust · strrep · strsplit · strtrim · upper
Core
blanks · char · compose · convertCharsToStrings · convertContainedStringsToChars · convertStringsToChars · genvarname · int2str · isletter · isspace · isStringScalar · isstrprop · mat2str · native2unicode · newline · num2str · sprintf · sscanf · str2double · str2num · strcmp · strcmpi · string · string.empty · strings · strlength · strncmp · strncmpi · strtok · unicode2native
Search
contains · endsWith · matches · startsWith · strfind
Pattern
digitsPattern · lettersPattern · pattern · regexpPattern · textBoundary · wildcardPattern
Open-source implementation
Unlike proprietary runtimes, every RunMat function is open-source. Read exactly how addPartOfSpeechDetails is executed, line by line, in Rust.
- View the source for addPartOfSpeechDetails in Rust on GitHub
- Learn how the RunMat runtime works
- Found a bug? Open an issue with a minimal reproduction.
About RunMat
RunMat is an open-source runtime that executes MATLAB-syntax code blazing on any GPU. It is licensed under the Apache 2.0 license.
- RunMat automatically optimizes your math for GPU execution on Apple, Nvidia, and AMD hardware. No code changes needed. Simulations that took hours now take minutes.
- Start running code in seconds. RunMat runs in the browser, on the desktop, or from the CLI. No license server, no IT ticket.