You can provide the terms that you want to provide information on in the text parameter. You can also return information about the terms in a document, by providing the document in the file, reference, or url parameter.
By default, Haven OnDemand compares the terms in your document to the terms in the English Wikipedia public text index. You can optionally specify another text index, by setting the indexes parameter. In this case, Haven OnDemand provides information on your terms based on the index or indexes that you specify.
The results of the Text Tokenization API include:
the weight that the term holds in the specified text indexes, based on Advanced Probabilistic Concept Modelling (APCM).
the number of documents that the term occurs in, in the specified text indexes.
the total number of times that the term occurs, in the specified text indexes.
For example:
/1/api/[async|sync]/tokenizetext/v1?text=probability+theory
The term included in the response is the term that Haven OnDemand uses after processing. The processing includes stemming (reducing plurals and verb forms of a word to the same stem) and transliteration (converting accented characters to non-accented forms).
{
"terms": [
{
"term": "PROBAB",
"weight": 77,
"documents": 49610,
"occurrences": 191725,
"case": 1,
"length": 6
},
{
"term": "THEOR",
"weight": 52,
"documents": 226510,
"occurrences": 1502845,
"case": 1,
"length": 5
}
]
}
When you specify multiple text indexes, the API returns each term once, with information from each text index reflected in the occurrence counts and weight information. If a term that you specify does not occur in the specified text indexes, the API returns the term with a default weight (the occurrence counts are zero).
If a term that you specify is a stop word (a very common term that does not provide much meaning, such as the, a, of in English), the API returns the term, but the occurrence counts and weights are all zero, because Haven OnDemand does not store these terms. For more information about stop words, see Stop Lists in Text Indexes.
You can use the Text Tokenization API to find information about the terms that Haven OnDemand matches in the Query Text Index API and other related APIs, such as Find Related Concepts and Find Similar. You can provide your complete query text in the text parameter, and the API ignores query syntax, such as Boolean and proximity operators, and provides information about the query terms. You can use this approach to find out the exact terms that Haven OnDemand matches in the query.
For example:
/1/api/[async|sync]/tokenizetext/v1?text=lunar+NEAR+crater&indexes;=news_eng
{
"terms": [
{
"term": "LUN",
"weight": 106,
"documents": 1310,
"occurrences": 2440,
"case": 1,
"length": 3
},
{
"term": "CRAT",
"weight": 132,
"documents": 369,
"occurrences": 617,
"case": 1,
"length": 4
}
]
}
If you want to tokenize the query operators as normal text, you can set the ignore_operators parameter to true. For example:
/1/api/[async|sync]/tokenizetext/v1?text=lunar+NEAR+crater&indexes;=news_eng&ignore_operators=true
{
"terms": [
{
"term": "LUN",
"weight": 106,
"documents": 1310,
"occurrences": 2440,
"case": 1,
"length": 3,
"start_pos": 1
},
{
"term": "NEAR",
"weight": 45,
"documents": 25815,
"occurrences": 31791,
"case": 0,
"length": 4,
"start_pos": 7
},
{
"term": "CRAT",
"weight": 132,
"documents": 369,
"occurrences": 617,
"case": 1,
"length": 4,
"start_pos": 12
}
]
}
For more information about text tokenization in Haven OnDemand, see Text Tokenization and Processing in Text Indexes.

