Changes

Jump to navigation Jump to search
265 bytes removed ,  15:36, 19 November 2011
==Languages==
The chosen language determines how the input text will be tokenized and stemmed. The basic code base and <code>jar</code> file of BaseX comes with built-in support for English and German. More languages are supported if the following libraries are placed found in the classpath ({{Version|7.0}}):
* [http://files.basex.org/maven/org/apache/lucene-stemmers/3.4.0/lucene-stemmers-3.4.0.jar lucene-stemmers-3.4.0.jar]: includes Snowball and Lucene stemmers and extends language support to the following languages: Arabic, Bulgarian, Catalan, Czech, Danish, Dutch, Finnish, French, Hindi, Hungarian, Italian, Latvian, Lithuanian, Norwegian, Portuguese, Romanian, Russian, Spanish, Swedish, Turkish.
* [http://en.sourceforge.jp/projects/igo/releases/ igo-0.4.3.jar]: A big thank Thank you goes out to [http://blog.infinite.jp Toshio HIRAI] for integrating the Japanese [http://igo.sourceforge.jp/ IGO lexer] in BaseX! In addition to the library, one of the following dictionary files must either be unzipped into the current directory, or into the <code>etc</code> sub-directory of the project’s . [[Configuration#Home DirectoryFull-Text/Japanese|Home DirectoryAn additional article]]::– IPA Dictionary: http://files.basex.org/etc/ipadic.zip:– NAIST Dictionary: http://files.basex.org/etc/naistdicexplains how IGO can be integrated, and how Japanese texts are tokenized and stemmed.zip
The JAR files can also be found in the <code>zip</code> and <code>exe</code> distribution files of BaseX.
Bureaucrats, editor, reviewer, Administrators
13,554

edits

Navigation menu