Package org.apache.lucene.analysis.uk
Class UkrainianMorfologikAnalyzer
java.lang.Object
org.apache.lucene.analysis.Analyzer
org.apache.lucene.analysis.StopwordAnalyzerBase
org.apache.lucene.analysis.uk.UkrainianMorfologikAnalyzer
- All Implemented Interfaces:
Closeable,AutoCloseable
public final class UkrainianMorfologikAnalyzer
extends org.apache.lucene.analysis.StopwordAnalyzerBase
A dictionary-based
Analyzer for Ukrainian.- Since:
- 6.2.0
-
Nested Class Summary
Nested classes/interfaces inherited from class org.apache.lucene.analysis.Analyzer
org.apache.lucene.analysis.Analyzer.ReuseStrategy, org.apache.lucene.analysis.Analyzer.TokenStreamComponents -
Field Summary
Fields inherited from class org.apache.lucene.analysis.StopwordAnalyzerBase
stopwordsFields inherited from class org.apache.lucene.analysis.Analyzer
GLOBAL_REUSE_STRATEGY, PER_FIELD_REUSE_STRATEGY -
Constructor Summary
ConstructorsConstructorDescriptionBuilds an analyzer with the default stop words.UkrainianMorfologikAnalyzer(org.apache.lucene.analysis.CharArraySet stopwords) Builds an analyzer with the given stop words.UkrainianMorfologikAnalyzer(org.apache.lucene.analysis.CharArraySet stopwords, org.apache.lucene.analysis.CharArraySet stemExclusionSet) Builds an analyzer with the given stop words. -
Method Summary
Modifier and TypeMethodDescriptionprotected org.apache.lucene.analysis.Analyzer.TokenStreamComponentscreateComponents(String fieldName) Creates aAnalyzer.TokenStreamComponentswhich tokenizes all the text in the providedReader.static org.apache.lucene.analysis.CharArraySetReturns the default stopword set for this analyzerprotected ReaderinitReader(String fieldName, Reader reader) Methods inherited from class org.apache.lucene.analysis.StopwordAnalyzerBase
getStopwordSet, loadStopwordSet, loadStopwordSet, loadStopwordSetMethods inherited from class org.apache.lucene.analysis.Analyzer
attributeFactory, close, getOffsetGap, getPositionIncrementGap, getReuseStrategy, initReaderForNormalization, normalize, normalize, tokenStream, tokenStream
-
Constructor Details
-
UkrainianMorfologikAnalyzer
public UkrainianMorfologikAnalyzer()Builds an analyzer with the default stop words. -
UkrainianMorfologikAnalyzer
public UkrainianMorfologikAnalyzer(org.apache.lucene.analysis.CharArraySet stopwords) Builds an analyzer with the given stop words.- Parameters:
stopwords- a stopword set
-
UkrainianMorfologikAnalyzer
public UkrainianMorfologikAnalyzer(org.apache.lucene.analysis.CharArraySet stopwords, org.apache.lucene.analysis.CharArraySet stemExclusionSet) Builds an analyzer with the given stop words. If a non-empty stem exclusion set is provided this analyzer will add aSetKeywordMarkerFilterbefore stemming.- Parameters:
stopwords- a stopword setstemExclusionSet- a set of terms not to be stemmed
-
-
Method Details
-
getDefaultStopwords
public static org.apache.lucene.analysis.CharArraySet getDefaultStopwords()Returns the default stopword set for this analyzer -
initReader
- Overrides:
initReaderin classorg.apache.lucene.analysis.Analyzer
-
createComponents
protected org.apache.lucene.analysis.Analyzer.TokenStreamComponents createComponents(String fieldName) Creates aAnalyzer.TokenStreamComponentswhich tokenizes all the text in the providedReader.- Specified by:
createComponentsin classorg.apache.lucene.analysis.Analyzer- Returns:
- A
Analyzer.TokenStreamComponentsbuilt from anStandardTokenizerfiltered withLowerCaseFilter,StopFilter,SetKeywordMarkerFilterif a stem exclusion set is provided andMorfologikFilteron the Ukrainian dictionary.
-