Metadata-Version: 1.0
Name: Products.UnicodeLexicon
Version: 2.2
Summary: Unicode aware lexicon for ZCTextIndex
Home-page: http://pypi.python.org/pypi/Products.UnicodeLexicon
Author: Stefan H. Holek
Author-email: stefan@epy.co.at
License: BSD
Description: ==============
        UnicodeLexicon
        ==============
        ---------------------------------------
        A Unicode aware lexicon for ZCTextIndex
        ---------------------------------------
        
        Motivation
        ==========
        
        The standard ZCTextIndex lexicon only deals with 8-bit strings (and only
        if you get the zope.conf *locale* setting right).
        It does not handle Unicode or UTF-8. UnicodeLexicon fills this gap.
        
        Installation
        ============
        
        This product adds a ZCTextIndex Unicode Lexicon type to Zope. The
        lexicon comes with word splitters, stop word removers, a case normalizer,
        and two accent normalizers.
        
        If you have GenericSetup installed, you can use the included extension
        profile to create a UnicodeLexicon in your portal_catalog and update
        the *Title*, *Description*, and *SearchableText* ZCTextIndexes.
        
        There is no upgrade path from UnicodeLexicon 1.0. If you have 1.0 on your
        system, you have to delete and recreate the lexicon.
        
        Pipeline Elements
        =================
        
        The splitter works with all languages that separate words with
        whitespace characters.
        
        The stop word remover knows about English language stop words only.
        
        The accent normalizer comes in two flavors. There is a normalizer for Latin
        and Western European text (fr, es, pt, it, en, nl), and one for German and
        Scandinavian text (de, dk, no, se, fi, is). The latter keeps the umlaut
        characters ä, ö, and ü in tact.
        
        Custom Pipeline Elements
        ------------------------
        
        Additional pipeline elements can be registered via ZCML. E.g.::
        
          <configure
            xmlns="http://namespaces.zope.org/zope"
            xmlns:unicodelexicon="http://namespaces.zope.org/unicodelexicon">
        
            <include package="Products.UnicodeLexicon" file="meta.zcml" />
        
            <unicodelexicon:registerPipelineElement
              group="Accent Normalizer"
              name="Normalize accented chars (Custom text)"
              factory="my.package.pipeline.MyCustomNormalizer"
              />
        
          </configure>
        
        Default Encoding
        ================
        
        The lexicon assumes either Unicode or UTF-8. If your application uses a
        different encoding, you can override the default by registering the encoding
        as a utility::
        
          <configure
            xmlns="http://namespaces.zope.org/zope">
        
            <utility
              provides="Products.UnicodeLexicon.interfaces.IDefaultEncoding"
              component="my.package.pipeline.defaultEncoding"
              />
        
          </configure>
        
        Related
        =======
        
        For CJK text you will want to use a ZCTextIndex lexicon in
        conjunction with CJKSplitter_. For Greek text you will want a
        ZCTextIndex lexicon with GRSplitter_.
        
        .. _CJKSplitter: http://www.zope.org/Members/panjunyong/CJKSplitter
        .. _GRSplitter: http://pypi.python.org/pypi/qi.GRSplitter
        
        
        Changelog
        =========
        
        2.2 - 2011-01-30
        ----------------
        
        - Allow to override the default encoding in ZCML.
          [stefan]
        
        2.1 - 2011-01-26
        ----------------
        
        - Add ability to register pipeline elements in ZCML.
          [stefan]
        
        - Fix a bug when updating PipelineFactory.
          [stefan]
        
        2.0 - 2011-01-21
        ----------------
        
        - Add an ordered PipelineFactory.
          [stefan]
        
        - Add an accent-normalizing pipeline element originally contributed by
          Marc-Auréle Darche.
          [stefan]
        
        - Release as Python egg.
          [stefan]
        
        1.0 - 2006-08-14
        ----------------
        
        - Initial release.
          [stefan]
        
Keywords: unicode zctextindex
Platform: UNKNOWN
Classifier: Development Status :: 5 - Production/Stable
Classifier: Framework :: Zope2
Classifier: License :: OSI Approved :: BSD License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 2
