Pages

Monday, May 19, 2014

org.alfresco.service.cmr.dictionary.DictionaryException: Could not import bootstrap model XXX.xml

To solve the above issue, check for typo/ syntax issues in model.xml

Error creating bean with name 'XXX' defined in class path resource [XXX/context/service-context.xml]: Invocation of init method failed; nested exception is java.lang.IllegalArgumentException: Class {XXX}XXX has not been defined in the data dictionary

I got the above error while adding custom behavior related to custom aspect using AMP (Alfresco Module Package) with Alfresco Maven SDK.

This is how I resolved that:

Each module requires a module context file (module-context.xml). This loads the Spring configuration for the module. 

Custom content models can be bootstrapped into the repository via Spring configuration added to the module context file.

You need to either import your custom-model-context.xml here or directly define the bean configuration for the custom model here as given below. 

  

References:

Module context file:http://docs.alfresco.com/4.2/index.jsp?topic=%2Fcom.alfresco.enterprise.doc%2Fconcepts%2Fdev-extensions-modules-module-context.html

Adding module data:
http://docs.alfresco.com/4.2/index.jsp?topic=%2Fcom.alfresco.enterprise.doc%2Fconcepts%2Fdev-extensions-modules-custom-model.html


Monday, April 14, 2014

How to show Sinhala characters in Mac OS X Terminal?

 
UTF-7 is the character encoding that represents Sinhala text.

    • Terminal > Preferences > Settings > Advanced 
    • International
    • Set Character encoding as Unicode (UTF-7)

    Tuesday, April 8, 2014

    Basic SOLR concepts explained with a simple use case

    This is an example scenario to understand the basic concepts behind SOLR/ Lucene indexing and search using advertising web site[4].

    Use case:

    Searcher: I want to search cars by different aspects such as car model, location, transmission, special features etc. Also, I want to see the similar cars that belong to same model as recommendations.

    SOLR uses index which is an optimized data structure for fast retrieval.

    To create an index, we need to come up with a set of documents with fields in it. How do we create a document for the following advertisement?



    Title: Toyota Rav4 for sale
    Category: Jeeps
    Location: Seeduwa
    Model: Toyota Rav4
    Transmission: Automatic
    Description: find later


    SOLR document:


    document 1

    Toyota Rav4 for sale


    Jeeps
    Seeduwa
    Toyota Rav4
    Automatic
    Brought Brand New By Toyota Lanka-Toyota Rav4 ACA21, YOM-2003, HG-XXXX Auto, done approx 79,500 Km excellent condition, Full Option, Alloy Wheels, Hood Railings,
    call No brokers please.


    Some more documents based on advertisements...

    document 2

    Nissan March for sale

    Cars
    Dankotuwa
    K11
    Automatic
    A/C, P/S, P/W, Center locking, registered year 1998, full option, Auto, New battery, Alloys, 4 doors, Home used car, Mint condition, Negotiable,

    document 3

    Nissan March K12 for rent
    Cars
    Galle
    K12
    Automatic
    A/C, P/S, P/W, Center locking, registered year 2004, full option, Auto, New battery, Alloys, 4 doors, cup holder, Doctor used car, Mint condition, Negotiable,

    Inverted Index


    Then SOLR creates an inverted index as given below: (Lets take example field as Title)

    toyota doc1(1x)
    rav4 doc1(1x)
    sale doc1(1x) doc2(1x)
    nissan doc2(1x)
    march doc2(1x)

    1x means the term frequency of the document for that particular field.


    Lucene Analyzers


    Note that “for” term here is eliminated during Lucene stop word removal process using Lucene text analysers. You can come up with your own analyser based on your preference as well.

    Field configuration and search

    You can configure, which fields your documents can contain, and how those fields should be dealt with when adding documents to the index, or when querying those fields using schema.xml.

    For example, if you need to index description field as well and the description value of the field should be retrievable during search, what you need to do is add the following line in schema.xml [1].

    Now, assume a user search for a vehicle.

    Search query: “nissan cars for rent”

    SOLR query would be /solr/select/?q=title:”nissan cars for rent"

    Ok what about the other fields (Category, location, transmission etc. ?)?

    By default, SOLR standard query parser can only search one field. To use multiple fields such as title and description and give them a weight to consider during retrieval based on their significance (boosts) we should use Dismax parser [2, 3]. Simply said, using Dismax parser you can make title field more important than description field. 



    Anatomy of a SOLR query


    q - main search statement
    fl - fields to be returned
    wt - response writer (response format)

    http://localhost:8983/solr/select?q=*:*&wt=json
    - select all the advertisements

    http://localhost:8983/solr/select?q=*:*&fl=title,category,location,transmission&sort=title desc
    - select title,category,location,transmission and sort by title in descending order

    wt parameter = response writer
    http://localhost:8983/solr/select?q=*:*&wt=json - Display results in json format
    http://localhost:8983/solr/select?q=*:*&wt=xml - Display results in XML format

    http://localhost:8983/solr/select?q=category:cars&fl=title,category,location,transmission -
    Give results related to cars only

    more option can be found at [5].

    Coming up next...

    • Extending SOLR functionality using RequestHandlers and Components
    • SOLR more like this
    References:
    [1] http://wiki.apache.org/solr/SchemaXml
    [2] https://wiki.apache.org/solr/DisMax
    [3] http://searchhub.org//2010/05/23/whats-a-dismax/
    [4] Ikman.lk
    [5] http://wiki.apache.org/solr/CommonQueryParameters

    Wednesday, April 2, 2014

    How would you decide if a class should be abstract class or interface?

    It depends :)

    In my opinion, to implement methods in abstract class you need to inherit the abstract class.One of the key benefits of inheritance is to minimise the amount of duplicate code by implement common functionalities in parent classes. so if the abstract class have some common generic behaviour that can be shared with its concrete classes, then using abstract class would be optimal.

    However, if all methods are abstract and those methods do not represent any unique/significant behaviour related to  the class instances, it may be better to use interface instead.

    Use abstract classes to define planned inheritance hierarchies. Classes with already defined inheritance hierarchy can extend their behavior in terms of the “roles” they can play, which are not common to its parents all the other children, using interfaces. Abstract classes will not help in this situation because of multiple inheritance restriction in  java language.

    How interfaces avoid “Deadly Diamond of Death” problem?

    A key difference between interface and abstract class is, “Interfaces simulate multiple inheritance” for languages where multiple inheritance is not supported due to “Deadly Diamond of Death” problem.

    How interfaces avoid “Deadly Diamond of Death” problem?

    Since interface methods do not have their underlying implementation, unlike the inherited class methods, there won’t be this problem as there can be multiple method signatures that are same, but there can be only one implementation for a particular class instance as duplicate methods cannot be compiled without any errors.

    Reference:
    Head First Java

    Thursday, March 13, 2014

    Constructor() has private access in Class

    If it is obvious for you that this has nothing to do with an issue on granting access, check for version incompatibilities of the .class or the related class.

    package java.nio.file does not exist in Mac OSX

    This is new addition in java 1.7, so if by default JDK is set as older version, this exception will be given. However, when I check java -version and it says java version "1.7.0_45".

    If you have java version specific code in your maven application add the following section in your pom.xml
               

        

       
       










    Still it will give the following error:
    [ERROR] Failed to execute goal X.plugins:maven-compiler-plugin:2.5.1:compile (default-compile) on project X: Compilation failure
    [ERROR] Failure executing javac, but could not parse the error:
    [ERROR] javac: invalid target release: 1.7
    [ERROR] Usage: javac
    [ERROR] use -help for a list of possible options


    To solve this issue, set the JAVA_HOME variable to the following using any of the following methods:

    // Set JAVA_HOME for one session
    export JAVA_HOME=/Library/Java/JavaVirtualMachines/jdk1.7.0_45.jdk/Contents/Home

    OR

    // Set JAVA_HOME for permanently
    vim ~/.bash_profile
    export JAVA_HOME=$(/usr/libexec/java_home)
    source .bash_profile
    echo $JAVA_HOME

    Now compile the application

    For those who are curious...

    When deciding which JVM to consider for compiling, path specified in JAVA_HOME is used. Here's how to check that.
    echo $JAVA_HOME

    If it is not specified in JAVA_HOME, using  the following command, you can see where JDK is located in your machine:
    which java

    It will give something like this: /usr/bin/java

    Try this to find where this command is heading to.
    ls -l /usr/bin/java

    This is a symbolic link to the path /System/Library/Frameworks/JavaVM.framework/Versions/Current/Commands

    Now try the following command:
    cd /System/Library/Frameworks/JavaVM.framework/Versions
    ls

    Check where "CurrentJDK" version is linked to. (Right click > Get info)
    Mine it was  /System/Library/Java/JavaVirtualMachines/1.6.0.jdk/Contents.

    Version specified as the "currentJDK" will determine which JVM should be used from the available JVMs.

    So, this is why I got the "package java.nio.file does not exist" at the first place, as the default referenced JDK is older than 1.7.

    How to point Current JDK to correct version?

    cd /System/Library/Frameworks/JavaVM.framework/Versions
    sudo rm CurrentJDK
    sudo ln -s /Library/Java/JavaVirtualMachines/jdk1.7.0_21.jdk/Contents/ CurrentJDK

    Additional info...

    Also, use the following command to verify from where the Java -version is read. (for fun!.. :))
    sudo dtrace -n 'syscall::posix_spawn:entry { trace(copyinstr(arg1)); }' -c "/usr/bin/java -version"

    It will output something like this:
    dtrace: description 'syscall::posix_spawn:entry ' matched 1 probe
    dtrace: pid 7584 has exited
    CPU     ID                    FUNCTION:NAME
      2    629                posix_spawn:entry   /Library/Java/JavaVirtualMachines/jdk1.7.0_45.jdk/Contents/Home/bin/java



    Saturday, February 22, 2014

    Named Entity Recognition using Conditional Random Fields (CRF)

    Named Entity Recognition

    Name Entity Recognition (NER) is a significant method for extracting structured information from unstructured text and organize information in a semantically accurate form for further inference and decision making.

    NER has been a key pre-processing step for most of the natural language processing applications such as information extraction, machine translation, information retrieval, topic detection, text summarization and automatic question answering tasks.

    In NER tasks frequently detected entities are Person, Location, Organization, Time, Currency , Percentage, Phone number, and ISBN.

    E.g., When translating "Sinhala text to English text", we need to figure out what are person names, locations in that text, so that we can avoid the overhead of finding corresponding English meaning for them. This is also helpful in question answering scenarios such as "Where Enrique was born?"

    Different methods such as rule based systems, statistical and gazetteers have been used in NER task, however, statistical approaches have been more prominent and other methods are used to refine the results as post processing mechanism.

    In computational statistics, NER has been identified as sequence labeling task and Conditional Random Fields (CRF ) has been successfully used to implement this.

    In this article I will use CRF++ to explain how to implement a named entity recognizer using a simple example.

    Consider the following input sentence:
    "Enrique was born in Spain"

    Now, by looking at this sentence any human can understand that Spain is a Location. But machines are unable to do so without previous learning.

    So, to learn the computer we need to identify a set of features that links the aspects of what we observe in this sentence with the class we want to predict, which in this case "Location". How can we do that?

    Considering the token/ word "Spain" itself is not sufficient to decide that it is a location in a generic manner. So, we consider its "context" as well which includes the previous/ next word, it's POS tag, previous NE tag etc. to infer the NE tag for token "Spain".

    Feature Template  and Training dataset

    In this example, I will use "previous word" as the feature. So, we will define this in feature template as given below:

    # Unigram
    U00:%x[-1,0]

    # Bigram
    B

    U00 is unique id to identify the feature. 

    I will explain %x[row, column]  using the following sentence that we going to train the model.
    I live in Colombo

    First, we need to define the sentence according to the following format. (training.data)
    I O
    live O
    in O
    Colombo Location

    current word: Colombo
    -1: in
    0: first column (Here, I have given only one column. But new columns are added when we define more features such as POS tag)
    In the above training file last column refers to the answers we give to model NE task.

    So, this feature indicates the model that after the word "in", it is "likely" to find a "Location".
    Now we train the model:

    crf_learn template train.data model

    Model file is generated using feature template and the training data file.

    Inference

    Now we need to know if the following sentence has any important entities such as Location.
    "Enrique was born in Spain"

    We need to format input file also according to the above format. (test.data)
    Enrique
    was
    born
    in
    Spain

    Now we use the following command to test the model.

    crf_test  -m model test.data

    Outcome would be the following:

    Enrique O
    was O
    born O
    in O
    Spain Location

    Likewise, the model will give predictions on entities present in the input files based on the given features and available training data.

    Note: Check the Unicode compatibility for different languages. E.g., for Sinhala Language it's UTF-7.

    Coming up next...
    • Probabilistic Graphic Models
    • Conditional Probability 
    • Finite State Automata
    • First order markov independence assumption

    Source code:
    https://bitbucket.org/jaywith/sinhala-named-entity-recognition

    Jayani Withanawasam

    Sunday, February 2, 2014

    Better approach to load resources using relative paths in Java

    FileInputStream (Absolute path)


    To load a resource file such as x.properties for program use, first thing that we would consider will be specifying the absolute file path as given below:

    InputStream input = new FileInputStream("/Users/jwithanawasam/some_dir/src/main/resources/
    config.properties”);

    However, when ever we moved the project to another location, this path has to be changed, which is not acceptable.

    FileInputStream (Relative path)


    So, the next option would be to use the relative file path as given below, instead of giving absolute file path:

    InputStream input = new FileInputStream("src/main/resources/config.properties”);

    This approach seems to solve the above mentioned concern.

    However, problem with this is the relative path is depending on the current working directory, which JVM is started. In this scenario, it is "/Users/jwithanawasam/some_dir". But, in a different deployment setting this may change, which leads to change the given relative path accordingly. Moreover, we, as developers do not have much control over JVMs current working directory.


    In any of the above cases, we will get java.io.FileNotFoundException error, which is a familiar exception for most java developers.


    class.getResourceAsStream


    At runtime, JVM checks the class path to locate any user defined  classes and packages. (In Maven, build artifacts and dependancies are stored under path given for M2_REPO class path variable. E.g., /Users/jwithanawasam/.m2/repository) The .jar file which is the deployable unit of the project will be located here.

    JVM uses class loader to load java libraries specified in class path.

    So, best thing we can do is load the resource specifying a path relative to its class path using class loader. Here, specified relative path will work  irrespective of the actual disk location the package is deployed.

    Following methods reads the file using class loader.

    InputStream input = Test.class.getResourceAsStream("/config.properties");

    Usually, in Java projects resources such as configuration files, images etc. are located in src/main/resources/ path. So, if we add a resource immediately inside this folder, during packaging, the file will be located in the immediate folder in .jar file.

    We can verify this using the following command to extract content of jar file:

    jar xf someproject.jar

    If you place the resources in another sub folder, then you have to specify the path relative to src/main/resources/ path.

    So, using this approach we can load resources using relative paths in a hard disk location independent manner. Once we package the application, it is ready to be deployed anywhere, as it it is, without the overhead of having to validate resource file paths, thus improving the portability of the application.

    ServletContext.getResourceAsStream for web applications


    For web applications, use the following method:

    ServletContext context = getServletContext();
        InputStream is = context.getResourceAsStream("/filename.txt");
     
    Here, file path is taken relative to your web application folder. (The unzipped version of the .war file)
    E.g., mywebapplication.war (unzipped) will have a hierarchy similar to the following.  
     
    mywebapplication
        META-INF
        WEB-INF
            classes
       filename.txt
     
    So, "/" means the root of this web application folder.  
    This method allows servlet containers to make a resource available to a servlet from any location, without using a class loader.