Tuesday, August 28, 2012

N-gram features for text classification

Traditionally, text classification has relied on bag-of-words count features. For some experiments, I was  wondering if using n-gram counts could make for a good feature set. Once I generated the features, I knew I was in trouble. For the WSJ corpus, I got about 20 million features for a trigram model. Just checked out the literature and found this paper that n-gram features don't help much:


A Study Using n-gram Features for Text Categorization, Johannes Furnkranz


Bigram and trigram features may give modest gains, but feature selection is obviously required. Feature selection based on document frequency, term frequency would be a simple approach.


Thursday, August 23, 2012

Origins of the Brahmi Script

This post is motivated by chapter 2 of James Gleick's book '', which discusses the evolution of writing.  Brahmi is the mother script from which the scripts of all modern Indian and South-East Asian languages have evolved. It was first seen in Emporor Ashoka's rock edicts dating tno the 3rd century B.C.  It is then one of the ancient world's "alphabets" - along with Greek, Phoenician and Aramaic.  The alphabet is based on the idea that symbols represent phonemes in contrast to other writing systems like logographic (e.g. Chinese which employs symbols for words) or syllabic (e.g. Japanese where symbols represent syllables). 

All the alphabetic scripts are said to be derived from a single script, the Phoenician. In fact, the very word 'alphabet' comes from the first two symbols in the Greek script 'Alpha' and 'Beta'. There is a lack of clarity on the origin of the Brahmi script, with two primary categories of theories. One propounds that the Brahmi evolved from the Aramaic script (itself an evolution over the Phoenician). This is based on the proposed orthographic similarities between symbols in the scripts. (See Figure).

The other theory proposes an indigenous development of the Brahmi script, based on the wide differences in how the writing systems work. I tend to favour this theory, though I must admit that my knowledge of this area is limited to reading a few articles and knowing some of the modern day descendants of these scripts. The modern day alphabet of Indian scripts are organized phonetically, and there is little ambiguity phonetically - as opposed to the Roman scripts. The earliest Semitic scripts (Phoenician, Aramaic) and even modern Arabic do not have vowels, whereas the so called "true" alphabets Greek and its modern Latin derivative scripts still have room for ambiguity. Even if there was some use of symbols from the Aramaic scripts, the design seems pretty novel to call it a new style of scripting. Is there an alternative line of evolution of the script? The Indus Valley script is still undecipered - could the Brahmi have evolved from there?    


Sunday, February 12, 2012

Indian English

From Chandan Mitra's weekly column in the Pioneer, some hilarious examples of English usage:


In a newspaper, describing a case of chain-snatching in which criminals shot dead the man who tried to resist and pursue the chain-snatchers, the reporter stated: “The deceased gave chase to the criminals who, however, managed to escape”!

Police notice: “Take care of belongings. You may be theft”


The article is interesting reading too.
http://dailypioneer.com/columnists/item/51044-dont-fast-you-may-be-theft-indlish-is-on-a-roll.html



Saturday, January 14, 2012

Yet Another Moses Installation Guide

Though Moses is a versatile MT system, its installation is still from stone age. Let me document here some of the key points to navigate through the installation of Moses. The intent is not to present a complete installation guide, but to highlight key issues that may crop up (as they cropped up for me). For a complete installation, this is probably the best guide. Another useful installation guide can be found here.

To install the Moses system, the following tools need to be installed. 
  • Language modelling toolkit (SRILM, IRSTLM, etc.)
  • GIZA++ package which contains GIZA++ and mkcls
  • Moses decoder (version 1.0 and above)

SRILM installation
  • The primary installation reference is the INSTALL document that ships with the tool.
  • Install all pre-requisites mentioned in the SRILM installation guide. On Ubuntu I had to install the following packages: csh, g++-multilib, tcl-dev
  • Set the environment variable SRILM to point to the base directory of the install package before building SRILM.
  • Following the instruction manual with the SRILM download should be enough once the pre-requisites are installed.    
  • The problems you may yet face are
    • Problem in identifying the architecture, especially if it a 64-bit machine. To make sure that the install script correctly identifies the architecture, set the variable MACHINE_TYPE in sbin/machine-type.
    • Problems with TCL compilation. You may not need the TCL user interfaces at all, so it may just be able ok to disable their compilation. Set the variable NO_TCL = X in the file common/your_architecture_specific_makefile.         
  • Make sure you have added the $SRILM/bin and $SRILM/bin/$MACHINE_TYPE to the PATH variable
  • Note: SRILM 1.7.1 and above are not compatible with Moses
IRSTLM installation
  • Ubuntu packages required: libtool make autoconf autotools-dev automake
  • The installation is pretty simple, just have to follow the installation guide
  • One caveat: Sometimes, it may be required to create a directory named 'm4' manually, if the first step mails

GIZA++ and mkcls installation
  • You get both if you download the giza-pp tool. 
  • Most straightforward installation. Download and 'make'.
  • Copy the binaries - GIZA++, mkcls, snt2cooc.out to a new directory. 
XMLRPC Server
  • XML RPC Server is required if you want to run a webservice providing translations. If you just want to get Moses running, you can skip this step.
  • Install the following packages: libxmlrpc-core-c3 libxmlrpc-core-c3-dev libxmlrpc-c3-dev libxmlrpc-c++4 libxmlrpc-c++4-dev 
Boost Library
The C++ Boost library  is required for installation of  Moses. Boost 1.48 has a serious bug which breaks Moses compilation. Unfornately, some Linux distributions (eg. Ubuntu 12.04) have broken versions of the Boost library.To fix this situation you can:
  • For Ubuntu 12.04: Remove boost 1.48 from your distribution and install Boost 1.46 which is available in the distribution. This works most of the time. If not, build Boost from source as described below.
  • To install Boost manually and making it work with Moses, follow the instructions in the section titled "Manually Installing Boost" on this page: http://www.statmt.org/moses/?n=Development.GetStarted
Moses installation
  • The primary installation reference is the INSTALL document that ships with the tool.
  • SRILM or IRSTLM need to be installed before Moses is installed
  • Make sure you have installed the packages  automake and libtool
  • Boost has to be installed
  • It is then a matter of just following the instructions. The command to be run is
  • /usr/bin/bjam --with-srilm=  --with-xmlrpc-c= --with-boost=
    • If the xml RPC is installed in /usr/bin, then the parameter would simply be '/usr'
    • --with-boost is required only when Boost is installed in a non-standard directory. The path should contain both lib/lib64 and include directories
Now Moses is ready to cross the Red Sea.


Alternative ways of installation Moses

If you fail to install from the source as mentioned above, then there are a couple of simpler alternatives you can try:

One, use the pre-compiled binaries provided by the Moses team: 
The pre-compiled version comes with IRSTLM and does not support XML-RPC to the best of my knowledge. However, it is handy to get started. 

If that too runs into trouble, then you can try using the virtual machine provided by the Moses team. 


If you are using Virtual Box, you can import the OVA images into VirtualBox. 
This guide many be useful for importing OVA images into VirtualBox:
http://www.maketecheasier.com/import-export-ova-files-in-virtualbox/

I have not tried the Virtual Images, so let me know if it works. 

Friday, September 23, 2011

Incorporating Linguistic Information into SMT Models

(Summary of the chapter 'Integrating Linguistic Information' in Philip Koehn's textbook 'Statistical Machine translation')


Traditional phrase based Statistical Machine Translation (SMT) has relied only on the surface form of words, but this can carry you only so far. Without considering any linguistic phenomena, there is no generalization possible and the SMT system ends up being a translation memory. Various kinds of linguistic information needs to be incorporated into the SMT process like: 

  • Name Transliteration and Number script conversions
  • Morphology changes - inflections, compounding, segmentation - these problems if not handled lead to data sparsity problems
  • Syntanctic phenomena like constituent structure, attachment, head-modifier re-orderings. Vanilla SMT is designed to handle local re-orderings but long range dependencies are not handled well. 

One way to handle them is to pre-process the parallel corpus before training and then run the SMT tools. Pre-processing could include:

  • Transliteration and back transliterations models need to be incorporated. An important problem is to identify the named entities in the first place.
  • Splitting words for a morphology rich input language. Compounding and segmentation can be handled similarly. 
  • Re-ordering worries can be handled by re-ordering the input language sentences in a pre-processing before feeding it to the SMT system. This re-ordering can be done either by handcrafted rules or learnt from data. This could be shallow like POS tag based re-ordering rules or full fledged parsed based. 

Similarly, some work may be done on the post processing side: 

  • If the output language is morphologically complex, then the morphological generation can take place in the post processing step after SMT. This assumes that the SMT system has generated enough information to be able to generate output morphology.
  • Alternatively, in order to ensure grammaticallity of the output sentences, we can do re-ranking of the candidate translations on the output side based on syntactic features like agreement and parse correctness. Note that a distinction has been made between correctness of syntactic parse quality as defined for parsing and as required for MT systems. 

The problem with such pre-processing and post-processing components is that these are themselves prone to error. The system does not handle all the errors in all components in an integrated framework, and necessitates the use of hard decision boundaries. A probabilistic approach which incorporates all these pre- and post-processing components would make a cleaner and more elegant approach. That is the motivation behind the factored translation model. In this model, the factors are basically annotations on the input and output words (e.g. morphology, POS factors).  Translation and generation functions are defined on the factors, and these are integrated using a log linear model. This provides the best way to test a diverse set of features in a structured way. Of course, the size of the phrase translation table will now grow, but this can be handled by using pre-compiled data structured. Decoding could also blow up, but pruning can be used to cut the search space.

Language Divergence between English and Hindi

Comparing two languages is interesting, especially for an application for machine translation. Languages exhibit so many differences, it mind-boggling to realize that we navigate between languages with ease. This paper, 'Interlingua-based English–Hindi Machine Translation and Language Divergence', summarizes the major differences between Hindi and English.

I have tried to tabulate the observations in the paper below, to make a handy reference:


Factor English Hindi



Word Order Subject-Verb-Object Subject-Object-Verb

Ram ate the mango राम ने आम खाया



Modifiers Post modifier Premodifier

The Prime Minister of India भारत का प्रधान मंत्री

play well अच्छे से खेलेंगे 



X-positions Prepositions Postpositions

of India भारत का 

Overloading

John ate rice with curd

John ate rice with a spoon



Compound Verbs not prevelant very common



Conjunct Verbs not prevelant very common


वह गाने लगे 


रुक जाओ 



Respect No special words Words indicating respect


आप, हम 



Person
Uses 2nd person for 3rd person

He obtained his degree आपने  अम्रीका से डिग्री प्राप्त की 



Gender Masculine, feminine, neuter Masculine, feminine



Gender specific possesive pronouns English has them Hindi lacks them

he, she वह



Morphology Poor Rich



Null subject divergence
Subject dropped in certain conditions

There was a king एक राजा था

I am going जा रहा हूँ 



Pleonastic divergence
Pleonastic dropped

It is raining बारिश हो रही है 



Conflational divergence
no appropriate word

Brutus stabbed Caesar ब्रूटस  ने सीसर को छुरे से मारा 



Categorical divergence
change in POS category

They are competing वे मुकाबला कर रहे है



Head swapping
Head and modifier are exchanged

The play is on खेल चल रहा है

Wednesday, September 21, 2011

Aligning Sentences to build a parallel corpus

This is a really old paper, from Gale & Church, on building a sentence aligned parallel corpus from a misaligned corpus. A dynamic programming formulation with a novel distance measure is used for alignment of the sentences. For a method as naive as this, the reported results are impressive on the Hansards corpus. Of course, the input corpus is paragraph aligned. 

The basic premise is simple: Sentences containing less number of characters in one language contain less characters in the other language, and correspondingly for for longer sentence. Based on this idea, the distance between 2 sentences is defined by a  random variable X: the number of charters in language L2 per character or language L1. 

I tried to see the behavior of this variable for the English-Hindi language pair. On a 14000 sentence parallel corpus, here are the results: 

mean(X) : 0.99, i.e. almost one Hindi character for an English character, which is in agreement with the paper's claims. Interesting thing is that if the whitespaces are not considered, the mean drops to 0.96. 
variance(X): 0.01979136 - very low, so the mean is very reliable. A linear fit can't get better than this: 



NLTK provides an implementation of the Gale-Church alignment algorithm. I tried running it on an absolutely parallel corpus, but the algorithm ends up misaligning the sentences. Reducing mean(X) to 0.9 also did not help. Wonder what's going on?