The other theory proposes an indigenous development of the Brahmi script, based on the wide differences in how the writing systems work. I tend to favour this theory, though I must admit that my knowledge of this area is limited to reading a few articles and knowing some of the modern day descendants of these scripts. The modern day alphabet of Indian scripts are organized phonetically, and there is little ambiguity phonetically - as opposed to the Roman scripts. The earliest Semitic scripts (Phoenician, Aramaic) and even modern Arabic do not have vowels, whereas the so called "true" alphabets Greek and its modern Latin derivative scripts still have room for ambiguity. Even if there was some use of symbols from the Aramaic scripts, the design seems pretty novel to call it a new style of scripting. Is there an alternative line of evolution of the script? The Indus Valley script is still undecipered - could the Brahmi have evolved from there?
Thursday, August 23, 2012
Origins of the Brahmi Script
The other theory proposes an indigenous development of the Brahmi script, based on the wide differences in how the writing systems work. I tend to favour this theory, though I must admit that my knowledge of this area is limited to reading a few articles and knowing some of the modern day descendants of these scripts. The modern day alphabet of Indian scripts are organized phonetically, and there is little ambiguity phonetically - as opposed to the Roman scripts. The earliest Semitic scripts (Phoenician, Aramaic) and even modern Arabic do not have vowels, whereas the so called "true" alphabets Greek and its modern Latin derivative scripts still have room for ambiguity. Even if there was some use of symbols from the Aramaic scripts, the design seems pretty novel to call it a new style of scripting. Is there an alternative line of evolution of the script? The Indus Valley script is still undecipered - could the Brahmi have evolved from there?
Sunday, February 12, 2012
Indian English
In a newspaper, describing a case of chain-snatching in which criminals shot dead the man who tried to resist and pursue the chain-snatchers, the reporter stated: “The deceased gave chase to the criminals who, however, managed to escape”!
Police notice: “Take care of belongings. You may be theft”
The article is interesting reading too.
http://dailypioneer.com/
Saturday, January 14, 2012
Yet Another Moses Installation Guide
- Language modelling toolkit (SRILM, IRSTLM, etc.)
- GIZA++ package which contains GIZA++ and mkcls
- Moses decoder (version 1.0 and above)
- The primary installation reference is the INSTALL document that ships with the tool.
- Install all pre-requisites mentioned in the SRILM installation guide. On Ubuntu I had to install the following packages: csh, g++-multilib, tcl-dev
- Set the environment variable SRILM to point to the base directory of the install package before building SRILM.
- Following the instruction manual with the SRILM download should be enough once the pre-requisites are installed.
- The problems you may yet face are:
- Problem in identifying the architecture, especially if it a 64-bit machine. To make sure that the install script correctly identifies the architecture, set the variable MACHINE_TYPE in sbin/machine-type.
- Problems with TCL compilation. You may not need the TCL user interfaces at all, so it may just be able ok to disable their compilation. Set the variable NO_TCL = X in the file common/your_architecture_specific_makefile.
- Make sure you have added the $SRILM/bin and $SRILM/bin/$MACHINE_TYPE to the PATH variable
- Note: SRILM 1.7.1 and above are not compatible with Moses
- Ubuntu packages required: libtool make autoconf autotools-dev automake
- The installation is pretty simple, just have to follow the installation guide
- One caveat: Sometimes, it may be required to create a directory named 'm4' manually, if the first step mails
- You get both if you download the giza-pp tool.
- Most straightforward installation. Download and 'make'.
- Copy the binaries - GIZA++, mkcls, snt2cooc.out to a new directory.
- XML RPC Server is required if you want to run a webservice providing translations. If you just want to get Moses running, you can skip this step.
- Install the following packages: libxmlrpc-core-c3 libxmlrpc-core-c3-dev libxmlrpc-c3-dev libxmlrpc-c++4 libxmlrpc-c++4-dev
The C++ Boost library is required for installation of Moses. Boost 1.48 has a serious bug which breaks Moses compilation. Unfornately, some Linux distributions (eg. Ubuntu 12.04) have broken versions of the Boost library.To fix this situation you can:
- For Ubuntu 12.04: Remove boost 1.48 from your distribution and install Boost 1.46 which is available in the distribution. This works most of the time. If not, build Boost from source as described below.
- To install Boost manually and making it work with Moses, follow the instructions in the section titled "Manually Installing Boost" on this page: http://www.statmt.org/moses/?n=Development.GetStarted
- The primary installation reference is the INSTALL document that ships with the tool.
- SRILM or IRSTLM need to be installed before Moses is installed
- Make sure you have installed the packages automake and libtool
- Boost has to be installed
- It is then a matter of just following the instructions. The command to be run is:
- /usr/bin/bjam --with-srilm=
--with-xmlrpc-c= --with-boost= - If the xml RPC is installed in /usr/bin, then the parameter would simply be '/usr'
- --with-boost is required only when Boost is installed in a non-standard directory. The path should contain both lib/lib64 and include directories
Alternative ways of installation Moses
http://www.maketecheasier.com/import-export-ova-files-in-virtualbox/
Friday, September 23, 2011
Incorporating Linguistic Information into SMT Models
- Name Transliteration and Number script conversions
- Morphology changes - inflections, compounding, segmentation - these problems if not handled lead to data sparsity problems
- Syntanctic phenomena like constituent structure, attachment, head-modifier re-orderings. Vanilla SMT is designed to handle local re-orderings but long range dependencies are not handled well.
- Transliteration and back transliterations models need to be incorporated. An important problem is to identify the named entities in the first place.
- Splitting words for a morphology rich input language. Compounding and segmentation can be handled similarly.
- Re-ordering worries can be handled by re-ordering the input language sentences in a pre-processing before feeding it to the SMT system. This re-ordering can be done either by handcrafted rules or learnt from data. This could be shallow like POS tag based re-ordering rules or full fledged parsed based.
- If the output language is morphologically complex, then the morphological generation can take place in the post processing step after SMT. This assumes that the SMT system has generated enough information to be able to generate output morphology.
- Alternatively, in order to ensure grammaticallity of the output sentences, we can do re-ranking of the candidate translations on the output side based on syntactic features like agreement and parse correctness. Note that a distinction has been made between correctness of syntactic parse quality as defined for parsing and as required for MT systems.
Language Divergence between English and Hindi
I have tried to tabulate the observations in the paper below, to make a handy reference:
| Factor | English | Hindi |
| Word Order | Subject-Verb-Object | Subject-Object-Verb |
| Ram ate the mango | राम ने आम खाया | |
| Modifiers | Post modifier | Premodifier |
| The Prime Minister of India | भारत का प्रधान मंत्री | |
| play well | अच्छे से खेलेंगे | |
| X-positions | Prepositions | Postpositions |
| of India | भारत का | |
| Overloading | ||
| John ate rice with curd | ||
| John ate rice with a spoon | ||
| Compound Verbs | not prevelant | very common |
| Conjunct Verbs | not prevelant | very common |
| वह गाने लगे | ||
| रुक जाओ | ||
| Respect | No special words | Words indicating respect |
| आप, हम | ||
| Person | Uses 2nd person for 3rd person | |
| He obtained his degree | आपने अम्रीका से डिग्री प्राप्त की | |
| Gender | Masculine, feminine, neuter | Masculine, feminine |
| Gender specific possesive pronouns | English has them | Hindi lacks them |
| he, she | वह | |
| Morphology | Poor | Rich |
| Null subject divergence | Subject dropped in certain conditions | |
| There was a king | एक राजा था | |
| I am going | जा रहा हूँ | |
| Pleonastic divergence | Pleonastic dropped | |
| It is raining | बारिश हो रही है | |
| Conflational divergence | no appropriate word | |
| Brutus stabbed Caesar | ब्रूटस ने सीसर को छुरे से मारा | |
| Categorical divergence | change in POS category | |
| They are competing | वे मुकाबला कर रहे है | |
| Head swapping | Head and modifier are exchanged | |
| The play is on | खेल चल रहा है |
Wednesday, September 21, 2011
Aligning Sentences to build a parallel corpus
Wednesday, August 31, 2011
Watson - The Quiz Champion
You must have heard of IBM's Watson system. It is, of course, the computer that won the Jeopardy competition against the show's previous champions. Jeopardy is a popular quiz show in which the competitors are provided clues and have to give questions that satisfy these clues. For example, a clue like 'This computer beat the reigning world chess champion' would elicit a question 'Who is Deep Blue?'. As you can see, the questions given by the competitors are easy questions of the nature 'What is', 'Who is', so the Jeopary question answer format can be considered like any other quiz show. The clues however are complex covering a wide array of topics, and could include puns, puzzles, and maths. The competitors also place bets on each questions. Competing at 'Jeopardy' thus requires the right combination of 'natural language understanding, broad knowledge, confidence and strategy'.
Watson's victory thus represents a major milestone for natural language processing, and particularly the sub-area known as 'Question-Answering'. Question-Answering systems have great practical use for building expert systems, customer support system, decision making tools, enterprise search systems.
Watch Watson's winning performance here:
This paper, Building Watson: An Overview of the DeepQA project, from IBM provides an overview of Watson and the DeepQA architecture that underlies it. The DeepQA architecture defines a framework for development of QA systems in an extensible and modular method, allowing different components to be customized, and to build robust QA systems that can be ported across domains. Figure 1 shows a high level diagram of the Watson's major components, and how queries are routed through it.
- Query Analysis: This is the first stage, where the input clue is analyzed to determine the question category (puzzle, pune, mathematical, numeric, logical, etc.) and the answer type (person, location, organization, etc.). Complex clues are also decomposed into simpler clues.
- Hypothesis Generation: Watson has at its disposal many sources of information like encyclopedias, books, lists of things like people, countries, etc. Watson does not attempt to get the correct answer straightaway. Instead, it first focusses on generating as many possible candidate answers, called 'hypotheses'. This is to ensure that good answers are not missed in the pursuit of the perfect answer. The attempt is to increase recall at this stage.
- Soft Filtering: Watson may generate hundreds and thousands of hypotheses, which then have to be analyzed in detail to find the correct answer. To limit this deep analysis to only the most relevant answers, Watson filters out the bad candidates by employing a few techniques like mismatch between the expected and candidate answer type.
- Hypothesis and Evidence scoring: Now Watson does a deep analysis of the candidate answers by employing sophisticated linguistic and statistical techniques, and looks to gather evidence for each hypothesis. This is one of the most critical parts of Watson since the evidence collected will determine how good the answer is and how confident Watson can be about it.
- Merging and Ranking: Once the evidence is collected, the confidence scores are generated for each candidate and candidates ranked. Now, looking at the answer's confidence level Watson decides if it should answer the question or not.
Figure 1: DeepQA Architecture (Source: The IBM paper)
The flexibility in the DeepQA architecture is achieved through the use of the UIMA text analysis framework. At one point in the trials, Watson was taking about two hours to generate an answer. The answer was to parallelize Watson with UIMA-AS and this got the response time down to the quiz show's average of 2 to 5 seconds. The improvement in accuracy is even more startling. When the IBM team stared working on Watson, the difference between the show's participants and early prototypes of Watson was huge. Figure 2 depicts the evolution in Watson's performance. It started from the baseline where the precision and recall were nowhere near the cloud of points corresponding to actual human competitors, but gradually reached human level performance.
Figure 2: Watson's accuracy over time (Source: The IBM paper)
What enabled Watson to reach this level of performance? Many of the underlying analysis algorithms aren't new, but have been around in the research community for a long time. More than groundbreaking original research, it is pragmatic engineering that lies at the core of Watson's success and the following are the salient contributory factors:
- Building an end-to-end system: Very early, the team build a baseline end-to-end system and then kept iterating and improving the system. They defined end-to-end evaluation metrics which captured the performance of the system as a whole, and not focusing only on the individual component accuracies at the initial stages. This helped make the correct trade-offs.
- Pervasive Confidence estimation: Every component in Watson gives a confidence estimate along with its response. This is critical since these confidence scores can be aggregated to get the final confidence on the answers and allows easy integration of components of varying accuracy. The rule is that no component is assumed to be perfect, but each makes available its confidence estimate of the answers.
- Many experts: There may be competing algorithms to do the same task. Rather than using the best, the system uses multiple algorithms so as to get diverse results and evidence. The confidence estimates help to blend the diverse results.
- Integrate shallow and deep knowledge: Balance the use of strict semantics and shallow semantics, leveraging many loosely formed ontologies.
- Massive parallelism: As mentioned, exploiting massive parallelism allows looking through a large number of hypotheses.
