Friday, November 28, 2014

Hello Hadoop: Welcoming Parallel Processing




Even the largest computers struggle with complex problems that have a lot of variables and large data sets. Imagine if one person had to sort through 26,000 boxes of large balls containing sets of 1,000 balls each with one letter of the alphabet: the task would take days. But if you separated the contents o the 1,000 unit boxes into 10 smaller equal boxes and asked 10 separate people to work on these smaller tasks, the job would be completed 10 times faster. This notion of parallel processing is one of the cornerstones of many Big Data projects.

Apache Hadoop (named after the creator Doug Cutting’s child’s toy elephant) is a free programming framework that supports the processing of large data sets in a distributed computing environment. Hadoop is part of the Apache project sponsored by the Apache Software Foundation and although it originally used Java, any programming language can be used to implement many parts of the system.

Hadoop was inspired by Google’s Map-Reduce, a software framework in which an application is broken down into numerous small parts. Any of these parts (also called fragments or blocks) can be run on any computer connected in an organised group called a cluster. Hadoop makes it possible to run applications on thousands of individual computers involving thousands of terabytes of data. Its distributed file system facilitates rapid data transfer rates among nodes and enables the system to continue operating uninterrupted in case of a node failure. This approach lowers the risk of catastrophic system failure, even if a significant number of computers become inoperative.

Saturday, November 15, 2014

Why use Hadoop? Any other solution to Large-Scale Data ?

Hadoop is a large-scale distributed batch processing infrastructure. While it can be used on a single machine, its true power lies in its ability to scale to hundreds or thousands of computers, each with several processor cores. Hadoop is also designed to efficiently distribute large amounts of work across a set of machines.


How large an amount of work? Orders of magnitude larger than many existing systems work with. Hundreds of gigabytes of data constitute the low end of Hadoop-scale. Actually Hadoop is built to process "web-scale" data on the order of hundreds of gigabytes to terabytes or petabytes. At this scale, it is likely that the input data set will not even fit on a single computer's hard drive, much less in memory. So Hadoop includes a distributed file system which breaks up input data and sends fractions of the original data to several machines in your cluster to hold. This results in the problem being processed in parallel using all of the machines in the cluster and computes output results as efficiently as possible.


Challenges at Large Scale



Performing large-scale computation is difficult. To work with this volume of data requires distributing parts of the problem to multiple machines to handle in parallel. Whenever multiple machines are used in cooperation with one another, the probability of failures rises. In a single-machine environment, failure is not something that program designers explicitly worry about very often: if the machine has crashed, then there is no way for the program to recover anyway.


In a distributed environment, however, partial failures are an expected and common occurrence. Networks can experience partial or total failure if switches and routers break down. Data may not arrive at a particular point in time due to unexpected network congestion. Individual compute nodes may overheat, crash, experience hard drive failures, or run out of memory or disk space. Data may be corrupted, or maliciously or improperly transmitted. Multiple implementations or versions of client software may speak slightly different protocols from one another. Clocks may become desynchronized, lock files may not be released, parties involved in distributed atomic transactions may lose their network connections part-way through, etc. In each of these cases, the rest of the distributed system should be able to recover from the component failure or transient error condition and continue to make progress. Of course, actually providing such resilience is a major software engineering challenge.


Different distributed systems specifically address certain modes of failure, while worrying less about others. Hadoop provides no security model, nor safeguards against maliciously inserted data. For example, it cannot detect a man-in-the-middle attack between nodes. On the other hand, it is designed to handle hardware failure and data congestion issues very robustly. Other distributed systems make different trade-offs, as they intend to be used for problems with other requirements (e.g., high security).


In addition to worrying about these sorts of bugs and challenges, there is also the fact that the compute hardware has finite resources available to it. The major resources include:



* Processor time
* Memory
* Hard drive space
* Network bandwidth


Individual machines typically only have a few gigabytes of memory. If the input data set is several terabytes, then this would require a thousand or more machines to hold it in RAM -- and even then, no single machine would be able to process or address all of the data.


Hard drives are much larger; a single machine can now hold multiple terabytes of information on its hard drives. But intermediate data sets generated while performing a large-scale computation can easily fill up several times more space than what the original input data set had occupied. During this process, some of the hard drives employed by the system may become full, and the distributed system may need to route this data to other nodes which can store the overflow.


Finally, bandwidth is a scarce resource even on an internal network. While a set of nodes directly connected by a gigabit Ethernet may generally experience high throughput between them, if all of the machines were transmitting multi-gigabyte data sets, they can easily saturate the switch's bandwidth capacity. Additionally if the machines are spread across multiple racks, the bandwidth available for the data transfer would be much less. Furthermore RPC requests and other data transfer requests using this channel may be delayed or dropped.


To be successful, a large-scale distributed system must be able to manage the above mentioned resources efficiently. Furthermore, it must allocate some of these resources toward maintaining the system as a whole, while devoting as much time as possible to the actual core computation.


Synchronization between multiple machines remains the biggest challenge in distributed system design. If nodes in a distributed system can explicitly communicate with one another, then application designers must be cognizant of risks associated with such communication patterns. It becomes very easy to generate more remote procedure calls (RPCs) than the system can satisfy! Performing multi-party data exchanges is also prone to deadlock or race conditions. Finally, the ability to continue computation in the face of failures becomes more challenging. For example, if 100 nodes are present in a system and one of them crashes, the other 99 nodes should be able to continue the computation, ideally with only a small penalty proportionate to the loss of 1% of the computing power. Of course, this will require re-computing any work lost on the unavailable node. Furthermore, if a complex communication network is overlaid on the distributed infrastructure, then determining how best to restart the lost computation and propagating this information about the change in network topology may be non trivial to implement.

Friday, November 14, 2014

Pig: the Silent Hero


SQL on Hadoop has been extensively covered in the media in the last year. Pig, being a well-established technology, has been largely overlooked though Pig as a Service was a noteworthy development. Considering Hadoop as a data platform though requires Pig and an understanding why and how it is important.  Data users are generally trained in using SQL, a declarative language, to query for data for reporting, analytic and ad-hoc explorations. SQL does not describe how the data is processed; it is more declarative and appeals to a lot of data users. ETL(Extract, Transform and Load)  processes, which are developed by data programmers, benefit and sometimes even require the ability to detail the data transformation steps. At times ETL programmers like a procedural language as opposed to a declarative language. Pig’s programming language, Pig Latin, is procedural and gives programmers control over every step of the processing.  Business users and programmers work on the same data set yet usually focus on different stages. The programmers commonly work on the whole ETL pipeline, i.e. they are responsible to clean and extract the raw data, transform it and load it into third party systems. Business users either access data on third party systems or access the extracted and transformed data for analysis and aggregation. The requirement of diverse tooling is therefore important as the interaction patterns with the same data set are divers.  Importantly, complex ETL workflows need management, extensibility, and test-ability to ensure stable and reliable data processing. Pig provides strong support on all aspects. Pig jobs can be scheduled and managed with workflow tools like Oozie to build and orchestrate large scale, graph-like data pipelines.  Pig achieves extensibility with UDFs (User Defined Function), which let programmers add functions written in one of many programming languages. The benefit of this model is that any kind of special functionality can be injected and that Pig and Hadoop manage the distribution and parallel execution of the function on potentially huge data sets in an efficient manner. This allows the programmers to focus on adding and solving specific domain problems, e.g. like rectifying specific data set anomalies or converting data formats, without worrying about the complexity of distributed computing.  Reliable data pipelines require testing before deployment in production to ensure correctness of the numerous data transformation and combination steps. Pig has features supporting easy and testable development of data pipelines. Pig supports unit tests, an interactive shell, and the option to run in a local mode, which allows it to execute programs in a fashion not requiring a Hadoop cluster. Programmers can use these to test their Pig programs in detail with test data sets before they ever enter production and also help them try out ideas quickly and inexpensively, which is essential for fast development cycles.  None of these features are particularly glamorous yet they are important to evaluate Hadoop and data processing with it. The choice of leveraging Pig for a big data project can easily make the difference between success and failure.  

Wednesday, November 12, 2014

Hadoop Distributed File System(HDFS)


**** BLOCKS ****
All filesystem has a block size, a block size refers to the minimum amount of data that it can read or write. Filesystem blocks are typically a few kilobytes of normally 512 bytes.
Whereas in hadoop filesystem are having a much larger size of blocks i.e 64MB by default.

Now a question arises why the block system of hadoop are of big size ?

To start this - HDFS is meant to handle large files. If you have small files, smaller block sizes are better. If you have large files, larger block sizes are better.
For large files, lets say you have a 1000MB file. With a 4k block size, you'd have to make 256,000 requests to get that file (1 request per block). In HDFS, those requests go across a network and come with a lot of overhead. Each request has to be processed by the Name Node to figure out where that block can be found. That's a lot of traffic! isn't it . If you use 64Mb blocks, the number of requests goes down to 16, greatly reducing the cost of overhead and load on the Name Node.

**** Namenodes and Datanodes ****
Hadoop cluster has two types of nodes which are operating in a master-slave pattern.
1) A Namenode (Master)
2) 'N' Datanodes (Slave)
* N is the number of machines.

The Namenode contains the filesystem namespace. It has the filesystem tree information and the metadata for all files and directories within the tree.This information is stored on the regular interval in the local disk in two formats :
1) Namespace image
2) Edit logs

NOTE : A Namenode also knows the datanodes on which the blocks for a given file is been stored. Namenode does not store the block location on the regular basis as the location of block is always rewriten when ever the system started.

A client always access the files with communicating with Namenode and Datanodes.

Datanode stores and retrieve the blocks when ever they asked to do by client or by Namenode. Dataode also report back to the Namenode about the information of list of block they are storing.
Without the namenode the filesystem or we can say the datanodes or in other words the blocks which contains our data cannot be accessed. This is the single point of failure till hadoop 1.x, but from hadoop 2.x we have different architecture of namenode configuration which contains the secondary namenode concepts.

The secondary Namenode also has serious bottleneck, which could not solve the Single Point of Failure of hadoop cluster.

In the upcoming post, I will explain more about the recent work about this issue.


Sunday, November 9, 2014

Let's start with Map-Reduce

As this is my first post regarding the Map-Reduce Programming, I will be trying to focus on the important points in a Map-Reduce framework.


Map-Reduce works by breaking the processing into two phases. 
1) The Map phase
2) The Reduce phase
Each phase (i.e map and reduce phase) has key-value pair as input and output.
The type of key value pair is been decided by a programmer as per the requirement.
The programmer also specify two function in each phase. The map function and reduce function respectively.
The map function is just the data preparation phase, setting up the data in such a way that the reducer function can do its work on it.
Note : The map function is a good place to drop the bad records.
The output from the map function is processed by the Map-Reduce framework before being sent to the reduce function.
Note : This processing Sorts and groups the key-value pairs by key.
The Mapper class is a generic type, with four formal type parameters that specifies the input key, input value, output key, output value types of the map function.
The map() method also provides an instance of Context to write the output.
In the reduce function again their are four formal type parameters are used to specify the input and output types.
Note : The input types of the reduce function must be of same match as of mapper output function.
Job - A job object forms the specification of the job and gives you control over how the job is run. When we run the job on Hadoop cluster we will package the code into jar file (which Hadoop will distribute around the cluster). Rather then explicitly specify the name of the jar file we can pass the class in the job's setJarByClass() method, which Hadoop will use to locate the relevant JAR file by looking for the JAR file containing this class.
In Job object, we specify the input and output paths. An input path is specified by calling the static addInputPath() method on FileInputFormat , and it can be a single file, a directory, or a file pattern.
Note : addInputPath() can be called more than once to use input from multiple paths.
Note : MultipleInputs class supports Map-Reduce jobs that have multiple input paths with a different InputFormat and Mapper for each path.
The output path (of which there is only one) is specified by the static setOutput Path() method on FileOutputFormat . It specifies a directory where the output files from the reducer functions are written. The directory shouldn’t exist before running the job because Hadoop will complain and not run the job.

Tuesday, November 4, 2014

Hashtable : Most useful java class to modify the meta-data management of Hadoop

Hashtables are an extremely useful mechanism for storing data. Hashtables work by mapping a key to a value, which is stored in an in-memory data structure. Rather than searching through all elements of the hashtable for a matching key, a hashing function analyses a key, and returns an index number. This index matches a stored value, and the data is then accessed. This is an extremely efficient data structure, and one all programmers should remember.
Hashtables are supported by Java, in the form of the java.util.Hashtable class. Hashtables accept as keys and values any Java object. You can use a String, for example, as a key, or perhaps a number such as an Integer. However, you can't use a primitive data type, so you'll need to instead use Char, Integer, Long, etc.
   // Use an Integer as a wrapper for an int
   Integer integer = new Integer ( i );
   hash.put( integer, data);
Data is placed into a hashtable through the put method, and can be accessed using the get method. It's important to know the key that maps to a value, otherwise its difficult to get the data back. If you want to process all the elements in a hashtable, you can always ask for an Enumeration of the hashtable's keys. The get method returns an object, which can then be cast back to the original object type.
   // Get all values with an enumeration of the keys
   for (Enumeration e = hash.keys(); e.hasMoreElements();)
   {
       String str = (String) hash.get( e.nextElement() );
       System.out.println (str);
   }
To demonstrate hashtables, I've written a little demo that adds one hundred strings to a hashtable. Each string is indexed by an Integer, which wraps the int primitive data type.  Individual elements can be returned, or the entire list can be displayed. Note that hashtables don't store keys sequentially, so there is no ordering to the list.

import java.util.*;

public class hash {
  public static void main (String args[]) throws Exception {
    // Start with ten, expand by ten when limit reached
    Hashtable hash = new Hashtable(10,10);

    for (int i = 0; i <= 100; i++)
    {
 Integer integer = new Integer ( i );
 hash.put( integer, "Number : " + i);
    }

    // Get value out again
    System.out.println (hash.get(new Integer(5)));
    // Get value out again
    System.out.println (hash.get(new Integer(21)));

    System.in.read();

    // Get all values
    for (Enumeration e = hash.keys(); e.hasMoreElements();)
    {
 System.out.println (hash.get(e.nextElement()));
    }
 }
}

Thursday, September 4, 2014

Benchmark Testing of Hadoop Cluster with TestDFSIO

This blogpost will help the newbie of Hadoop to learn about the performance measurement of Hadoop cluster.
In this article I will give the details of an important benchmarking tools that is included in the Apache Hadoop distribution. Namely, we look at the benchmarks TestDFSIO. TestDFSIO is one of best industry standard benchmarks used in recent days.


Prerequisites

Before we start few things should be installed/configured in your system.  One needs to configure Hadoop cluster on his system. For that download Hadoop from here.
Configure either of these two.
i) Single user Mode.
ii) Multi Cluster Mode.

TestDFSIO :

The TestDFSIO benchmark is a read and write test for HDFS. It is helpful for tasks such as stress testing HDFS, to discover performance bottlenecks in your network, to shake out the hardware, OS and Hadoop setup of your cluster machines (particularly the NameNode and the DataNodes) and to give you a first impression of how fast your cluster is in terms of I/O.

A source code of the of the documentation, can be found here.

Run write tests before read tests :

Always perform the write test in HDFS before the read test.
The read test of TestDFSIO does not generate its own input files. For this reason, it is a convenient  to run a write test   and then follow-up with reading the same data.


Run a write test :

The command to run a write test that generates 10000 output files each  of 1 GB is:

$ hadoop jar hadoop-*test*.jar TestDFSIO -write -nrFiles 10000 -fileSize 1000

Run a read test :

The command will read the generated 10000 output files each of size 1GB is:

$ hadoop jar hadoop-*test*.jar TestDFSIO -read -nrFiles 10000 -fileSize 1000


Cleaning up and remove test data :

$ hadoop jar hadoop-*test*.jar TestDFSIO -clean

This will clear the data of the directory /benchmarks/TestDFSIO on HDFS

TestDFSIO results :


----- TestDFSIO ----- : write
           Date & time: Fri Apr 08 2011
       Number of files: 10000
Total MBytes processed: 10000000
     Throughput mb/sec: 4.989
Average IO rate mb/sec: 5.185
 IO rate std deviation: 0.960
    Test exec time sec: 1813.53

----- TestDFSIO ----- : read
           Date & time: Fri Apr 08 2011
       Number of files: 10000
Total MBytes processed: 10000000
     Throughput mb/sec: 11.349
Average IO rate mb/sec: 22.341
 IO rate std deviation: 119.231
    Test exec time sec: 1144.842


 Here, the most notable metrics are Throughput mb/sec and Average IO rate mb/sec. Both of them are based on the file size written (or read) by the individual map tasks and the elapsed time to do so.

  I will come up with more article on Hadoop. Thanks 


Sunday, August 31, 2014

Going through the Cloud

Various higher education institutions face increasing challenges due to shrinking revenues, budget restrictions and limited funds for R & D. Therefore it leads to a serious issue, regarding the stipends of research scholars.
Some colleges are looking toward business transformation strategies by organizing workshops, conferences and this brings up an ideological shift to enterprise culture to address the people and process sides of that paradigm. On the technology side, one of the most significant potential avenues for transformation is the adoption of Cloud Computing to reduce costs, boost performance and productivity and increase revenue. But unlike other countries, India is still lacking a lot for set the infrastructure of clouds in the institutions.

Basically the complete transformation has to address all three elements; but I will focus here on the technology aspect. Cloud computing is an effective and efficient technology that can provide education and prepare students as per the latest requirements of the job market, everything at a low cost and with no reduction of quality in the scope of education.

Most of the educational institutions of India host their IT services on premises. Shifting away from this model to cloud, be it Iaas or Paas, would allow them to focus on their core business— education—instead of having to allocate resources for things such as technical support, storage, help desks, e-labs and e-assessments. All the institutions could settle for a “pay as you go” model as to share the computational loads with various colleges, which would prevent the need to host and maintain dedicated infrastructures.

Though cloud computing can bring many benefits to higher education institutions, some issues still need to be considered and addressed, among which are:
  Security and the degree to which an institution is willing to relinquish a certain degree of control over that security
·     The legal issues surrounding data sharing
·     The service provider’s ability to ensure privacy controls and protect data ownership
·     The service provider’s ability to provide adequate and satisfactory support
·     The continuity and availability of data

Cloud computing is an attractive option for cost-conscious institutions seeking to reduce financial and environmental outlays while transforming education. This technological advancement can make knowledge available to entire communities, bridge the social and economic divide and prepare future generations of students and teachers to face the challenges of the global economy.

Sunday, February 16, 2014

Sound Processing in MATLAB

In this semester, I am pursuing the subject, "Speech Processing". A real fascinating subject .  I also liked it in every manner as my B.Tech final year project was actually related to Speech Processing. ( See my previous blog posts) .




 A little try with the sound processing of  Matlab. Hope you will like it.

What is digital sound data?






Getting a pre-recorded sound files (digital sound data)

Click here to access a .wav file



Download any .wav file from here name it anything. In my case , let it be 'road.wav'


Loading Sound files into MATLAB

·        I want to read the digital sound data from the .wav file into an array in our MATLAB workspace.  I can then listen to it, plot it, manipulate, etc.  Use the following command at the MATLAB prompt:

[road,fs]=wavread('road.wav');

·       The array road now contains the stereo sound data and fs is the sampling frequency.  This data is sampled at the same rate as that on a music CD (fs=44,100 samples/second).

·       See the size of road:  
size(road)

·       The left and right channel signals are the two columns of the road array:


 left=road(:,1);
right=road(:,2);


·       Let’s plot the left data versus time.  Note that the plot will look solid because there are so many data points and the screen resolution can’t show them all.  This picture shows you where the signal is strong and weak over time.

time = (1/44100)*length(left);
t=linspace(0,time,length(left));
plot(t,left)xlabel('time (sec)');
ylabel('relative signal strength'

·       Let’s plot a small portion so you can see some details

  
time=(1/44100)*2000;
t=linspace(0,time,2000);
plot(t,left(1:2000))
xlabel('time (sec)');
ylabel('relative signal strength')


 ·       Let’s listen to the data (plug in your headphones).  Click on the speaker icon in the lower right hand corner of your screen to adjust the volume.  Enter these commands below one at a time.  Wait until the sound stops from one command before you enter another sound command!

soundsc(left,fs)       % plays left channel as mono

soundsc(right,fs)    % plays right channel mono (sound nearly the same)

soundsc(road,fs)     % plays stereo 















Sunday, December 15, 2013

Recover Virus Infected Hidden Files from Pen Drive or Flash Drive

For Advanced users:
open cmd
type i: (Your Drive Letter) and enter
type attrib -h -s /s /d  and press enter


A brief explanation for the above mentioned procedure.


 Today my camera's memory card is hit by Trozan Viruses. I tried with the anti-virus softwares to clean it. But still it was the showing the same error. 


 The whole disk space is empty. But in the properties, size was showing as 3 GB approx. I tried 'Show hidden files " options in Folder options, but of no use.


I asked few of my friends about this. Actually its a very common type of bugs you will find with your memory card or any pen drives also. 


I did some research and came to know that these files are hidden by Trojan Viruses within registry values.


The only way , I reckon you change it , is by using Command Prompt.



Open Run. Type cmd and press enter 





A command Prompt will appear, type cd / and hit enter.. This command will let the command applicable to Root directories and sub folders also.





Now, we have to move to the drive, which is your flash drive. In my case, its E drive. So, type E: and press enter. 





Now type  attrib -h -s /s /d  and press enter. { attrib space -h space /s space /d }

After a few seconds you can see your driver letter again this means the process is completed.

You will get all your data back :) 

NOTE :


     This process is applicable not only for Pen Drive or Flash dives like External Hard Disc, but also applicable for your system drives also.

    After applying this command, some of the strange folders you have never seen may appear but don't worry, they are system related files, just hide them if you don't want or try delete if possible.


 If you encounter any problem, mail me at dev.dipayan16@gmail.com

Thank you for visiting ! 


















Sunday, June 24, 2012

First try on HTML 5 . :) :) Awesome experience :)


Here, I have implemented a simple drag and drop in action. J  . I’ll de­fine a drag­gable image that can be dropped into a circle.


The Complete source code ::

<!DOCTYPE html>

<html lang="en">
<head>
    <title>Basic drag and drop example</title>
    <style>

        .drop-div {
            width: 150px;
            height: 150px;
            border: 3px dashed #224163;
            background-color: #AABACC;
            margin-top: 15px;
            border-radius: 80px;
            text-align: center;
        }

        [draggable=true] {
            -khtml-user-drag: element;
            -webkit-user-drag: element;
            -khtml-user-select: none;
            -webkit-user-select: none;
            cursor: pointer;
            margin-top: 50px;
        }
    </style>
    <script>

        function dragStartHandler(event) {
            event.dataTransfer.setData('Text', 'text');
        }
        function dropHandler(event) {
            preventDefaults(event);
            if (!event.target) {
                event.target = event.srcElement
            }
            event.target.appendChild(document.getElementById('draggable-img'));
        }
        function dragOverHandler(event) {
            preventDefaults(event);
        }

        function preventDefaults(event) {
            if (event.preventDefault) {
                event.preventDefault();
            }

            try {
                event.returnValue = false;
            }
            catch (exception) {}
        }
    </script>
</head>

<body>

<img draggable="true"
     ondragstart="dragStartHandler(event);"
     class="draggable-img"
     id="draggable-img"
     src="C:\Users\Dipayan\Desktop\ball.png "/>
<div class="drop-div"
     ondragover="dragOverHandler(event);"
     ondrop="dropHandler(event)"
     id="drop-div"></div>

</body>
</html>


What It Looks Like


















You can then click on the red ball , drag the ball to in­side the circle, and drop it to pro­duce this:












Saturday, May 5, 2012

DSN-Less MSAccess connection

public Connection getConnection() throws Exception {  
 try{
                // Load the driver
                    Class.forName("sun.jdbc.odbc.JdbcOdbcDriver");
                        String filename = "C:/Users/Dipayan/Desktop/dic.mdb";
    String database = "jdbc:odbc:Driver={Microsoft Access Driver (*.mdb)};DBQ=";
    database+=filename.trim() + ";DriverID=22;READONLY=true}";
    Connection conn=DriverManager.getConnection(database,"","");    
                        System.out.println("Connected to the dictionary");  
                     // Create a Statement
                        Statement stmt = conn.createStatement();
                        System.out.println("Statement Created"); 
                        // Create a query String
                        
                        ResultSet rs = stmt.executeQuery("SELECT meaning FROM dictionary where word='"+str+"'");
                        System.out.println("Query Executed"); 
                        while(rs.next())
                        {
                          text = rs.getString(1).toString();
                         // System.out.print(t);
                         
                        } 
                              
                        stmt.close();
                        conn.close();
                        System.out.println("Disconnected from database");
                }
                catch(SQLException e) {
                            e.printStackTrace();
                } catch (ClassNotFoundException e) {
// TODO Auto-generated catch block
e.printStackTrace();
}


Friday, April 13, 2012

Expanding Dictionary Of Acoustic Model



 Today I’m going to tell you how to expand dictionary of acoustic model for Sphinx4. In simple words, This tutorial will tell you how you can add more words in Sphinx’s words database (Dictionary) and let it recognize those words, which are not available in default acoustic models provided by CMU Sphinx. This tutorial is based on “HelloWorld” example provided by CMU Sphinx.


Important Files in this example :
1 ) HelloWorld.java
2) hello.gram
3) helloworld.config.xml
Acoustic Model used in this example :
WSJ_8gau_13dCep_16k_40mel_130Hz_6800Hz.jar
Lets say, We are creating a SR system for ABC National airlines. Everything will go fine and Sphinx will recognize most of the words except the name of cities and states of India.  Now, I will tell you, How to add name of cities and states in dictionary.
PART ONE
Step 1 : Create a txt file “words.txt”, Write all the names of cities and states in it and save.
Step 2 : Open this link : http://www.speech.cs.cmu.edu/tools/lmtool.html
Step 3 : On that page, go to “Sentence corpus file:” section, Browse to “words.txt” file and click “Compile Knowledge Base”.
Step 4 : On next page, Click on “Dictionary” link and save that .DIC file.
PART TWO
Step 1 : Extract WSJ_8gau_13dCep_16k_40mel_130Hz_6800Hz.jar file.
Step 2 : Go to edu\cmu\sphinx\model\acoustic\WSJ_8gau_13dCep_16k_40mel_130Hz_6800Hz\dict folder.
Step 3 : Open “cmudict.0.6d” file in that folder.
Step 4 : Copy data from .DIC file, you have downloaded in PART ONE, paste it in “cmudict.0.6d” file and save.
Step 5 : Zip the extracted hierarchy back as it was and Zip file named should be same as JAR file.
Now, remove “WSJ_8gau_13dCep_16k_40mel_130Hz_6800Hz.jar” file from Project’s CLASSPATH and add “WSJ_8gau_13dCep_16k_40mel_130Hz_6800Hz.zip” instead of it.
That’s it ! We are done.  Now Sphinx will also recognize all name of cities and states that we wrote in “words.txt” file.
Now, FAQ time. I will be posting FAQ and few important notes in comments. 
:)
If you have any quires, Please feel free to ask.
Regards,