Saturday, October 1, 2016

Big data analytics

Big data analytics is the process of examining large data sets to uncover hidden patterns, unknown correlations, market trends, customer preferences and other useful business information.With today’s technology, it’s possible to analyze your data and get answers from it almost immediately – an effort that’s slower and less efficient with more traditional business intelligence solutions.

Why is big data analytics important?

Big data analytics helps organizations harness their data and use it to identify new opportunities. That, in turn, leads to smarter business moves, more efficient operations, higher profits and happier customers. IIA Director of Research Tom Davenport interviewed more than 50 businesses to understand how they used big data. He found that they got value in the following ways:
  1. Cost reduction. Big data technologies such as Hadoop and cloud-based analytics bring significant cost advantages when it comes to storing large amounts of data – plus they can identify more efficient ways of doing business.
  2. Faster, better decision making. With the speed of Hadoop and in-memory analytics, combined with the ability to analyze new sources of data, businesses are able to analyze information immediately – and make decisions based on what they’ve learned.
  3. New products and services. With the ability to gauge customer needs and satisfaction through analytics comes the power to give customers what they want. Davenport points out that with big data analytics, more companies are creating new products to meet customers’ needs.

The most important research topics in the Big Data field

Here are the major research fields  where BigData is involved

1) Improving Data analytic techniques- Gather all datas,filter them out on certain constraints and use them to take confident decisions.

2) Natural Language Processing methods - Use NL-processing techniques on Big Data to find out the current sentimental trend and it can be used on business,politics,finance ...etc

3) BigData tools and deployment platforms -Conventional tools are inefficient to handle Bigdata, Lots of research is needed in these fields.

4) Better datamining techniques-Data mining is the method to grab data from various platforms.Improved distributed crawling techniques and algorithms are need for scrape data from multiple platforms.

5) Algorithms for Data visualization-In order to visualize the required information from a pool of random  data, powerful algorithms are crucial for accurate result.

6) Lots more...


Here are the research topics that might be relevant to healthcare and bigdata:
  1. Sentiment analysis
  2. Live drug response analysis
  3. Heterogeneous information integration at large volume of data
  4. Security and privacy issues related to Healthcare infomation exchange.
  5. Metadata management
  6. Information retrieval tools for efficient data searching.
  7. Fraud detection
 

Tuesday, September 27, 2016

CS 6712 GRID AND CLOUD COMPUTING LAB MANUAL

CS 6712 GRID AND CLOUD COMPUTING LAB MANUAL

CS 6712 GRID AND CLOUD COMPUTING LAB MANUAL

Anna University 2013 Regulation


CS 6712 Grid and Cloud Computing Lab Manual can be downloaded from the below Links


Lab Manual 1

https://drive.google.com/open?id=0ByI3h-WZRk-ndS1ZbUpmUkM2OVU
   
Lab Manual 2

https://drive.google.com/open?id=0ByI3h-WZRk-nTlQwSk4zb04yRnM

Lab Manual 3

https://drive.google.com/open?id=0ByI3h-WZRk-nX3N2ZFNJVDdrWEU

 Lab Manual 4

 

Lab Manual 5

https://drive.google.com/open?id=0ByI3h-WZRk-nSXB2YnlYaUh2VVU 
  

Saturday, September 24, 2016

CS 6712 GRID lab Prerequisites



GRID LAB Exercises - Prerequisites

1. Install Java
2. Install GCC
 $ sudo add-apt-repository ppa:ubuntu-toolchain-r/test
 $ sudo apt-get update
 $ sudo apt-get install gcc-4.9  
3. Installing Perl
 $ sudo apt-get install perl

4. Installing Grid Essential

 $ sudo dpkg -i globus-toolkit-repo_latest_all.deb

if error comes  ===> $ sudo apt-get update

 $ sudo apt-get install globus-data-management-client
 $ sudo  apt-get install globus-gridftp
 $ sudo  apt-get install globus-gram5
 $ sudo  apt-get install globus-gsi
 $ sudo  apt-get install globus-data-management-server
 
 $ sudo  apt-get install globus-data-management-sdk
 $ sudo  apt-get install globus-resource-management-server
 $ sudo  apt-get install globus-resource-management-client
 $ sudo  apt-get install globus-resource-management-sdk
 $ sudo apt-get install myproxy
 $ sudo apt-get install gsi-openssh
 $ sudo apt-get install globus-gridftp globus-gram5 globus-gsi myproxy myproxy-server myproxy-admin

5. Installing Eclipse or Netbeans
 $ chmod +x netbeans-8.1-javaee-linux.sh
 $./netbeans-8.1-javaee-linux.sh

6. Installing Apache Axis

In eclipse --> windows--> preference -> add the axis file---> apply --> ok

Download tomcat and install and start the service
in terminal go to tomcat folder $ bin/startup.sh 
in webbrowser --> localhost:8080

Thursday, September 22, 2016

CS6703 GRID AND CLOUD COMPUTING SYLLABUS



CS6703                       GRID AND CLOUD COMPUTING                      L T P C               3 0 0 3

OBJECTIVES:

The student should be made to:
·         Understand how Grid computing helps in solving large scale scientific problems.
·         Gain knowledge on the concept of virtualization that is fundamental to cloud computing.
·         Learn how to program the grid and the cloud.
·         Understand the security issues in the grid and the cloud environment.

UNIT I                                                                   INTRODUCTION                                             9
Evolution of Distributed computing: Scalable computing over the Internet – Technologies for network based systems – clusters of cooperative computers – Grid computing Infrastructures – cloud computing – service oriented architecture – Introduction to Grid Architecture and standards –
Elements of Grid – Overview of Grid Architecture.

UNIT II                                                                 GRID SERVICES                                               9
Introduction to Open Grid Services Architecture (OGSA) – Motivation – Functionality Requirements – Practical & Detailed view of OGSA/OGSI – Data intensive grid service models – OGSA services.

UNIT III                                                            VIRTUALIZATION                                               9
Cloud deployment models: public, private, hybrid, community – Categories of cloud computing: Everything as a service: Infrastructure, platform, software – Pros and Cons of cloud computing – Implementation levels of virtualization – virtualization structure – virtualization of CPU, Memory and I/O devices – virtual clusters and Resource Management – Virtualization for data center automation.

UNIT IV                                                    PROGRAMMING MODEL                                          9
Open source grid middleware packages – Globus Toolkit (GT4) Architecture , Configuration – Usage of Globus – Main components and Programming model – Introduction to Hadoop Framework – Mapreduce, Input splitting, map and reduce functions, specifying input and output parameters,
configuring and running a job – Design of Hadoop file system, HDFS concepts, command line and java interface, dataflow of File read & File write.

UNIT V                                                                    SECURITY                                                       9
Trust models for Grid security environment – Authentication and Authorization methods – Grid security infrastructure – Cloud Infrastructure security: network, host and application level – aspects of data security, provider data and its security, Identity and access management architecture, IAM practices in the cloud, SaaS, PaaS, IaaS availability in the cloud, Key privacy issues in the cloud.
                                                                                                                           TOTAL: 45 PERIODS

OUTCOMES:

At the end of the course, the student should be able to:
 Apply grid computing techniques to solve large scale scientific problems.
 Apply the concept of virtualization.
 Use the grid and cloud tool kits.
 Apply the security models in the grid and the cloud environment.

TEXT BOOK:
1. Kai Hwang, Geoffery C. Fox and Jack J. Dongarra, “Distributed and Cloud Computing: Clusters, Grids, Clouds and the Future of Internet”, First Edition, Morgan Kaufman Publisher, an Imprint of Elsevier, 2012.

REFERENCES:
1. Jason Venner, “Pro Hadoop- Build Scalable, Distributed Applications in the Cloud”, A Press, 2009
2. Tom White, “Hadoop The Definitive Guide”, First Edition. O’Reilly, 2009.
3. Bart Jacob (Editor), “Introduction to Grid Computing”, IBM Red Books, Vervante, 2005
4. Ian Foster, Carl Kesselman, “The Grid: Blueprint for a New Computing Infrastructure”, 2nd Edition, Morgan Kaufmann.
5. Frederic Magoules and Jie Pan, “Introduction to Grid Computing” CRC Press, 2009.
6. Daniel Minoli, “A Networking Approach to Grid Computing”, John Wiley Publication, 2005.
7. Barry Wilkinson, “Grid Computing: Techniques and Applications”, Chapman and Hall, CRC, Taylor and Francis Group, 2010.

CS6712 GRID AND CLOUD COMPUTING LAB SYLLABUS



CS6712          GRID AND CLOUD COMPUTING LABORATORY          L T P C           0 0 3 2

OBJECTIVES:

The student should be made to:
 Be exposed to tool kits for grid and cloud environment.
 Be familiar with developing web services/Applications in grid framework
 Learn to run virtual machines of different configuration.
 Learn to use Hadoop

LIST OF EXPERIMENTS:

GRID COMPUTING LAB:
Use Globus Toolkit or equivalent and do the following:
1. Develop a new Web Service for Calculator.
2. Develop new OGSA-compliant Web Service.
3. Using Apache Axis develop a Grid Service.
4. Develop applications using Java or C/C++ Grid APIs
5. Develop secured applications using basic security mechanisms available in Globus Toolkit.
6. Develop a Grid portal, where user can submit a job and get the result. Implement it with and without GRAM concept.

CLOUD COMPUTING LAB:
Use Eucalyptus or Open Nebula or equivalent to set up the cloud and demonstrate:
1. Find procedure to run the virtual machine of different configuration. Check how many virtual machines can be utilized at particular time.
2. Find procedure to attach virtual block to the virtual machine and check whether it holds the data even after the release of the virtual machine.
3. Install a C compiler in the virtual machine and execute a sample program.
4. Show the virtual machine migration based on the certain condition from one node to the other.
5. Find procedure to install storage controller and interact with it.
6. Find procedure to set up the one node Hadoop cluster.
7. Mount the one node Hadoop cluster using FUSE.
8. Write a program to use the API’s of Hadoop to interact with it.
9. Write a wordcount program to demonstrate the use of Map and Reduce tasks
                                                                                                                       TOTAL: 45 PERIODS

OUTCOMES:
At the end of the course, the student should be able to:
 Use the grid and cloud tool kits.
 Design and implement applications on the Grid.
 Design and Implement applications on the Cloud.

LIST OF EQUIPMENT FOR A BATCH OF 30 STUDENTS:

SOFTWARE:
Globus Toolkit or equivalent Eucalyptus or Open Nebula or equivalent

HARDWARE:
Standalone desktops 30 Nos

TCS-SEARS Mega Drive on 24th September