Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

TeraSort-Local-Hadoop-MR-Spark

Refer http://sortbenchmark.org/

Three version of program are made.

input file is:

One is local one, where program uses external sort to sort the data and uses im-memory data-stucture.

Hadoop-MR and Spark uses 16 nodes of d2.xlarge on aws to sort the 1 TB data.

For deployment use the script in hadoopsetup repo while follow manual steps for Spark.

About

Sorting 1TB of data using Hadoop Map Reduce, Apache Spark and custom Java solution

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages