diff --git a/Hadoop-Compatible-File-System.md b/Hadoop-Compatible-File-System.md index 12d3e42..32e9043 100644 --- a/Hadoop-Compatible-File-System.md +++ b/Hadoop-Compatible-File-System.md @@ -1,15 +1,31 @@ +HDFS is optimized for large files. The scalability of the single HDFS namenode is limited by the number of files. It is hard for HDFS to store lots of small files. + +SeaweedFS excels on small files. Now it is possible to enable Hadoop jobs to read from and write to SeaweedFS. + +# Build SeaweedFS Hadoop Client Jar +``` +$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/client +$ mvn install +$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/hdfs +$ mvn package +$ ls -al target/seaweedfs-hadoop-client-1.0-SNAPSHOT.jar + +``` + # Test SeaweedFS on Hadoop Suppose you are getting a new Hadoop installation. Here are the minimum steps to get SeaweedFS to run. -You would need to start a weed filer first, download the seaweedfs-hadoop-client-xxx.jar, and do the following: +You would need to start a weed filer first, build the seaweedfs-hadoop-client-xxx.jar, and do the following: ``` $ cd ${HADOOP_HOME} +# create etc/hadoop/mapred-site.xml, just to satisfy hdfs dfs. skip this if the file already exists. $ echo "" > etc/hadoop/mapred-site.xml $ bin/hdfs dfs -Dfs.defaultFS=seaweedfs://localhost:8888 \ -Dfs.seaweedfs.impl=seaweed.hdfs.SeaweedFileSystem \ -libjars ./seaweedfs-hadoop-client-1.0-SNAPSHOT.jar \ -ls / -``` \ No newline at end of file +``` +Both reads and writes are working fine. \ No newline at end of file