From 5c928ffff7f3f7ca45be4875009431d396277aad Mon Sep 17 00:00:00 2001 From: Chris Lu Date: Tue, 4 Dec 2018 21:56:33 -0800 Subject: [PATCH] Updated Hadoop Compatible File System (markdown) --- Hadoop-Compatible-File-System.md | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/Hadoop-Compatible-File-System.md b/Hadoop-Compatible-File-System.md index 12d3e42..32e9043 100644 --- a/Hadoop-Compatible-File-System.md +++ b/Hadoop-Compatible-File-System.md @@ -1,15 +1,31 @@ +HDFS is optimized for large files. The scalability of the single HDFS namenode is limited by the number of files. It is hard for HDFS to store lots of small files. + +SeaweedFS excels on small files. Now it is possible to enable Hadoop jobs to read from and write to SeaweedFS. + +# Build SeaweedFS Hadoop Client Jar +``` +$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/client +$ mvn install +$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/hdfs +$ mvn package +$ ls -al target/seaweedfs-hadoop-client-1.0-SNAPSHOT.jar + +``` + # Test SeaweedFS on Hadoop Suppose you are getting a new Hadoop installation. Here are the minimum steps to get SeaweedFS to run. -You would need to start a weed filer first, download the seaweedfs-hadoop-client-xxx.jar, and do the following: +You would need to start a weed filer first, build the seaweedfs-hadoop-client-xxx.jar, and do the following: ``` $ cd ${HADOOP_HOME} +# create etc/hadoop/mapred-site.xml, just to satisfy hdfs dfs. skip this if the file already exists. $ echo "" > etc/hadoop/mapred-site.xml $ bin/hdfs dfs -Dfs.defaultFS=seaweedfs://localhost:8888 \ -Dfs.seaweedfs.impl=seaweed.hdfs.SeaweedFileSystem \ -libjars ./seaweedfs-hadoop-client-1.0-SNAPSHOT.jar \ -ls / -``` \ No newline at end of file +``` +Both reads and writes are working fine. \ No newline at end of file