Updated Hadoop Compatible File System (markdown)

Chris Lu committed 2018-12-04 21:56:33 -08:00
1 parent 120d4752a1
commit 5c928ffff7
1 file changed
+18 -2
+18 -2
@@ -1,15 +1,31 @@
HDFS is optimized for large files. The scalability of the single HDFS namenode is limited by the number of files. It is hard for HDFS to store lots of small files.
SeaweedFS excels on small files. Now it is possible to enable Hadoop jobs to read from and write to SeaweedFS.
# Build SeaweedFS Hadoop Client Jar
```
$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/client
$ mvn install
$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/hdfs
$ mvn package
$ ls -al target/seaweedfs-hadoop-client-1.0-SNAPSHOT.jar
```
# Test SeaweedFS on Hadoop
Suppose you are getting a new Hadoop installation. Here are the minimum steps to get SeaweedFS to run.
You would need to start a weed filer first, download the seaweedfs-hadoop-client-xxx.jar, and do the following:
You would need to start a weed filer first, build the seaweedfs-hadoop-client-xxx.jar, and do the following:
```
$ cd ${HADOOP_HOME}
# create etc/hadoop/mapred-site.xml, just to satisfy hdfs dfs. skip this if the file already exists.
$ echo "<configuration></configuration>" > etc/hadoop/mapred-site.xml
$ bin/hdfs dfs -Dfs.defaultFS=seaweedfs://localhost:8888 \
-Dfs.seaweedfs.impl=seaweed.hdfs.SeaweedFileSystem \
-libjars ./seaweedfs-hadoop-client-1.0-SNAPSHOT.jar \
-ls /
```
```
Both reads and writes are working fine.