mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-11 00:37:52 +02:00
Updated Hadoop Compatible File System (markdown)
1 parent
120d4752a1
commit
5c928ffff7
1 file changed
+18
-2
@@ -1,15 +1,31 @@
|
||||
HDFS is optimized for large files. The scalability of the single HDFS namenode is limited by the number of files. It is hard for HDFS to store lots of small files.
|
||||
|
||||
SeaweedFS excels on small files. Now it is possible to enable Hadoop jobs to read from and write to SeaweedFS.
|
||||
|
||||
# Build SeaweedFS Hadoop Client Jar
|
||||
```
|
||||
$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/client
|
||||
$ mvn install
|
||||
$cd $GOPATH/src/github.com/chrislusf/seaweedfs/other/java/hdfs
|
||||
$ mvn package
|
||||
$ ls -al target/seaweedfs-hadoop-client-1.0-SNAPSHOT.jar
|
||||
|
||||
```
|
||||
|
||||
# Test SeaweedFS on Hadoop
|
||||
|
||||
Suppose you are getting a new Hadoop installation. Here are the minimum steps to get SeaweedFS to run.
|
||||
|
||||
You would need to start a weed filer first, download the seaweedfs-hadoop-client-xxx.jar, and do the following:
|
||||
You would need to start a weed filer first, build the seaweedfs-hadoop-client-xxx.jar, and do the following:
|
||||
|
||||
```
|
||||
$ cd ${HADOOP_HOME}
|
||||
# create etc/hadoop/mapred-site.xml, just to satisfy hdfs dfs. skip this if the file already exists.
|
||||
$ echo "<configuration></configuration>" > etc/hadoop/mapred-site.xml
|
||||
$ bin/hdfs dfs -Dfs.defaultFS=seaweedfs://localhost:8888 \
|
||||
-Dfs.seaweedfs.impl=seaweed.hdfs.SeaweedFileSystem \
|
||||
-libjars ./seaweedfs-hadoop-client-1.0-SNAPSHOT.jar \
|
||||
-ls /
|
||||
|
||||
```
|
||||
```
|
||||
Both reads and writes are working fine.
|
||||
Reference in new issue
Block a user