From d843b2da79a63dd3323f6f5a19737f1d0a2386e5 Mon Sep 17 00:00:00 2001 From: Chris Lu Date: Thu, 13 Dec 2018 11:44:09 -0800 Subject: [PATCH] Updated Hadoop Compatible File System (markdown) --- Hadoop-Compatible-File-System.md | 27 +++++++++++++++++++++++++++ 1 file changed, 27 insertions(+) diff --git a/Hadoop-Compatible-File-System.md b/Hadoop-Compatible-File-System.md index f1026d3..c9835fc 100644 --- a/Hadoop-Compatible-File-System.md +++ b/Hadoop-Compatible-File-System.md @@ -77,8 +77,35 @@ Follow instructions on spark doc: * https://spark.apache.org/docs/latest/configuration.html#inheriting-hadoop-cluster-configuration * https://spark.apache.org/docs/latest/configuration.html#custom-hadoophive-configuration +## installation inheriting from Hadoop cluster configuration + Inheriting from Hadoop cluster configuration should be the easiest way. +To make these files visible to Spark, set HADOOP_CONF_DIR in $SPARK_HOME/conf/spark-env.sh to a location containing the configuration file `core-site.xml`, usually `/etc/hadoop/conf` + +## installation not inheriting from Hadoop cluster configuration + +Copy the seaweedfs-hadoop-client-x.x.x.jar to all executor machines. + +Add the following to spark/conf/spark-defaults.conf on every node running Spark +``` +spark.driver.extraClassPath /path/to/seaweedfs-hadoop-client-x.x.x.jar +spark.executor.extraClassPath /path/to/seaweedfs-hadoop-client-x.x.x.jar +``` + +And modify the configuration at runntime: + +``` +./bin/spark-submit \ + --name "My app" \ + --master local[4] \ + --conf spark.eventLog.enabled=false \ + --conf "spark.executor.extraJavaOptions=-XX:+PrintGCDetails -XX:+PrintGCTimeStamps" \ + --conf spark.hadoop.fs.seaweedfs.impl=seaweed.hdfs.SeaweedFileSystem \ + --conf spark.hadoop.fs.defaultFS=seaweedfs://localhost:8888 \ + myApp.jar +``` + # Installation for HBase If HBase is used, create a folder and configure the HBase root directory in `etc/hbase/conf/hbase-site.xml`: