{"templateId":"markdown","sharedDataIds":{},"props":{"metadata":{"markdoc":{"tagList":[]},"type":"markdown"},"seo":{"title":"Amazon Elastic MapReduce","description":"Amazon Elastic MapReduceでTreasure AIのtd-sparkドライバを使用し、Sparkからデータにアクセスする方法を解説する技術ドキュメントページです。","siteUrl":"https://docs.treasure.ai","lang":"en-US","jsonLd":{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.treasure.ai/","name":"Treasure AI","url":"https://www.treasure.ai/","logo":"https://www.treasure.ai/hubfs/assets/images/logos/primary-logo.svg"},{"@type":"WebSite","@id":"https://docs.treasure.ai/#website","name":"Treasure AI Documentation","url":"https://docs.treasure.ai/","inLanguage":["en","ja"],"publisher":{"@id":"https://www.treasure.ai/"}}]},"llmstxt":{"hide":false,"sections":[{"title":"Table of contents","includeFiles":["**/*"],"excludeFiles":[]}],"excludeFiles":[]}},"dynamicMarkdocComponents":[],"compilationErrors":[],"ast":{"$$mdtype":"Tag","name":"article","attributes":{},"children":[{"$$mdtype":"Tag","name":"Heading","attributes":{"level":1,"id":"amazon-elastic-mapreduce","__idx":0},"children":["Amazon Elastic MapReduce"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Amazon Elastic MapReduce (EMR)でApache Spark Driver for Treasure Data(td-sparkとも呼ばれます)を使用できます。最適なパフォーマンスを得るためにAmazon EC2の",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["us-east"]},"リージョンの使用をお勧めしますが、他のSpark環境でもtd-sparkを使用できます。"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"概要","__idx":1},"children":["概要"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["td-sparkの概要については、",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/products/customer-data-platform/data-workbench/queries/trino/query_faqs"},"children":["TD Spark FAQs"]},"を参照してください。"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"td-sparkで何ができるか","__idx":2},"children":["td-sparkで何ができるか?"]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":["ScalaとPython (PySpark)のSparkからTreasure Dataにアクセスします。"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Treasure DataテーブルをSparkにDataFrameとして取り込みます(TDクエリは発行されず、TD保存データとSparkクラスター間の最短のレイテンシパスを提供します)。"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Presto、Hive、またはSparkSQLクエリを発行し、結果をSpark DataFrameとして返します。"]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["このドライバは、Spark 2.4.4以降での使用が推奨されます"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"使用に関する推奨事項","__idx":3},"children":["使用に関する推奨事項"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["最速のデータアクセスと最低のデータ転送コストのために、AWSのus-eastリージョンでSparkクラスターをセットアップすることをお勧めします。他のAWSリージョンまたは処理環境を使用すると、データ転送コストが非常に高くなる可能性があります。"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"emr上のtd-spark-driver","__idx":4},"children":["EMR上のTD Spark Driver"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"emr-sparkクラスターの作成","__idx":5},"children":["EMR Sparkクラスターの作成"]},{"$$mdtype":"Tag","name":"ol","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Sparkサポート付きのEMRクラスターを作成します。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["S3からのデータ転送パフォーマンスを最大化するために、",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["us-east"]},"リージョンが推奨されます。"]}]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"p","attributes":{},"children":["新しいEMRのマスターノードアドレスを確認します。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["テーブルデータをSpark DataFrameとして読み取ります"]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-9-26.c6538e30038a6b2dd513a89175ca7d459cba989c4e8b5491b087998ef2f4e30e.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-11-11.ce52ded660f570f751fb8d590bf08c5bc1575a6656533e0c4587e5a23ce1ec39.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["デフォルトのセキュリティグループ(ElasticMapReduce-master)でEMRを作成した場合は、環境からのインバウンドアクセスを許可する必要があります。",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"http://docs.aws.amazon.com/ElasticMapReduce/latest/DeveloperGuide/emr-man-sec-groups.html"},"children":["Amazon EMR-Managed Security Groups"]},"を参照してください。"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"リファレンス","__idx":6},"children":["リファレンス"]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"http://docs.aws.amazon.com/ElasticMapReduce/latest/ReleaseGuide/emr-spark-launch.html"},"children":["Create An EMR Cluster with Spark"]}]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"emrクラスターへのログイン","__idx":7},"children":["EMRクラスターへのログイン"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"http://docs.aws.amazon.com//ElasticMapReduce/latest/ManagementGuide/emr-connect-master-node-ssh.html"},"children":["Connect to EMR Master node with SSH"]}]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"# SOCKS5プロキシポート8157を使用して、EMR Sparkジョブ履歴ページ(ポート18080)、Zeppelinノートブック(ポート8890)などにアクセスできるようにします。\n$ ssh -i (your AWS key pair file. .pem) -D8157 hadoop@ec2-xxx-xxx-xxx-xxx.compute-1.amazonaws.com\n     __|  __|_  )\n       _|  (     /   Amazon Linux AMI\n      ___|\\___|___|\nhttps://aws.amazon.com/amazon-linux-ami/2016.09-release-notes/\n4 package(s) needed for security, out of 6 available\nRun \"sudo yum update\" to apply all updates.\n\nEEEEEEEEEEEEEEEEEEEE MMMMMMMM           MMMMMMMM RRRRRRRRRRRRRRR\nE::::::::::::::::::E M:::::::M         M:::::::M R::::::::::::::R\nEE:::::EEEEEEEEE:::E M::::::::M       M::::::::M R:::::RRRRRR:::::R\n  E::::E       EEEEE M:::::::::M     M:::::::::M RR::::R      R::::R\n  E::::E             M::::::M:::M   M:::M::::::M   R:::R      R::::R\n  E:::::EEEEEEEEEE   M:::::M M:::M M:::M M:::::M   R:::RRRRRR:::::R\n  E::::::::::::::E   M:::::M  M:::M:::M  M:::::M   R:::::::::::RR\n  E:::::EEEEEEEEEE   M:::::M   M:::::M   M:::::M   R:::RRRRRR::::R\n  E::::E             M:::::M    M:::M    M:::::M   R:::R      R::::R\n  E::::E       EEEEE M:::::M     MMM     M:::::M   R:::R      R::::R\nEE:::::EEEEEEEE::::E M:::::M             M:::::M   R:::R      R::::R\nE::::::::::::::::::E M:::::M             M:::::M RR::::R      R::::R\nEEEEEEEEEEEEEEEEEEEE MMMMMMM             MMMMMMM RRRRRRR      RRRRRR\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"td-spark-integrationのセットアップ","__idx":8},"children":["TD Spark Integrationのセットアップ"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["td-spark jarファイルをダウンロードします:"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"[hadoop@ip-x-x-x-x]$ wget https://s3.amazonaws.com/td-spark/td-spark-assembly_2.11-0.4.0.jar\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["マスターノードに",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["td.conf"]},"ファイルを作成します:"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"# ここにTD APIキーを記述します\nspark.td.apikey (your TD API key)\nspark.td.site (your site name: us, jp, ap02, eu01, etc.)\n# (推奨) より高速なパフォーマンスのためにKryoSerializerを使用\nspark.serializer org.apache.spark.serializer.KryoSerializer\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"emrでspark-shellを使用","__idx":9},"children":["EMRでspark-shellを使用"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"[hadoop@ip-x-x-x-x]$ spark-shell --master yarn --jars td-spark-assembly-latest.jar --properties-file td.conf\nWelcome to\n      ____              __\n     / __/__  ___ _____/ /__\n    _\\ \\/ _ \\/ _ `/ __/  '_/\n   /___/ .__/\\_,_/_/ /_/\\_\\   version 2.1.0\n      /_/\nscala> import com.treasuredata.spark._\nscala> val td = spark.td\nscala> val d = td.table(\"sample_datasets.www_access\").df\nscala> d.show\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\n|user|           host|                path|             referer|code|               agent|size|method|      time|\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\n|null|136.162.131.221|    /category/health|   /category/cameras| 200|Mozilla/5.0 (Wind...|  77|   GET|1412373596|\n|null| 172.33.129.134|      /category/toys|   /item/office/4216| 200|Mozilla/5.0 (comp...| 115|   GET|1412373585|\n|null| 220.192.77.135|  /category/software|                   -| 200|Mozilla/5.0 (comp...| 116|   GET|1412373574|\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\nonly showing top 3 rows\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"emrでzeppelin-notebookを使用","__idx":10},"children":["EMRでZeppelin Notebookを使用"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"td-spark用にzeppelinを設定","__idx":11},"children":["td-spark用にZeppelinを設定"]},{"$$mdtype":"Tag","name":"ol","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":["EMRクラスターへのSSHトンネルを作成します。"]}]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"$ ssh -i (your AWS key pair file. .pem) -D8157 hadoop@ec2-xxx-xxx-xxx-xxx.compute-1.amazonaws.com\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["(Chromeユーザー向け) Proxy Switchy Sharp Chrome Extensionをインストールします。EMRマスターにアクセスする際にEMRのproxy-switchをオンにします"," ","2. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["http://(your EMR master node public address):8890/"]},"を開きます"," ","3. td-sparkを設定するために",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Interpreters"]},"ページを開きます。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-15-28.a4f96f2069e1df24d7c95be0d35e0d01c8821747176e796e484e3c15a686c98f.d33fa2c4.png","alt":""},"children":[]}," ","4. プロファイルの詳細を編集し、",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Save"]},"を選択します。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-16-54.cb15bd5327e74c96e696f3d16f94c0eca3c92fc777e25e082d78cf668caca70b.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"td内のデータセットにdataframeとしてアクセス","__idx":12},"children":["TD内のデータセットにDataFrameとしてアクセス"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["テーブルデータをSpark DataFrameとして読み取ることができます。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-17-56.82c18313f4714d4d7a00e721ac7cd60cd21bf1ec4ab882db2e1be59b85254977.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"prestoクエリの実行","__idx":13},"children":["Prestoクエリの実行"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-18-48.0e69e43eeae788cd1b81f341a8ac3874bfcd86cb55306d80b9166ad729cf9b5a.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"spark-history-serverの確認","__idx":14},"children":["Spark History Serverの確認"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["EMRマスターノードのパブリックアドレスを使用して、",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["History Server"]},"でイベント履歴を確認できます。"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["http://(your EMR master node public address):18080/"]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"img","attributes":{"src":"/assets/image2020-11-18_9-21-52.42a636aa3768897c576b4224cfb24be36c330e2135b7c1f7a7b9c5558e55b395.d33fa2c4.png","alt":""},"children":[]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"pysparkとsparksqlでのtd-spark-driverの使用","__idx":15},"children":["PySparkとSparkSQLでのTD Spark Driverの使用"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"pyspark","__idx":16},"children":["PySpark"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"bash","header":{"controls":{"copy":{}}},"source":"$ ./bin/pyspark  --driver-class-path ~/work/git/td-spark/td-spark/target/td-spark-assembly-0.1-SNAPSHOT.jar --properties-file ../td-dev.conf\n\n>>> df = spark.read.format(\"com.treasuredata.spark\").load(\"sample_datasets.www_access\")\n>>> df.show(10)\n2016-07-19 16:34:15-0700  info [TDRelation] Fetching www_access within time range:[-9223372036854775808,9223372036854775807) - (TDRelation.scala:82)\n2016-07-19 16:34:16-0700  info [TDRelation] Retrieved 19 PlazmaAPI entries - (TDRelation.scala:85)\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\n|user|           host|                path|             referer|code|               agent|size|method|      time|\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\n|null| 148.165.90.106|/category/electro...|     /category/music| 200|Mozilla/4.0 (comp...|  66|   GET|1412333993|\n|null|  144.105.77.66|/item/electronics...| /category/computers| 200|Mozilla/5.0 (iPad...| 135|   GET|1412333977|\n|null| 108.54.178.116|/category/electro...|  /category/software| 200|Mozilla/5.0 (Wind...|  69|   GET|1412333961|\n|null|104.129.105.202|/item/electronics...|     /item/games/394| 200|Mozilla/5.0 (comp...|  83|   GET|1412333945|\n|null|   208.48.26.63|  /item/software/706|/search/?c=Softwa...| 200|Mozilla/5.0 (comp...|  76|   GET|1412333930|\n|null|  108.78.209.95|/item/giftcards/4879|      /item/toys/197| 200|Mozilla/5.0 (Wind...| 137|   GET|1412333914|\n|null| 108.198.97.206|/item/computers/4785|                   -| 200|Mozilla/5.0 (Wind...|  69|   GET|1412333898|\n|null| 172.195.185.46|     /category/games|     /category/games| 200|Mozilla/5.0 (Maci...|  41|   GET|1412333882|\n|null|   88.24.72.177|/item/giftcards/4410|                   -| 200|Mozilla/4.0 (comp...|  72|   GET|1412333866|\n|null|  24.129.141.79|/category/electro...|/category/networking| 200|Mozilla/5.0 (comp...|  73|   GET|1412333850|\n+----+---------------+--------------------+--------------------+----+--------------------+----+------+----------+\nonly showing top 10 rows\n\n## Prestoジョブの送信\n>>> df = spark.read.format(\"com.treasuredata.spark\").options(sql=\"select 1\").load(\"sample_datasets\")\n2016-07-19 16:56:56-0700  info [TDSparkContext]\nSubmitted job 515990:\nselect 1 - (TDSparkContext.scala:70)\n>>> df.show(10)\n+-----+\n|_col0|\n+-----+\n|    1|\n+-----+\n\n## ジョブ結果の読み取り\n>>> df = sqlContext.read.format(\"com.treasuredata.spark\").load(\"job_id:515990\")\n","lang":"bash"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"sparksql","__idx":17},"children":["SparkSQL"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"scala","header":{"controls":{"copy":{}}},"source":"# DataFrameを一時テーブルとして登録\nscala> td.df(\"hb_tiny.rankings\").createOrReplaceTempView(\"rankings\")\n\nscala> val q1 = spark.sql(\"select page_url, page_rank from rankings where page_rank > 100\")\nq1: org.apache.spark.sql.DataFrame = [page_url: string, page_rank: bigint]\n\nscala> q1.show\n2016-07-20 11:27:11-0700  info [TDRelation] Fetching rankings within time range:[-9223372036854775808,9223372036854775807) - (TDRelation.scala:82)\n2016-07-20 11:27:12-0700  info [TDRelation] Retrieved 2 PlazmaAPI entries - (TDRelation.scala:85)\n+--------------------+---------+\n|            page_url|page_rank|\n+--------------------+---------+\n|xjhmjsuqolfklbvxn...|      251|\n|seozvzwkcfgnfuzfd...|      165|\n|fdgvmwbrjlmvuoquy...|      132|\n|gqghyyardomubrfsv...|      108|\n|qtqntqkvqioouwfuj...|      278|\n|wrwgqnhxviqnaacnc...|      135|\n|  cxdmunpixtrqnvglnt|      146|\n| ixgiosdefdnhrzqomnf|      126|\n|xybwfjcuhauxiopfi...|      112|\n|ecfuzdmqkvqktydvi...|      237|\n|dagtwwybivyiuxmkh...|      177|\n|emucailxlqlqazqru...|      134|\n|nzaxnvjaqxapdjnzb...|      119|\n|       ffygkvsklpmup|      332|\n|hnapejzsgqrzxdswz...|      171|\n|rvbyrwhzgfqvzqkus...|      148|\n|knwlhzmcyolhaccqr...|      104|\n|nbizrgdziebsaecse...|      665|\n|jakofwkgdcxmaaqph...|      187|\n|kvhuvcjzcudugtidf...|      120|\n+--------------------+---------+\nonly showing top 20 rows\n","lang":"scala"},"children":[]}]},"headings":[{"value":"Amazon Elastic MapReduce","id":"amazon-elastic-mapreduce","depth":1},{"value":"概要","id":"概要","depth":2},{"value":"td-sparkで何ができるか?","id":"td-sparkで何ができるか","depth":2},{"value":"使用に関する推奨事項","id":"使用に関する推奨事項","depth":2},{"value":"EMR上のTD Spark Driver","id":"emr上のtd-spark-driver","depth":2},{"value":"EMR Sparkクラスターの作成","id":"emr-sparkクラスターの作成","depth":3},{"value":"リファレンス","id":"リファレンス","depth":3},{"value":"EMRクラスターへのログイン","id":"emrクラスターへのログイン","depth":2},{"value":"TD Spark Integrationのセットアップ","id":"td-spark-integrationのセットアップ","depth":2},{"value":"EMRでspark-shellを使用","id":"emrでspark-shellを使用","depth":2},{"value":"EMRでZeppelin Notebookを使用","id":"emrでzeppelin-notebookを使用","depth":2},{"value":"td-spark用にZeppelinを設定","id":"td-spark用にzeppelinを設定","depth":3},{"value":"TD内のデータセットにDataFrameとしてアクセス","id":"td内のデータセットにdataframeとしてアクセス","depth":3},{"value":"Prestoクエリの実行","id":"prestoクエリの実行","depth":3},{"value":"Spark History Serverの確認","id":"spark-history-serverの確認","depth":3},{"value":"PySparkとSparkSQLでのTD Spark Driverの使用","id":"pysparkとsparksqlでのtd-spark-driverの使用","depth":2},{"value":"PySpark","id":"pyspark","depth":3},{"value":"SparkSQL","id":"sparksql","depth":3}],"frontmatter":{"seo":{"title":"Amazon Elastic MapReduce","description":"Amazon Elastic MapReduceでTreasure AIのtd-sparkドライバを使用し、Sparkからデータにアクセスする方法を解説する技術ドキュメントページです。"}},"lastModified":"2026-06-17T07:22:53.000Z","pagePropGetterError":{"message":"","name":""}},"slug":"/ja/int/amazon-elastic-mapreduce","userData":{"isAuthenticated":false,"teams":["anonymous"]},"isPublic":true}