{
  "id": 131286,
  "title": "Downloading Data on AWS",
  "url": "/competitions/deepfake-detection-challenge/discussion/131286",
  "author_name": "",
  "post_date": "2020-02-19T06:19:19.499756400Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi Everyone,</p>\n\n<p>I wanted to download the data directly to AWS as I have limitation on the upload speed making it hard for me to upload the data to AWS after downloading it locally to my laptop. As I'll need to use AWS for training.</p>\n\n<p>I tried most of the command line proposed techniques, but seems that Kaggle team has closed all of the. </p>\n\n<p>Here's the proposed setup that I'm using. I have 2 volumes:\n1. \"System Volume\" which I use spot instances to start such a volume to save money. This volume is usually small like 8GB or even less.\n2. \"Data Volume\" which a large drive that contains the data and the anaconda setup for the python environment. This way I don't have to reinstall the packages every time I start a new spot instance.</p>\n\n<p>I use Ubuntu 18.04 for the spot instance. I use different instances based on the need. Some times I use P3 instances if GPUs are required or regular instances like t2 and others for other processing.</p>\n\n<p>For downloading the data, I did the following:\n1. Started an Ubuntu 18.04 instance.\n2. Created the \"Data Volume\" and attached it to the newly created instance.\n3. Partition and mount the new volume using the following commands: <a href=\"https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html\">https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html</a>\n4. Installed VNC on the instance. Here is a link on how to do it: <a href=\"https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/\">https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/</a>\n5. You might need to install those packages as well: sudo apt install xfce4 xfce4-goodies\n6. To connect to the VNC, you have 2 choices, either tunnel port 5901 to your laptop from the instance or you could just open the port on AWS security group allowing access to the VNC session. BTW, 5901 is for vnc session 1 and 5902 for session 2, etc.. Keep that in mind.\n7. Download <a href=\"https://www.realvnc.com/en/connect/download/viewer/\">https://www.realvnc.com/en/connect/download/viewer/</a> to be able to open the VNC session.\n8. After viewing the VNC, open firefox and login to your account at Kaggle and download the data to your Data volume.</p>\n\n<p>Hope that helps.</p>\n\n<p>Regards,\nAmro</p>",
  "messages": [
    {
      "id": "750102",
      "postDate": "02/19/2020 06:19:19",
      "content": "<p>Hi Everyone,</p>\n\n<p>I wanted to download the data directly to AWS as I have limitation on the upload speed making it hard for me to upload the data to AWS after downloading it locally to my laptop. As I'll need to use AWS for training.</p>\n\n<p>I tried most of the command line proposed techniques, but seems that Kaggle team has closed all of the. </p>\n\n<p>Here's the proposed setup that I'm using. I have 2 volumes:\n1. \"System Volume\" which I use spot instances to start such a volume to save money. This volume is usually small like 8GB or even less.\n2. \"Data Volume\" which a large drive that contains the data and the anaconda setup for the python environment. This way I don't have to reinstall the packages every time I start a new spot instance.</p>\n\n<p>I use Ubuntu 18.04 for the spot instance. I use different instances based on the need. Some times I use P3 instances if GPUs are required or regular instances like t2 and others for other processing.</p>\n\n<p>For downloading the data, I did the following:\n1. Started an Ubuntu 18.04 instance.\n2. Created the \"Data Volume\" and attached it to the newly created instance.\n3. Partition and mount the new volume using the following commands: <a href=\"https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html\">https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html</a>\n4. Installed VNC on the instance. Here is a link on how to do it: <a href=\"https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/\">https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/</a>\n5. You might need to install those packages as well: sudo apt install xfce4 xfce4-goodies\n6. To connect to the VNC, you have 2 choices, either tunnel port 5901 to your laptop from the instance or you could just open the port on AWS security group allowing access to the VNC session. BTW, 5901 is for vnc session 1 and 5902 for session 2, etc.. Keep that in mind.\n7. Download <a href=\"https://www.realvnc.com/en/connect/download/viewer/\">https://www.realvnc.com/en/connect/download/viewer/</a> to be able to open the VNC session.\n8. After viewing the VNC, open firefox and login to your account at Kaggle and download the data to your Data volume.</p>\n\n<p>Hope that helps.</p>\n\n<p>Regards,\nAmro</p>",
      "rawMarkdown": "Hi Everyone,\n\nI wanted to download the data directly to AWS as I have limitation on the upload speed making it hard for me to upload the data to AWS after downloading it locally to my laptop. As I'll need to use AWS for training.\n\nI tried most of the command line proposed techniques, but seems that Kaggle team has closed all of the. \n\nHere's the proposed setup that I'm using. I have 2 volumes:\n1. \"System Volume\" which I use spot instances to start such a volume to save money. This volume is usually small like 8GB or even less.\n2. \"Data Volume\" which a large drive that contains the data and the anaconda setup for the python environment. This way I don't have to reinstall the packages every time I start a new spot instance.\n\nI use Ubuntu 18.04 for the spot instance. I use different instances based on the need. Some times I use P3 instances if GPUs are required or regular instances like t2 and others for other processing.\n\nFor downloading the data, I did the following:\n1. Started an Ubuntu 18.04 instance.\n2. Created the \"Data Volume\" and attached it to the newly created instance.\n3. Partition and mount the new volume using the following commands: https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html\n4. Installed VNC on the instance. Here is a link on how to do it: https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/\n5. You might need to install those packages as well: sudo apt install xfce4 xfce4-goodies\n6. To connect to the VNC, you have 2 choices, either tunnel port 5901 to your laptop from the instance or you could just open the port on AWS security group allowing access to the VNC session. BTW, 5901 is for vnc session 1 and 5902 for session 2, etc.. Keep that in mind.\n7. Download https://www.realvnc.com/en/connect/download/viewer/ to be able to open the VNC session.\n8. After viewing the VNC, open firefox and login to your account at Kaggle and download the data to your Data volume.\n\nHope that helps.\n\nRegards,\nAmro",
      "votes": null
    },
    {
      "id": "750304",
      "postDate": "02/19/2020 09:16:21",
      "content": "<p>Install CurlWget extension in chrome and download zip file in local system (interrupt after u got url from CurlWget). Copy that url from CurlWget and keep running on AWS. it will takes 2 hours to download all data.</p>\n\n<p>make sure u have extra 500GB for extracting zip files. <a href=\"/am1to2\">@am1to2</a> </p>",
      "rawMarkdown": "Install CurlWget extension in chrome and download zip file in local system (interrupt after u got url from CurlWget). Copy that url from CurlWget and keep running on AWS. it will takes 2 hours to download all data.\n\nmake sure u have extra 500GB for extracting zip files. @am1to2",
      "votes": null
    },
    {
      "id": "750311",
      "postDate": "02/19/2020 09:24:25",
      "content": "<p>Thanks <a href=\"/seshurajup\">@seshurajup</a> </p>",
      "rawMarkdown": "Thanks @seshurajup",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 750304,
      "author_name": "seshurajup",
      "author_url": "",
      "post_date": "02/19/2020 09:16:21",
      "content": "<p>Install CurlWget extension in chrome and download zip file in local system (interrupt after u got url from CurlWget). Copy that url from CurlWget and keep running on AWS. it will takes 2 hours to download all data.</p>\n\n<p>make sure u have extra 500GB for extracting zip files. <a href=\"/am1to2\">@am1to2</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 750311,
          "author_name": "am1to2",
          "author_url": "",
          "post_date": "02/19/2020 09:24:25",
          "content": "<p>Thanks <a href=\"/seshurajup\">@seshurajup</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "750102": "Hi Everyone,\n\nI wanted to download the data directly to AWS as I have limitation on the upload speed making it hard for me to upload the data to AWS after downloading it locally to my laptop. As I'll need to use AWS for training.\n\nI tried most of the command line proposed techniques, but seems that Kaggle team has closed all of the. \n\nHere's the proposed setup that I'm using. I have 2 volumes:\n1. \"System Volume\" which I use spot instances to start such a volume to save money. This volume is usually small like 8GB or even less.\n2. \"Data Volume\" which a large drive that contains the data and the anaconda setup for the python environment. This way I don't have to reinstall the packages every time I start a new spot instance.\n\nI use Ubuntu 18.04 for the spot instance. I use different instances based on the need. Some times I use P3 instances if GPUs are required or regular instances like t2 and others for other processing.\n\nFor downloading the data, I did the following:\n1. Started an Ubuntu 18.04 instance.\n2. Created the \"Data Volume\" and attached it to the newly created instance.\n3. Partition and mount the new volume using the following commands: https://smallbusiness.chron.com/partitioning-hard-drive-ubuntu-live-cd-38975.html\n4. Installed VNC on the instance. Here is a link on how to do it: https://ehikioya.com/vnc-for-ubuntu-amazon-ec2-gui-server/\n5. You might need to install those packages as well: sudo apt install xfce4 xfce4-goodies\n6. To connect to the VNC, you have 2 choices, either tunnel port 5901 to your laptop from the instance or you could just open the port on AWS security group allowing access to the VNC session. BTW, 5901 is for vnc session 1 and 5902 for session 2, etc.. Keep that in mind.\n7. Download https://www.realvnc.com/en/connect/download/viewer/ to be able to open the VNC session.\n8. After viewing the VNC, open firefox and login to your account at Kaggle and download the data to your Data volume.\n\nHope that helps.\n\nRegards,\nAmro",
    "750304": "Install CurlWget extension in chrome and download zip file in local system (interrupt after u got url from CurlWget). Copy that url from CurlWget and keep running on AWS. it will takes 2 hours to download all data.\n\nmake sure u have extra 500GB for extracting zip files. @am1to2",
    "750311": "Thanks @seshurajup"
  },
  "source": "meta"
}