{
  "id": 264694,
  "title": "How Can I Download Only New Data With Kaggle Command Line?",
  "url": "/competitions/seti-breakthrough-listen/discussion/264694",
  "author_name": "Chris Deotte",
  "post_date": "2021-08-12T21:48:08.578000",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I would like to download just the new data with Kaggle command line tool. I do not want to download the additional 70GB of old data. Does anyone know how to do this?</p>\n<p>To download a single file, we type</p>\n<pre><code>kaggle competitions download -f [FILE_NAME] -c [competition]\n</code></pre>\n<p>How do we download a folder of images?</p>",
  "messages": [
    {
      "id": 1469361,
      "postDate": "2021-08-12T21:48:08.580Z",
      "content": "<p>I would like to download just the new data with Kaggle command line tool. I do not want to download the additional 70GB of old data. Does anyone know how to do this?</p>\n<p>To download a single file, we type</p>\n<pre><code>kaggle competitions download -f [FILE_NAME] -c [competition]\n</code></pre>\n<p>How do we download a folder of images?</p>",
      "rawMarkdown": "I would like to download just the new data with Kaggle command line tool. I do not want to download the additional 70GB of old data. Does anyone know how to do this?\n\nTo download a single file, we type\n\n    kaggle competitions download -f [FILE_NAME] -c [competition]\n\nHow do we download a folder of images?",
      "votes": 4
    },
    {
      "id": 1479727,
      "postDate": "2021-08-18T15:54:30.120Z",
      "content": "<p>It's a late response.<br>\nI use wget [url].<br>\nTo get the url of a folder, just try to download the folder with any Browser(like Chrome), and you can get the url in \"download\" by right clicking it.</p>",
      "rawMarkdown": "It's a late response.\nI use wget [url].\nTo get the url of a folder, just try to download the folder with any Browser(like Chrome), and you can get the url in \"download\" by right clicking it.",
      "votes": 1
    },
    {
      "id": 1470054,
      "postDate": "2021-08-13T09:02:25.177Z",
      "content": "<p>I downloaded train folder from competitions data page, then I uploaded the zip file to the server.<br>\nYou can select train folder, then click the download button, you can download only train folder.</p>",
      "rawMarkdown": "I downloaded train folder from competitions data page, then I uploaded the zip file to the server.\nYou can select train folder, then click the download button, you can download only train folder.",
      "votes": 2,
      "replies": [
        {
          "id": 1470124,
          "postDate": "2021-08-13T10:04:22.310Z",
          "content": "<p>Selecting \"train\" and \"test\" and downloading as zip was also the way i did it.<br>\nNow, if you are working with remote machines, that reset every time you run them up, you probably want an automatic download script for the data. Selecting a subset would save a lot of bandwidth, that is not always needed. Adding this functionality to the kaggle command line tool would be greatly appreciated. Hosting the already available data on an additional server doesn't really sound like an optimal solution. </p>",
          "rawMarkdown": "Selecting \"train\" and \"test\" and downloading as zip was also the way i did it.\nNow, if you are working with remote machines, that reset every time you run them up, you probably want an automatic download script for the data. Selecting a subset would save a lot of bandwidth, that is not always needed. Adding this functionality to the kaggle command line tool would be greatly appreciated. Hosting the already available data on an additional server doesn't really sound like an optimal solution. ",
          "votes": 2
        },
        {
          "id": 1470242,
          "postDate": "2021-08-13T11:54:10.820Z",
          "content": "<p>Thanks. Kaggle should allow wildcards in their CLI like below (and package the files into single zip download):</p>\n<pre><code>kaggle competitions download -f  train/*/* -c seti-breakthrough-listen\n</code></pre>\n<p>in the same way that zip allows wildcards to avoid unzipping old data (if we are required to download both)</p>\n<pre><code>unzip seti-breakthrough-listen.zip -x old_leaky_data/*/*/*\n</code></pre>",
          "rawMarkdown": "Thanks. Kaggle should allow wildcards in their CLI like below (and package the files into single zip download):\n\n    kaggle competitions download -f  train/*/* -c seti-breakthrough-listen\n\nin the same way that zip allows wildcards to avoid unzipping old data (if we are required to download both)\n\n    unzip seti-breakthrough-listen.zip -x old_leaky_data/*/*/*",
          "votes": 5
        },
        {
          "id": 1470254,
          "postDate": "2021-08-13T12:08:55.493Z",
          "content": "<p>Great concept!<br>\nIf you ask for such feature in <a href=\"https://www.kaggle.com/product-feedback\" target=\"_blank\">product feedback</a>, count me in for supporting the proposal.</p>",
          "rawMarkdown": "Great concept!\nIf you ask for such feature in [product feedback](https://www.kaggle.com/product-feedback), count me in for supporting the proposal.",
          "votes": 3
        },
        {
          "id": 1470289,
          "postDate": "2021-08-13T12:45:01.587Z",
          "content": "<p>I posted a feature request here. Everyone who thinks this would be helpful should upvote. Thanks<br>\n<a href=\"https://www.kaggle.com/product-feedback/264838\" target=\"_blank\">https://www.kaggle.com/product-feedback/264838</a></p>",
          "rawMarkdown": "I posted a feature request here. Everyone who thinks this would be helpful should upvote. Thanks\nhttps://www.kaggle.com/product-feedback/264838",
          "votes": 7
        },
        {
          "id": 1472209,
          "postDate": "2021-08-14T17:51:55.763Z",
          "content": "<p>Upvoted <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
          "rawMarkdown": "Upvoted @cdeotte !",
          "votes": 1
        }
      ]
    },
    {
      "id": 1469377,
      "postDate": "2021-08-12T22:18:26.770Z",
      "content": "<p>What about <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253912\" target=\"_blank\">this solution</a> ?</p>",
      "rawMarkdown": "What about [this solution](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253912) ?",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1479727,
      "author_name": "Last Scene",
      "author_url": "",
      "post_date": "2021-08-18T15:54:30.120000",
      "content": "<p>It's a late response.<br>\nI use wget [url].<br>\nTo get the url of a folder, just try to download the folder with any Browser(like Chrome), and you can get the url in \"download\" by right clicking it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1470054,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2021-08-13T09:02:25.177000",
      "content": "<p>I downloaded train folder from competitions data page, then I uploaded the zip file to the server.<br>\nYou can select train folder, then click the download button, you can download only train folder.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1470124,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-13T10:04:22.310000",
          "content": "<p>Selecting \"train\" and \"test\" and downloading as zip was also the way i did it.<br>\nNow, if you are working with remote machines, that reset every time you run them up, you probably want an automatic download script for the data. Selecting a subset would save a lot of bandwidth, that is not always needed. Adding this functionality to the kaggle command line tool would be greatly appreciated. Hosting the already available data on an additional server doesn't really sound like an optimal solution. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1470242,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-13T11:54:10.820000",
          "content": "<p>Thanks. Kaggle should allow wildcards in their CLI like below (and package the files into single zip download):</p>\n<pre><code>kaggle competitions download -f  train/*/* -c seti-breakthrough-listen\n</code></pre>\n<p>in the same way that zip allows wildcards to avoid unzipping old data (if we are required to download both)</p>\n<pre><code>unzip seti-breakthrough-listen.zip -x old_leaky_data/*/*/*\n</code></pre>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1470254,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-08-13T12:08:55.493000",
          "content": "<p>Great concept!<br>\nIf you ask for such feature in <a href=\"https://www.kaggle.com/product-feedback\" target=\"_blank\">product feedback</a>, count me in for supporting the proposal.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1470289,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-13T12:45:01.587000",
          "content": "<p>I posted a feature request here. Everyone who thinks this would be helpful should upvote. Thanks<br>\n<a href=\"https://www.kaggle.com/product-feedback/264838\" target=\"_blank\">https://www.kaggle.com/product-feedback/264838</a></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1472209,
          "author_name": "Old Monk",
          "author_url": "",
          "post_date": "2021-08-14T17:51:55.763000",
          "content": "<p>Upvoted <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1469377,
      "author_name": "kkiller",
      "author_url": "",
      "post_date": "2021-08-12T22:18:26.770000",
      "content": "<p>What about <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253912\" target=\"_blank\">this solution</a> ?</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1469361": "I would like to download just the new data with Kaggle command line tool. I do not want to download the additional 70GB of old data. Does anyone know how to do this?\n\nTo download a single file, we type\n\n    kaggle competitions download -f [FILE_NAME] -c [competition]\n\nHow do we download a folder of images?",
    "1479727": "It's a late response.\nI use wget [url].\nTo get the url of a folder, just try to download the folder with any Browser(like Chrome), and you can get the url in \"download\" by right clicking it.",
    "1470054": "I downloaded train folder from competitions data page, then I uploaded the zip file to the server.\nYou can select train folder, then click the download button, you can download only train folder.",
    "1469377": "What about [this solution](https://www.kaggle.com/c/seti-breakthrough-listen/discussion/253912) ?"
  }
}