{
  "id": 36485,
  "title": "sample.tar.gz - Have you tried using it?",
  "url": "/competitions/passenger-screening-algorithm-challenge/discussion/36485",
  "author_name": "",
  "post_date": "2017-07-17T03:48:03.204709100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Has anyone tried working with the provided, downloadable data file sample.tar.gz, and if yes, were you successful? </p>\n\n<p>I'm using R 3.4 on an up-to-date Windows 10 (sorry!) with RStudio.</p>\n\n<p>I downloaded sample.tar.gz and found a way, I thought, to list the files within without unzipping by  using the untar() command. I'd rather leave the file in its .tar.gz state if possible due to the its size when extracted. I am not, however, out of space on my hard disk. I checked! (767 GB free ...)</p>\n\n<p>untar() does appear to start off listing contained files. For example, here's an excerpt:</p>\n\n<p>[17] \"sample/00360f79fd6e02781457eda48f85da90.a3d\" <br>\n[18] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Extended header uid=457615 is out of range 0..32767\" <br>\n[19] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Ignoring unknown extended header keyword <code>SCHILY.dev'\" <br>\n[20] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword</code>SCHILY.ino'\" <br>\n[21] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Ignoring unknown extended header keyword `SCHILY.nlink'\"           </p>\n\n<p>At the end of untar()'s attempt, however, is </p>\n\n<p>*<strong>... tar.exe: Error exit delayed from previous errors\" <br>\nattr(,\"status\")\n[1] 2\nWarning message:\nrunning command 'tar.exe -tf \"sample.tar.gz\"' had status 2 *</strong></p>\n\n<p>I haven't entirely figured out what this means yet, but it doesn't  all sound good. I have been able to unzip a specific .aps file within the tar.gz, but would like to be able to see a list of all the contained *.aps files. If untar()'s list is failing before it finishes listing everything, I don't know what I'm missing. I tried specifying a wildcard to untar() using the file argument, but that didn't do anything. Perhaps I don't have the syntax right ...</p>\n\n<p>untar(\"sample.tar.gz\", files = \"sample/*.aps\", list = FALSE, exdir = \".\", compressed = NA, extras = NULL, verbose = TRUE, restore_times =  TRUE, tar = Sys.getenv(\"TAR\"))</p>\n\n<p>I know there's a Google Cloud area for accessing the data files, and I will eventually graduate to that. For now, I'd really like to just start with the sample.tar.gz, if possible.</p>\n\n<p>Thanks in advance for your thoughts on how I might be able to get around the above mentioned errors and/or warnings.</p>",
  "messages": [
    {
      "id": "203935",
      "postDate": "07/17/2017 03:48:03",
      "content": "<p>Has anyone tried working with the provided, downloadable data file sample.tar.gz, and if yes, were you successful? </p>\n\n<p>I'm using R 3.4 on an up-to-date Windows 10 (sorry!) with RStudio.</p>\n\n<p>I downloaded sample.tar.gz and found a way, I thought, to list the files within without unzipping by  using the untar() command. I'd rather leave the file in its .tar.gz state if possible due to the its size when extracted. I am not, however, out of space on my hard disk. I checked! (767 GB free ...)</p>\n\n<p>untar() does appear to start off listing contained files. For example, here's an excerpt:</p>\n\n<p>[17] \"sample/00360f79fd6e02781457eda48f85da90.a3d\" <br>\n[18] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Extended header uid=457615 is out of range 0..32767\" <br>\n[19] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Ignoring unknown extended header keyword <code>SCHILY.dev'\" <br>\n[20] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword</code>SCHILY.ino'\" <br>\n[21] \"C:\\PROGRA~2\\TESSER~1\\tar.exe: Ignoring unknown extended header keyword `SCHILY.nlink'\"           </p>\n\n<p>At the end of untar()'s attempt, however, is </p>\n\n<p>*<strong>... tar.exe: Error exit delayed from previous errors\" <br>\nattr(,\"status\")\n[1] 2\nWarning message:\nrunning command 'tar.exe -tf \"sample.tar.gz\"' had status 2 *</strong></p>\n\n<p>I haven't entirely figured out what this means yet, but it doesn't  all sound good. I have been able to unzip a specific .aps file within the tar.gz, but would like to be able to see a list of all the contained *.aps files. If untar()'s list is failing before it finishes listing everything, I don't know what I'm missing. I tried specifying a wildcard to untar() using the file argument, but that didn't do anything. Perhaps I don't have the syntax right ...</p>\n\n<p>untar(\"sample.tar.gz\", files = \"sample/*.aps\", list = FALSE, exdir = \".\", compressed = NA, extras = NULL, verbose = TRUE, restore_times =  TRUE, tar = Sys.getenv(\"TAR\"))</p>\n\n<p>I know there's a Google Cloud area for accessing the data files, and I will eventually graduate to that. For now, I'd really like to just start with the sample.tar.gz, if possible.</p>\n\n<p>Thanks in advance for your thoughts on how I might be able to get around the above mentioned errors and/or warnings.</p>",
      "rawMarkdown": "Has anyone tried working with the provided, downloadable data file sample.tar.gz, and if yes, were you successful? \n\nI'm using R 3.4 on an up-to-date Windows 10 (sorry!) with RStudio.\n\nI downloaded sample.tar.gz and found a way, I thought, to list the files within without unzipping by  using the untar() command. I'd rather leave the file in its .tar.gz state if possible due to the its size when extracted. I am not, however, out of space on my hard disk. I checked! (767 GB free ...)\n\nuntar() does appear to start off listing contained files. For example, here's an excerpt:\n\n[17] \"sample/00360f79fd6e02781457eda48f85da90.a3d\"                                                        \n[18] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Extended header uid=457615 is out of range 0..32767\"               \n[19] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.dev'\"             \n[20] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.ino'\"             \n[21] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.nlink'\"           \n\nAt the end of untar()'s attempt, however, is \n\n***... tar.exe: Error exit delayed from previous errors\"                           \nattr(,\"status\")\n[1] 2\nWarning message:\nrunning command 'tar.exe -tf \"sample.tar.gz\"' had status 2 ***\n\nI haven't entirely figured out what this means yet, but it doesn't  all sound good. I have been able to unzip a specific .aps file within the tar.gz, but would like to be able to see a list of all the contained *.aps files. If untar()'s list is failing before it finishes listing everything, I don't know what I'm missing. I tried specifying a wildcard to untar() using the file argument, but that didn't do anything. Perhaps I don't have the syntax right ...\n\nuntar(\"sample.tar.gz\", files = \"sample/*.aps\", list = FALSE, exdir = \".\", compressed = NA, extras = NULL, verbose = TRUE, restore_times =  TRUE, tar = Sys.getenv(\"TAR\"))\n\nI know there's a Google Cloud area for accessing the data files, and I will eventually graduate to that. For now, I'd really like to just start with the sample.tar.gz, if possible.\n\nThanks in advance for your thoughts on how I might be able to get around the above mentioned errors and/or warnings.",
      "votes": null
    },
    {
      "id": "204088",
      "postDate": "07/17/2017 15:22:33",
      "content": "<p>Yes, I have been able to download and uncompress it without any problems. It uses 5.7GB compressed and uncompressed also uses 5.7GB.</p>\n\n<p>Maybe you could try to download again the file.</p>",
      "rawMarkdown": "Yes, I have been able to download and uncompress it without any problems. It uses 5.7GB compressed and uncompressed also uses 5.7GB.\n\nMaybe you could try to download again the file.",
      "votes": null
    },
    {
      "id": "204113",
      "postDate": "07/17/2017 16:12:57",
      "content": "<p>Thank you, ironbar! After playing around with it some more, I too have been able to download and uncompress all of sample.tar.gz. The warnings and errors still display when running untar(), but I think in actuality they are just threats and I have the 16 files to work with. (4 files for each of the 4 file types)</p>\n\n<p>My next question is ... Do you happen to know the difference between the files that <em>do</em> and the files that <em>do not</em> start with a \".\" in this sample download? I am able to work with 8 of the files just fine, but the 8 starting with a \"period\" character are throwing an error in my read-data function. [ i.e. filename.aps vs. .filename.aps ]</p>\n\n<p>UPDATE: I may be able to answer my own question. The one without the leading \".\" is the data file, and the one with the leading \".\" is the header file. If I'm wrong about this, please let me know!</p>",
      "rawMarkdown": "Thank you, ironbar! After playing around with it some more, I too have been able to download and uncompress all of sample.tar.gz. The warnings and errors still display when running untar(), but I think in actuality they are just threats and I have the 16 files to work with. (4 files for each of the 4 file types)\n\nMy next question is ... Do you happen to know the difference between the files that *do* and the files that *do not* start with a \".\" in this sample download? I am able to work with 8 of the files just fine, but the 8 starting with a \"period\" character are throwing an error in my read-data function. [ i.e. filename.aps vs. .filename.aps ]\n\nUPDATE: I may be able to answer my own question. The one without the leading \".\" is the data file, and the one with the leading \".\" is the header file. If I'm wrong about this, please let me know!",
      "votes": null
    },
    {
      "id": "204779",
      "postDate": "07/19/2017 16:46:35",
      "content": "<p>I don't have files starting by \".\" on my uncompressed folder. </p>\n\n<p>And also reading the kernel it seems that the header is inside the file, not outside in another file. <br>\n<a href=\"https://www.kaggle.com/wcukierski/reading-images\">https://www.kaggle.com/wcukierski/reading-images</a></p>",
      "rawMarkdown": "I don't have files starting by \".\" on my uncompressed folder. \n\nAnd also reading the kernel it seems that the header is inside the file, not outside in another file.   \nhttps://www.kaggle.com/wcukierski/reading-images",
      "votes": null
    },
    {
      "id": "223075",
      "postDate": "09/21/2017 03:22:02",
      "content": "<p>If it's still relevant: I definitely have the files beginning with a period, as well.  Also, the warnings when expanding the tar file seem to be caused by the file being marked as \"quarantined,\" which the creator's OS seems to include in the metadata for the tarfile (it just means the origin is untrusted, and is apparently used to indicate to any program that subsequently tries to use the file that it should treat it cautiously.)  The warnings are just indicating that there are unexpected and unknown fields in the header (related to quarantine metadata), but it works fine anyway.</p>\n\n<p>In the Unix world, files starting with a period are usually hidden from directory listings.  They're used for things like configuration files in a user's home directory that people usually don't want cluttering their file lists; I suspect that may be why they appear to not exist.</p>\n\n<p>It's also possible, of course, that they really aren't there; stranger things have happened, after all.</p>",
      "rawMarkdown": "If it's still relevant: I definitely have the files beginning with a period, as well.  Also, the warnings when expanding the tar file seem to be caused by the file being marked as \"quarantined,\" which the creator's OS seems to include in the metadata for the tarfile (it just means the origin is untrusted, and is apparently used to indicate to any program that subsequently tries to use the file that it should treat it cautiously.)  The warnings are just indicating that there are unexpected and unknown fields in the header (related to quarantine metadata), but it works fine anyway.\n\nIn the Unix world, files starting with a period are usually hidden from directory listings.  They're used for things like configuration files in a user's home directory that people usually don't want cluttering their file lists; I suspect that may be why they appear to not exist.\n\nIt's also possible, of course, that they really aren't there; stranger things have happened, after all.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 204088,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "07/17/2017 15:22:33",
      "content": "<p>Yes, I have been able to download and uncompress it without any problems. It uses 5.7GB compressed and uncompressed also uses 5.7GB.</p>\n\n<p>Maybe you could try to download again the file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 204113,
          "author_name": "keesiewonder",
          "author_url": "",
          "post_date": "07/17/2017 16:12:57",
          "content": "<p>Thank you, ironbar! After playing around with it some more, I too have been able to download and uncompress all of sample.tar.gz. The warnings and errors still display when running untar(), but I think in actuality they are just threats and I have the 16 files to work with. (4 files for each of the 4 file types)</p>\n\n<p>My next question is ... Do you happen to know the difference between the files that <em>do</em> and the files that <em>do not</em> start with a \".\" in this sample download? I am able to work with 8 of the files just fine, but the 8 starting with a \"period\" character are throwing an error in my read-data function. [ i.e. filename.aps vs. .filename.aps ]</p>\n\n<p>UPDATE: I may be able to answer my own question. The one without the leading \".\" is the data file, and the one with the leading \".\" is the header file. If I'm wrong about this, please let me know!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 204779,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "07/19/2017 16:46:35",
          "content": "<p>I don't have files starting by \".\" on my uncompressed folder. </p>\n\n<p>And also reading the kernel it seems that the header is inside the file, not outside in another file. <br>\n<a href=\"https://www.kaggle.com/wcukierski/reading-images\">https://www.kaggle.com/wcukierski/reading-images</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 223075,
          "author_name": "mmiron",
          "author_url": "",
          "post_date": "09/21/2017 03:22:02",
          "content": "<p>If it's still relevant: I definitely have the files beginning with a period, as well.  Also, the warnings when expanding the tar file seem to be caused by the file being marked as \"quarantined,\" which the creator's OS seems to include in the metadata for the tarfile (it just means the origin is untrusted, and is apparently used to indicate to any program that subsequently tries to use the file that it should treat it cautiously.)  The warnings are just indicating that there are unexpected and unknown fields in the header (related to quarantine metadata), but it works fine anyway.</p>\n\n<p>In the Unix world, files starting with a period are usually hidden from directory listings.  They're used for things like configuration files in a user's home directory that people usually don't want cluttering their file lists; I suspect that may be why they appear to not exist.</p>\n\n<p>It's also possible, of course, that they really aren't there; stranger things have happened, after all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "203935": "Has anyone tried working with the provided, downloadable data file sample.tar.gz, and if yes, were you successful? \n\nI'm using R 3.4 on an up-to-date Windows 10 (sorry!) with RStudio.\n\nI downloaded sample.tar.gz and found a way, I thought, to list the files within without unzipping by  using the untar() command. I'd rather leave the file in its .tar.gz state if possible due to the its size when extracted. I am not, however, out of space on my hard disk. I checked! (767 GB free ...)\n\nuntar() does appear to start off listing contained files. For example, here's an excerpt:\n\n[17] \"sample/00360f79fd6e02781457eda48f85da90.a3d\"                                                        \n[18] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Extended header uid=457615 is out of range 0..32767\"               \n[19] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.dev'\"             \n[20] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.ino'\"             \n[21] \"C:\\\\PROGRA~2\\\\TESSER~1\\\\tar.exe: Ignoring unknown extended header keyword `SCHILY.nlink'\"           \n\nAt the end of untar()'s attempt, however, is \n\n***... tar.exe: Error exit delayed from previous errors\"                           \nattr(,\"status\")\n[1] 2\nWarning message:\nrunning command 'tar.exe -tf \"sample.tar.gz\"' had status 2 ***\n\nI haven't entirely figured out what this means yet, but it doesn't  all sound good. I have been able to unzip a specific .aps file within the tar.gz, but would like to be able to see a list of all the contained *.aps files. If untar()'s list is failing before it finishes listing everything, I don't know what I'm missing. I tried specifying a wildcard to untar() using the file argument, but that didn't do anything. Perhaps I don't have the syntax right ...\n\nuntar(\"sample.tar.gz\", files = \"sample/*.aps\", list = FALSE, exdir = \".\", compressed = NA, extras = NULL, verbose = TRUE, restore_times =  TRUE, tar = Sys.getenv(\"TAR\"))\n\nI know there's a Google Cloud area for accessing the data files, and I will eventually graduate to that. For now, I'd really like to just start with the sample.tar.gz, if possible.\n\nThanks in advance for your thoughts on how I might be able to get around the above mentioned errors and/or warnings.",
    "204088": "Yes, I have been able to download and uncompress it without any problems. It uses 5.7GB compressed and uncompressed also uses 5.7GB.\n\nMaybe you could try to download again the file.",
    "204113": "Thank you, ironbar! After playing around with it some more, I too have been able to download and uncompress all of sample.tar.gz. The warnings and errors still display when running untar(), but I think in actuality they are just threats and I have the 16 files to work with. (4 files for each of the 4 file types)\n\nMy next question is ... Do you happen to know the difference between the files that *do* and the files that *do not* start with a \".\" in this sample download? I am able to work with 8 of the files just fine, but the 8 starting with a \"period\" character are throwing an error in my read-data function. [ i.e. filename.aps vs. .filename.aps ]\n\nUPDATE: I may be able to answer my own question. The one without the leading \".\" is the data file, and the one with the leading \".\" is the header file. If I'm wrong about this, please let me know!",
    "204779": "I don't have files starting by \".\" on my uncompressed folder. \n\nAnd also reading the kernel it seems that the header is inside the file, not outside in another file.   \nhttps://www.kaggle.com/wcukierski/reading-images",
    "223075": "If it's still relevant: I definitely have the files beginning with a period, as well.  Also, the warnings when expanding the tar file seem to be caused by the file being marked as \"quarantined,\" which the creator's OS seems to include in the metadata for the tarfile (it just means the origin is untrusted, and is apparently used to indicate to any program that subsequently tries to use the file that it should treat it cautiously.)  The warnings are just indicating that there are unexpected and unknown fields in the header (related to quarantine metadata), but it works fine anyway.\n\nIn the Unix world, files starting with a period are usually hidden from directory listings.  They're used for things like configuration files in a user's home directory that people usually don't want cluttering their file lists; I suspect that may be why they appear to not exist.\n\nIt's also possible, of course, that they really aren't there; stranger things have happened, after all."
  },
  "source": "meta"
}