{
  "id": 56787,
  "title": "Training files unavailable in kernels?",
  "url": "/competitions/trackml-particle-identification/discussion/56787",
  "author_name": "",
  "post_date": "2018-05-15T01:27:17.711085900Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I'm attempting to use data from train_2, train_3, etc. in the same way I've used data from train 1, but when working in a kernel I'm getting errors, it is like the .zip is empty for every file except train_1.</p>\n\n<p>I am able to run code with train_1, and am loading in the provided library. Looking at train_2 locally, it appears to be the same as train_1.</p>\n\n<p>The following code snippets are from public notebooks that fail when attempting to use alternative zip files:</p>\n\n<hr>\n\n<blockquote>\n  <p>path_to_train = \"../input/train_\" event_prefix = \"event000001000\"\n  hits, cells, particles, truth = load_event(os.path.join(path_to_train,\n  event_prefix))</p>\n</blockquote>\n\n<p><em>from kNN approach by Mikhail Hushchyn</em></p>\n\n<hr>\n\n<blockquote>\n  <p>train = np.unique([p.split('-')[0] for p in sorted(glob.glob('../input/train_4/**'))])</p>\n</blockquote>\n\n<p><em>from The Martian by the1owl</em></p>\n\n<hr>\n\n<p>Any help appreciated. </p>",
  "messages": [
    {
      "id": "328736",
      "postDate": "05/15/2018 01:27:17",
      "content": "<p>I'm attempting to use data from train_2, train_3, etc. in the same way I've used data from train 1, but when working in a kernel I'm getting errors, it is like the .zip is empty for every file except train_1.</p>\n\n<p>I am able to run code with train_1, and am loading in the provided library. Looking at train_2 locally, it appears to be the same as train_1.</p>\n\n<p>The following code snippets are from public notebooks that fail when attempting to use alternative zip files:</p>\n\n<hr>\n\n<blockquote>\n  <p>path_to_train = \"../input/train_\" event_prefix = \"event000001000\"\n  hits, cells, particles, truth = load_event(os.path.join(path_to_train,\n  event_prefix))</p>\n</blockquote>\n\n<p><em>from kNN approach by Mikhail Hushchyn</em></p>\n\n<hr>\n\n<blockquote>\n  <p>train = np.unique([p.split('-')[0] for p in sorted(glob.glob('../input/train_4/**'))])</p>\n</blockquote>\n\n<p><em>from The Martian by the1owl</em></p>\n\n<hr>\n\n<p>Any help appreciated. </p>",
      "rawMarkdown": "I'm attempting to use data from train_2, train_3, etc. in the same way I've used data from train 1, but when working in a kernel I'm getting errors, it is like the .zip is empty for every file except train_1.\n\nI am able to run code with train_1, and am loading in the provided library. Looking at train_2 locally, it appears to be the same as train_1.\n\nThe following code snippets are from public notebooks that fail when attempting to use alternative zip files:\n\n\n----------\n\n\n&gt; path_to_train = \"../input/train_\" event_prefix = \"event000001000\"\n&gt; hits, cells, particles, truth = load_event(os.path.join(path_to_train,\n&gt; event_prefix))\n\n*from kNN approach by Mikhail Hushchyn*\n\n\n----------\n\n\n&gt; train = np.unique([p.split('-')[0] for p in sorted(glob.glob('../input/train_4/**'))])\n\n *from The Martian by the1owl*\n\n\n----------\n\n\nAny help appreciated.",
      "votes": null
    },
    {
      "id": "329671",
      "postDate": "05/16/2018 23:55:28",
      "content": "<p>Sounds like a bug! I've passed this on to the engineering team. Thanks for reporting.</p>",
      "rawMarkdown": "Sounds like a bug! I've passed this on to the engineering team. Thanks for reporting.",
      "votes": null
    },
    {
      "id": "352747",
      "postDate": "07/05/2018 02:54:19",
      "content": "<p>Is this question been answered yet?  I observed the same thing as @macfaII.</p>",
      "rawMarkdown": "Is this question been answered yet?  I observed the same thing as @macfaII.",
      "votes": null
    },
    {
      "id": "352750",
      "postDate": "07/05/2018 03:03:09",
      "content": "<p>I don't think it has been addressed but the last time I checked the issue was that the additional files aren't available on the kernel. You can list available files and they aren't there...</p>\n\n<p>You can either work with the files locally or just use train 1, there is plenty of data and in a different thread the organizers claimed each set should be similar / unbiased.</p>",
      "rawMarkdown": "I don't think it has been addressed but the last time I checked the issue was that the additional files aren't available on the kernel. You can list available files and they aren't there...\n\nYou can either work with the files locally or just use train 1, there is plenty of data and in a different thread the organizers claimed each set should be similar / unbiased.",
      "votes": null
    },
    {
      "id": "352751",
      "postDate": "07/05/2018 03:07:19",
      "content": "<blockquote>\n  <p>Note: Due to the number and size of the data files, we've only included the files from \"train_1\" in the Kernels environment (along with the full Test set, etc.)</p>\n</blockquote>\n\n<p>in the welcome thread</p>",
      "rawMarkdown": "&gt; Note: Due to the number and size of the data files, we've only included the files from \"train_1\" in the Kernels environment (along with the full Test set, etc.)\n\nin the welcome thread",
      "votes": null
    },
    {
      "id": "352791",
      "postDate": "07/05/2018 06:00:11",
      "content": "<p>@macfaII thanks. You are right train_1 is good enough but I am curious as to how my models will perform on a different data. </p>",
      "rawMarkdown": "macfaII thanks. You are right train_1 is good enough but I am curious as to how my models will perform on a different data.",
      "votes": null
    },
    {
      "id": "352945",
      "postDate": "07/05/2018 13:51:46",
      "content": "<p>... train_1 2 etc... all randomly extracted from the same master dataset. Your model should behave exactly the same, except for statistical variance.</p>",
      "rawMarkdown": "... train_1 2 etc... all randomly extracted from the same master dataset. Your model should behave exactly the same, except for statistical variance.",
      "votes": null
    },
    {
      "id": "353079",
      "postDate": "07/05/2018 21:00:14",
      "content": "<p>Merci beaucoup, monsieur David Rousseau.</p>",
      "rawMarkdown": "Merci beaucoup, monsieur David Rousseau.",
      "votes": null
    },
    {
      "id": "355195",
      "postDate": "07/11/2018 07:00:30",
      "content": "<p><strong>UPDATE</strong></p>\n\n<p>I have downloaded train_2.zip and observed the following:-</p>\n\n<p>For anyone else having the same issue, the reason kernels do not recognize the files in train_{2,3,4,5}.zip is not because they are empty. I believe there is some kind of invisible characters in the file names. Once I renamed them to the same names by retyping, they are readable. I had the same issue reading them locally until I did the renaming.</p>",
      "rawMarkdown": "**UPDATE**\n\nI have downloaded train_2.zip and observed the following:-\n\nFor anyone else having the same issue, the reason kernels do not recognize the files in train_{2,3,4,5}.zip is not because they are empty. I believe there is some kind of invisible characters in the file names. Once I renamed them to the same names by retyping, they are readable. I had the same issue reading them locally until I did the renaming.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 329671,
      "author_name": "colinmorris",
      "author_url": "",
      "post_date": "05/16/2018 23:55:28",
      "content": "<p>Sounds like a bug! I've passed this on to the engineering team. Thanks for reporting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 352747,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "07/05/2018 02:54:19",
          "content": "<p>Is this question been answered yet?  I observed the same thing as @macfaII.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352750,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "07/05/2018 03:03:09",
          "content": "<p>I don't think it has been addressed but the last time I checked the issue was that the additional files aren't available on the kernel. You can list available files and they aren't there...</p>\n\n<p>You can either work with the files locally or just use train 1, there is plenty of data and in a different thread the organizers claimed each set should be similar / unbiased.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352791,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "07/05/2018 06:00:11",
          "content": "<p>@macfaII thanks. You are right train_1 is good enough but I am curious as to how my models will perform on a different data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 352945,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/05/2018 13:51:46",
          "content": "<p>... train_1 2 etc... all randomly extracted from the same master dataset. Your model should behave exactly the same, except for statistical variance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 353079,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "07/05/2018 21:00:14",
          "content": "<p>Merci beaucoup, monsieur David Rousseau.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355195,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "07/11/2018 07:00:30",
          "content": "<p><strong>UPDATE</strong></p>\n\n<p>I have downloaded train_2.zip and observed the following:-</p>\n\n<p>For anyone else having the same issue, the reason kernels do not recognize the files in train_{2,3,4,5}.zip is not because they are empty. I believe there is some kind of invisible characters in the file names. Once I renamed them to the same names by retyping, they are readable. I had the same issue reading them locally until I did the renaming.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 352751,
      "author_name": "outrunner",
      "author_url": "",
      "post_date": "07/05/2018 03:07:19",
      "content": "<blockquote>\n  <p>Note: Due to the number and size of the data files, we've only included the files from \"train_1\" in the Kernels environment (along with the full Test set, etc.)</p>\n</blockquote>\n\n<p>in the welcome thread</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328736": "I'm attempting to use data from train_2, train_3, etc. in the same way I've used data from train 1, but when working in a kernel I'm getting errors, it is like the .zip is empty for every file except train_1.\n\nI am able to run code with train_1, and am loading in the provided library. Looking at train_2 locally, it appears to be the same as train_1.\n\nThe following code snippets are from public notebooks that fail when attempting to use alternative zip files:\n\n\n----------\n\n\n&gt; path_to_train = \"../input/train_\" event_prefix = \"event000001000\"\n&gt; hits, cells, particles, truth = load_event(os.path.join(path_to_train,\n&gt; event_prefix))\n\n*from kNN approach by Mikhail Hushchyn*\n\n\n----------\n\n\n&gt; train = np.unique([p.split('-')[0] for p in sorted(glob.glob('../input/train_4/**'))])\n\n *from The Martian by the1owl*\n\n\n----------\n\n\nAny help appreciated.",
    "329671": "Sounds like a bug! I've passed this on to the engineering team. Thanks for reporting.",
    "352747": "Is this question been answered yet?  I observed the same thing as @macfaII.",
    "352750": "I don't think it has been addressed but the last time I checked the issue was that the additional files aren't available on the kernel. You can list available files and they aren't there...\n\nYou can either work with the files locally or just use train 1, there is plenty of data and in a different thread the organizers claimed each set should be similar / unbiased.",
    "352751": "&gt; Note: Due to the number and size of the data files, we've only included the files from \"train_1\" in the Kernels environment (along with the full Test set, etc.)\n\nin the welcome thread",
    "352791": "macfaII thanks. You are right train_1 is good enough but I am curious as to how my models will perform on a different data.",
    "352945": "... train_1 2 etc... all randomly extracted from the same master dataset. Your model should behave exactly the same, except for statistical variance.",
    "353079": "Merci beaucoup, monsieur David Rousseau.",
    "355195": "**UPDATE**\n\nI have downloaded train_2.zip and observed the following:-\n\nFor anyone else having the same issue, the reason kernels do not recognize the files in train_{2,3,4,5}.zip is not because they are empty. I believe there is some kind of invisible characters in the file names. Once I renamed them to the same names by retyping, they are readable. I had the same issue reading them locally until I did the renaming."
  },
  "source": "meta"
}