{
  "id": 180491,
  "title": "extra data on level5 website?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/180491",
  "author_name": "cj",
  "post_date": "2020-09-05T09:00:44.634000",
  "votes": 6,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi, I came across this page on the level5 site. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2048731%2F509abbb69edb130c783a87f0449f1142%2Flyft-level5-site.png?generation=1599296306729582&amp;alt=media\" alt=\"\"></p>\n<p>Is there even more training data available on level5 site than what is available through kaggle? The training set part2 download is 70 GB. Are we permitted to use this too?</p>",
  "messages": [
    {
      "id": 999000,
      "postDate": "2020-09-05T09:00:44.633Z",
      "content": "<p>Hi, I came across this page on the level5 site. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2048731%2F509abbb69edb130c783a87f0449f1142%2Flyft-level5-site.png?generation=1599296306729582&amp;alt=media\" alt=\"\"></p>\n<p>Is there even more training data available on level5 site than what is available through kaggle? The training set part2 download is 70 GB. Are we permitted to use this too?</p>",
      "rawMarkdown": "Hi, I came across this page on the level5 site. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2048731%2F509abbb69edb130c783a87f0449f1142%2Flyft-level5-site.png?generation=1599296306729582&alt=media)\n\nIs there even more training data available on level5 site than what is available through kaggle? The training set part2 download is 70 GB. Are we permitted to use this too?\n",
      "votes": 6
    },
    {
      "id": 1003092,
      "postDate": "2020-09-08T16:50:42.447Z",
      "content": "<p>Just to clarify for future kagglers:<br>\nData on Kaggle match the ones on Lyft's website, so no problems in using those as they are the same :) </p>",
      "rawMarkdown": "Just to clarify for future kagglers:\nData on Kaggle match the ones on Lyft's website, so no problems in using those as they are the same :) ",
      "votes": 1,
      "replies": [
        {
          "id": 1003101,
          "postDate": "2020-09-08T16:59:25.460Z",
          "content": "<p>I haven't invested too much time investigating the differences, but going by the raw size of the sets, how can they be the same?<br>\nKaggle: 22 GB<br>\nLyft: 86 GB</p>",
          "rawMarkdown": "I haven't invested too much time investigating the differences, but going by the raw size of the sets, how can they be the same?\nKaggle: 22 GB\nLyft: 86 GB"
        },
        {
          "id": 1003129,
          "postDate": "2020-09-08T17:19:14.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> You can find the \"missing\" data <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "@ilu000 You can find the \"missing\" data [here](https://www.kaggle.com/philculliton/lyft-full-training-set)",
          "votes": 2
        },
        {
          "id": 1003176,
          "postDate": "2020-09-08T17:54:09.563Z",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , yeah I didn't add that to the list as I wasn't sure about possible overlaps.<br>\nkaggle full_train: 71.7 GB</p>\n<p>so, there is still some overlap. And i'd love to know where from the hosts (@lucabergamini ), so I don't have to investigate myself ;)</p>\n<p>Kaggle train/val/test + kaggle full train = ~94 GB<br>\nKaggle only train + kaggle full train = ~86 GB? <br>\nLyft: 86 GB</p>\n<p>so is Kaggle train + kaggle full train the same as the data on the lyft website including validation? If so, what is validation on kaggle?</p>",
          "rawMarkdown": "thanks @pestipeti , yeah I didn't add that to the list as I wasn't sure about possible overlaps.\nkaggle full_train: 71.7 GB\n\nso, there is still some overlap. And i'd love to know where from the hosts (@lucabergamini ), so I don't have to investigate myself ;)\n\nKaggle train/val/test + kaggle full train = ~94 GB\nKaggle only train + kaggle full train = ~86 GB? \nLyft: 86 GB\n\nso is Kaggle train + kaggle full train the same as the data on the lyft website including validation? If so, what is validation on kaggle?"
        },
        {
          "id": 1003202,
          "postDate": "2020-09-08T18:26:42.557Z",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> should be kaggle full train + kaggle val = lyft dataset?</p>",
          "rawMarkdown": "@ilu000 should be kaggle full train + kaggle val = lyft dataset?"
        },
        {
          "id": 1003206,
          "postDate": "2020-09-08T18:31:31.323Z",
          "content": "<pre><code>Kaggle validate        9 GB\nKaggle full_train   71.7 GB\n---------------------------\nTotal:              80.7 GB\n\nLyft dataset including validation = 86.6 GB\n</code></pre>\n<p>It's close but not equal.</p>",
          "rawMarkdown": "```\nKaggle validate        9 GB\nKaggle full_train   71.7 GB\n---------------------------\nTotal:              80.7 GB\n\nLyft dataset including validation = 86.6 GB\n```\nIt's close but not equal."
        },
        {
          "id": 1003254,
          "postDate": "2020-09-08T19:15:39.920Z",
          "content": "<p>Update to release 1.1 of the dataset:</p>\n<p><strong>Lyft Website</strong></p>\n<ul>\n<li>train</li>\n<li>train_full</li>\n<li>validation</li>\n</ul>\n<p><strong>Kaggle</strong></p>\n<ul>\n<li>train</li>\n<li>train_full as additional data</li>\n<li>validation</li>\n</ul>",
          "rawMarkdown": "Update to release 1.1 of the dataset:\n\n**Lyft Website**\n- train\n- train_full\n- validation\n\n**Kaggle**\n- train\n- train_full as additional data\n- validation",
          "votes": 1
        },
        {
          "id": 1003260,
          "postDate": "2020-09-08T19:19:01.847Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> </p>\n<blockquote>\n  <p>Validation is not included on Lyft's website</p>\n</blockquote>\n<p>What is validation dataset (8.2 GB) on Lyft website then?</p>",
          "rawMarkdown": "Thanks a lot @lucabergamini \n\n> Validation is not included on Lyft's website\n\nWhat is validation dataset (8.2 GB) on Lyft website then?"
        },
        {
          "id": 1003283,
          "postDate": "2020-09-08T19:37:36.120Z",
          "content": "<p>Ah I should have checked it! We changed it in the last release. I'll edit my answer!</p>",
          "rawMarkdown": "Ah I should have checked it! We changed it in the last release. I'll edit my answer!",
          "votes": 1
        },
        {
          "id": 1003294,
          "postDate": "2020-09-08T19:47:20.793Z",
          "content": "<p>Thank you for the clarification <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> , highly appreciated!</p>\n<p>So:<br>\nLyft train part 1 is train on kaggle<br>\nLyft train part 2 is train_full on kaggle<br>\nLyft validation is validation on kaggle</p>\n<p>edit: is sample.zarr included in the training set?<br>\nand despite the name, train_full is actually not the full training set? it's just part 2 (admittedly the much larger part) of the training set?</p>",
          "rawMarkdown": "Thank you for the clarification @lucabergamini , highly appreciated!\n\nSo:\nLyft train part 1 is train on kaggle\nLyft train part 2 is train_full on kaggle\nLyft validation is validation on kaggle\n\nedit: is sample.zarr included in the training set?\nand despite the name, train_full is actually not the full training set? it's just part 2 (admittedly the much larger part) of the training set?",
          "votes": 1
        },
        {
          "id": 1003766,
          "postDate": "2020-09-09T08:25:45.730Z",
          "content": "<p><code>sample.zarr</code> is just a subset of train part 1 (first 100 scenes I think), so train_1 also includes scenes from <code>sample.zarr</code>.<br>\n<code>train_full</code> is the full train set actually (train part 1 is the first part of it).</p>\n<p>we should probably clarify it even further somewhere as can I see people are getting a little bit confused by this</p>",
          "rawMarkdown": "`sample.zarr` is just a subset of train part 1 (first 100 scenes I think), so train_1 also includes scenes from `sample.zarr`.\n`train_full` is the full train set actually (train part 1 is the first part of it).\n\nwe should probably clarify it even further somewhere as can I see people are getting a little bit confused by this",
          "votes": 2
        },
        {
          "id": 1004308,
          "postDate": "2020-09-09T16:00:32.550Z",
          "content": "<p>yes, complete clarification would be great. <br>\nSpecifically, is part1 (lyft website) included in part2 (lyft website)? Is part2 (lyft website) the same as train_full (kaggle dataset)?</p>",
          "rawMarkdown": "yes, complete clarification would be great. \nSpecifically, is part1 (lyft website) included in part2 (lyft website)? Is part2 (lyft website) the same as train_full (kaggle dataset)?"
        },
        {
          "id": 1004401,
          "postDate": "2020-09-09T17:32:50.287Z",
          "content": "<p>I'm just wondering what the best place for that would be….</p>\n<p>In the meantime:</p>\n<ul>\n<li>part1 is included in part2</li>\n<li>part2 is the same as train_full</li>\n<li>sample is included in part1 (and therefore in part2)</li>\n</ul>",
          "rawMarkdown": "I'm just wondering what the best place for that would be....\n\nIn the meantime:\n- part1 is included in part2\n- part2 is the same as train_full\n- sample is included in part1 (and therefore in part2)",
          "votes": 2
        },
        {
          "id": 1004605,
          "postDate": "2020-09-09T20:56:59.557Z",
          "content": "<p>\"part1 is included in part2\" is a surprise))</p>",
          "rawMarkdown": "\"part1 is included in part2\" is a surprise))",
          "votes": 1
        },
        {
          "id": 1004623,
          "postDate": "2020-09-09T21:25:40.567Z",
          "content": "<p>yeah, somewhat misleading wording here</p>",
          "rawMarkdown": "yeah, somewhat misleading wording here"
        },
        {
          "id": 1004678,
          "postDate": "2020-09-09T23:39:27.983Z",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> Yes I second <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> the numbers are bit confusing. Grateful for the dataset and competition but even on the published papers numbers don't add up. The claim is there are 170k scene. But tabular description on paper says 162k scene. Only thing constant is hrs and Km.Think we need to find out for ourselves. I recommend a separate thread created by host to clarify size and the number of scene.   </p>",
          "rawMarkdown": "@lucabergamini Yes I second @zaharch and @ilu000 the numbers are bit confusing. Grateful for the dataset and competition but even on the published papers numbers don't add up. The claim is there are 170k scene. But tabular description on paper says 162k scene. Only thing constant is hrs and Km.Think we need to find out for ourselves. I recommend a separate thread created by host to clarify size and the number of scene.   ",
          "votes": 1
        },
        {
          "id": 1005247,
          "postDate": "2020-09-10T10:59:52.600Z",
          "content": "<p>We re-sourced the dataset between the paper release and the competition start to include also traffic lights. We tried to retain the original splits, but it was not trivial with more than 100k scenes. </p>\n<p>I'll start a post to clarify the current structure :) </p>",
          "rawMarkdown": "We re-sourced the dataset between the paper release and the competition start to include also traffic lights. We tried to retain the original splits, but it was not trivial with more than 100k scenes. \n\nI'll start a post to clarify the current structure :) ",
          "votes": 2
        }
      ]
    },
    {
      "id": 999216,
      "postDate": "2020-09-05T13:38:16.057Z",
      "content": "<p>This should be the one mentioned under the \"Additional Files\" paragraph in the \"Data\"-Tab, also hosted on <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">kaggle</a><br>\nGiven that it is mentioned under the Data-Tab, I would certainly assume that we can use it for training.</p>",
      "rawMarkdown": "This should be the one mentioned under the \"Additional Files\" paragraph in the \"Data\"-Tab, also hosted on [kaggle](https://www.kaggle.com/philculliton/lyft-full-training-set)\nGiven that it is mentioned under the Data-Tab, I would certainly assume that we can use it for training.",
      "votes": 1,
      "replies": [
        {
          "id": 999373,
          "postDate": "2020-09-05T15:43:45.617Z",
          "content": "<p>wow, I completely missed that. Thank you <a href=\"https://www.kaggle.com/hanshaas\" target=\"_blank\">@hanshaas</a> </p>",
          "rawMarkdown": "wow, I completely missed that. Thank you @hanshaas ",
          "votes": 1
        }
      ]
    },
    {
      "id": 999365,
      "postDate": "2020-09-05T15:39:41.097Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1003092,
      "author_name": "Luca Bergamini",
      "author_url": "",
      "post_date": "2020-09-08T16:50:42.447000",
      "content": "<p>Just to clarify for future kagglers:<br>\nData on Kaggle match the ones on Lyft's website, so no problems in using those as they are the same :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1003101,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-08T16:59:25.460000",
          "content": "<p>I haven't invested too much time investigating the differences, but going by the raw size of the sets, how can they be the same?<br>\nKaggle: 22 GB<br>\nLyft: 86 GB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003129,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-09-08T17:19:14.137000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> You can find the \"missing\" data <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1003176,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-08T17:54:09.563000",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> , yeah I didn't add that to the list as I wasn't sure about possible overlaps.<br>\nkaggle full_train: 71.7 GB</p>\n<p>so, there is still some overlap. And i'd love to know where from the hosts (@lucabergamini ), so I don't have to investigate myself ;)</p>\n<p>Kaggle train/val/test + kaggle full train = ~94 GB<br>\nKaggle only train + kaggle full train = ~86 GB? <br>\nLyft: 86 GB</p>\n<p>so is Kaggle train + kaggle full train the same as the data on the lyft website including validation? If so, what is validation on kaggle?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003202,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-09-08T18:26:42.557000",
          "content": "<p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> should be kaggle full train + kaggle val = lyft dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003206,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-08T18:31:31.323000",
          "content": "<pre><code>Kaggle validate        9 GB\nKaggle full_train   71.7 GB\n---------------------------\nTotal:              80.7 GB\n\nLyft dataset including validation = 86.6 GB\n</code></pre>\n<p>It's close but not equal.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003254,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-08T19:15:39.920000",
          "content": "<p>Update to release 1.1 of the dataset:</p>\n<p><strong>Lyft Website</strong></p>\n<ul>\n<li>train</li>\n<li>train_full</li>\n<li>validation</li>\n</ul>\n<p><strong>Kaggle</strong></p>\n<ul>\n<li>train</li>\n<li>train_full as additional data</li>\n<li>validation</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1003260,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-08T19:19:01.847000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> </p>\n<blockquote>\n  <p>Validation is not included on Lyft's website</p>\n</blockquote>\n<p>What is validation dataset (8.2 GB) on Lyft website then?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1003283,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-08T19:37:36.120000",
          "content": "<p>Ah I should have checked it! We changed it in the last release. I'll edit my answer!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1003294,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-08T19:47:20.793000",
          "content": "<p>Thank you for the clarification <a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> , highly appreciated!</p>\n<p>So:<br>\nLyft train part 1 is train on kaggle<br>\nLyft train part 2 is train_full on kaggle<br>\nLyft validation is validation on kaggle</p>\n<p>edit: is sample.zarr included in the training set?<br>\nand despite the name, train_full is actually not the full training set? it's just part 2 (admittedly the much larger part) of the training set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1003766,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-09T08:25:45.730000",
          "content": "<p><code>sample.zarr</code> is just a subset of train part 1 (first 100 scenes I think), so train_1 also includes scenes from <code>sample.zarr</code>.<br>\n<code>train_full</code> is the full train set actually (train part 1 is the first part of it).</p>\n<p>we should probably clarify it even further somewhere as can I see people are getting a little bit confused by this</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1004308,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-09T16:00:32.550000",
          "content": "<p>yes, complete clarification would be great. <br>\nSpecifically, is part1 (lyft website) included in part2 (lyft website)? Is part2 (lyft website) the same as train_full (kaggle dataset)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1004401,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-09T17:32:50.287000",
          "content": "<p>I'm just wondering what the best place for that would be….</p>\n<p>In the meantime:</p>\n<ul>\n<li>part1 is included in part2</li>\n<li>part2 is the same as train_full</li>\n<li>sample is included in part1 (and therefore in part2)</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1004605,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-09-09T20:56:59.557000",
          "content": "<p>\"part1 is included in part2\" is a surprise))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1004623,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2020-09-09T21:25:40.567000",
          "content": "<p>yeah, somewhat misleading wording here</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1004678,
          "author_name": "The Brown Iceman",
          "author_url": "",
          "post_date": "2020-09-09T23:39:27.983000",
          "content": "<p><a href=\"https://www.kaggle.com/lucabergamini\" target=\"_blank\">@lucabergamini</a> Yes I second <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> and <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> the numbers are bit confusing. Grateful for the dataset and competition but even on the published papers numbers don't add up. The claim is there are 170k scene. But tabular description on paper says 162k scene. Only thing constant is hrs and Km.Think we need to find out for ourselves. I recommend a separate thread created by host to clarify size and the number of scene.   </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1005247,
          "author_name": "Luca Bergamini",
          "author_url": "",
          "post_date": "2020-09-10T10:59:52.600000",
          "content": "<p>We re-sourced the dataset between the paper release and the competition start to include also traffic lights. We tried to retain the original splits, but it was not trivial with more than 100k scenes. </p>\n<p>I'll start a post to clarify the current structure :) </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 999216,
      "author_name": "Johann Haas",
      "author_url": "",
      "post_date": "2020-09-05T13:38:16.057000",
      "content": "<p>This should be the one mentioned under the \"Additional Files\" paragraph in the \"Data\"-Tab, also hosted on <a href=\"https://www.kaggle.com/philculliton/lyft-full-training-set\" target=\"_blank\">kaggle</a><br>\nGiven that it is mentioned under the Data-Tab, I would certainly assume that we can use it for training.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 999373,
          "author_name": "cj",
          "author_url": "",
          "post_date": "2020-09-05T15:43:45.617000",
          "content": "<p>wow, I completely missed that. Thank you <a href=\"https://www.kaggle.com/hanshaas\" target=\"_blank\">@hanshaas</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 999365,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-05T15:39:41.097000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "999000": "Hi, I came across this page on the level5 site. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2048731%2F509abbb69edb130c783a87f0449f1142%2Flyft-level5-site.png?generation=1599296306729582&alt=media)\n\nIs there even more training data available on level5 site than what is available through kaggle? The training set part2 download is 70 GB. Are we permitted to use this too?\n",
    "1003092": "Just to clarify for future kagglers:\nData on Kaggle match the ones on Lyft's website, so no problems in using those as they are the same :) ",
    "999216": "This should be the one mentioned under the \"Additional Files\" paragraph in the \"Data\"-Tab, also hosted on [kaggle](https://www.kaggle.com/philculliton/lyft-full-training-set)\nGiven that it is mentioned under the Data-Tab, I would certainly assume that we can use it for training.",
    "999365": ""
  }
}