{
  "id": 269439,
  "title": "Stratified tfrecords",
  "url": "/competitions/g2net-gravitational-wave-detection/discussion/269439",
  "author_name": "",
  "post_date": "2021-08-31T17:17:36.787859600Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey Kagglers! ;-) Are you looking for a bunch of tfrecords that were stratified on the target values? Then these datasets might be helpful:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/allunia/1-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/1-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/2-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/2-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/3-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/3-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/4-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/4-g2net-data-balanced</a></li>\n</ul>\n<p>As this data is quite noisy I hope that this further balance is helpful to find a stable CV. I have split them as they took too much space to create all tfrecords at once. I'm not able to download the data at the moment and haven't found a way to combine all of them in one dataset only using the kaggle platform. If you know a way please explain it to me. ;-)</p>",
  "messages": [
    {
      "id": "1498125",
      "postDate": "08/31/2021 17:17:36",
      "content": "<p>Hey Kagglers! ;-) Are you looking for a bunch of tfrecords that were stratified on the target values? Then these datasets might be helpful:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/allunia/1-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/1-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/2-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/2-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/3-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/3-g2net-data-balanced</a></li>\n<li><a href=\"https://www.kaggle.com/allunia/4-g2net-data-balanced\" target=\"_blank\">https://www.kaggle.com/allunia/4-g2net-data-balanced</a></li>\n</ul>\n<p>As this data is quite noisy I hope that this further balance is helpful to find a stable CV. I have split them as they took too much space to create all tfrecords at once. I'm not able to download the data at the moment and haven't found a way to combine all of them in one dataset only using the kaggle platform. If you know a way please explain it to me. ;-)</p>",
      "rawMarkdown": "Hey Kagglers! ;-) Are you looking for a bunch of tfrecords that were stratified on the target values? Then these datasets might be helpful:\n\n- https://www.kaggle.com/allunia/1-g2net-data-balanced\n- https://www.kaggle.com/allunia/2-g2net-data-balanced\n- https://www.kaggle.com/allunia/3-g2net-data-balanced\n- https://www.kaggle.com/allunia/4-g2net-data-balanced\n\nAs this data is quite noisy I hope that this further balance is helpful to find a stable CV. I have split them as they took too much space to create all tfrecords at once. I'm not able to download the data at the moment and haven't found a way to combine all of them in one dataset only using the kaggle platform. If you know a way please explain it to me. ;-)",
      "votes": null
    },
    {
      "id": "1498408",
      "postDate": "09/01/2021 00:56:11",
      "content": "<p><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> I've read there are performance benefit from having multiple TFrecord files. This is mentioned in a few articles  but I can't find the original one I read which was very clear. </p>\n<p>From memory the benefit are</p>\n<ul>\n<li>Very large files can be slower to read.</li>\n<li>With multiple TFrecord files you can parallelize the reads </li>\n</ul>\n<p>From memory rule of thumb was to have the TFrecords under a under 500MB and have at least 10. </p>\n<p>I've used this approach in my work and get fantastic performance with the TPU. </p>\n<p>However I'm a newbie so take this as unverified. </p>",
      "rawMarkdown": "allunia I've read there are performance benefit from having multiple TFrecord files. This is mentioned in a few articles  but I can't find the original one I read which was very clear. \n\nFrom memory the benefit are\n* Very large files can be slower to read.\n* With multiple TFrecord files you can parallelize the reads \n\nFrom memory rule of thumb was to have the TFrecords under a under 500MB and have at least 10. \n\nI've used this approach in my work and get fantastic performance with the TPU. \n\nHowever I'm a newbie so take this as unverified.",
      "votes": null
    },
    {
      "id": "1498722",
      "postDate": "09/01/2021 06:44:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kevinmcisaac\" target=\"_blank\">@kevinmcisaac</a> , </p>\n<p>thank you for pointing it out. So far I haven't tried it to use even more tfrecords to improve my performance, I just split them such that they are somehow \"easy to use\" for cross validation. Do you have a resource where I can read more about it? As my TPU quota is already consumed and I don't want to use colab for this competition at the moment, do you have any further tips for improving the performance? I already used mixed precision. </p>\n<p>My motivation for stratification on the target was simply that this data is extremely noisy. When looking at my model predictions I can really say that there are sometimes predictions like 0.89 with one model and 0.15 with another by only changing a bit of \"randomness\". They had the same architecture, training procedure etc.. This is really bad, the variance is really high. So my hope was that it can be helpful to find some balance in every step I can find. And this includes the stratification on the targets.</p>",
      "rawMarkdown": "Hi @kevinmcisaac , \n\nthank you for pointing it out. So far I haven't tried it to use even more tfrecords to improve my performance, I just split them such that they are somehow \"easy to use\" for cross validation. Do you have a resource where I can read more about it? As my TPU quota is already consumed and I don't want to use colab for this competition at the moment, do you have any further tips for improving the performance? I already used mixed precision. \n\nMy motivation for stratification on the target was simply that this data is extremely noisy. When looking at my model predictions I can really say that there are sometimes predictions like 0.89 with one model and 0.15 with another by only changing a bit of \"randomness\". They had the same architecture, training procedure etc.. This is really bad, the variance is really high. So my hope was that it can be helpful to find some balance in every step I can find. And this includes the stratification on the targets.",
      "votes": null
    },
    {
      "id": "1498735",
      "postDate": "09/01/2021 06:59:25",
      "content": "<p>You can find my public data sets for these TFRs</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass</a></li>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii</a></li>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions</a></li>\n</ul>\n<p>My set up is Keras/TF and I use</p>\n<ul>\n<li>Continusious Wavelets Transforms (CTW) from <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">here</a> </li>\n<li>Efficientnet. Generally B0 for speed while testing things.<br>\nWith a TPU and a CTW that is 128x128 an epoch takes about 3min.</li>\n</ul>",
      "rawMarkdown": "You can find my public data sets for these TFRs\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions\n\nMy set up is Keras/TF and I use\n* Continusious Wavelets Transforms (CTW) from [here](https://github.com/Kevin-McIsaac/cmorlet-tensorflow) \n* Efficientnet. Generally B0 for speed while testing things.\nWith a TPU and a CTW that is 128x128 an epoch takes about 3min.",
      "votes": null
    },
    {
      "id": "1559950",
      "postDate": "10/27/2021 08:36:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "rawMarkdown": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1498408,
      "author_name": "kevinmcisaac",
      "author_url": "",
      "post_date": "09/01/2021 00:56:11",
      "content": "<p><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a> I've read there are performance benefit from having multiple TFrecord files. This is mentioned in a few articles  but I can't find the original one I read which was very clear. </p>\n<p>From memory the benefit are</p>\n<ul>\n<li>Very large files can be slower to read.</li>\n<li>With multiple TFrecord files you can parallelize the reads </li>\n</ul>\n<p>From memory rule of thumb was to have the TFrecords under a under 500MB and have at least 10. </p>\n<p>I've used this approach in my work and get fantastic performance with the TPU. </p>\n<p>However I'm a newbie so take this as unverified. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1498722,
          "author_name": "allunia",
          "author_url": "",
          "post_date": "09/01/2021 06:44:05",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kevinmcisaac\" target=\"_blank\">@kevinmcisaac</a> , </p>\n<p>thank you for pointing it out. So far I haven't tried it to use even more tfrecords to improve my performance, I just split them such that they are somehow \"easy to use\" for cross validation. Do you have a resource where I can read more about it? As my TPU quota is already consumed and I don't want to use colab for this competition at the moment, do you have any further tips for improving the performance? I already used mixed precision. </p>\n<p>My motivation for stratification on the target was simply that this data is extremely noisy. When looking at my model predictions I can really say that there are sometimes predictions like 0.89 with one model and 0.15 with another by only changing a bit of \"randomness\". They had the same architecture, training procedure etc.. This is really bad, the variance is really high. So my hope was that it can be helpful to find some balance in every step I can find. And this includes the stratification on the targets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1498735,
          "author_name": "kevinmcisaac",
          "author_url": "",
          "post_date": "09/01/2021 06:59:25",
          "content": "<p>You can find my public data sets for these TFRs</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass</a></li>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii</a></li>\n<li><a href=\"https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions\" target=\"_blank\">https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions</a></li>\n</ul>\n<p>My set up is Keras/TF and I use</p>\n<ul>\n<li>Continusious Wavelets Transforms (CTW) from <a href=\"https://github.com/Kevin-McIsaac/cmorlet-tensorflow\" target=\"_blank\">here</a> </li>\n<li>Efficientnet. Generally B0 for speed while testing things.<br>\nWith a TPU and a CTW that is 128x128 an epoch takes about 3min.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1559950,
      "author_name": "zerafachris",
      "author_url": "",
      "post_date": "10/27/2021 08:36:44",
      "content": "<p>Hey All,</p>\n<p>Thank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey <a href=\"https://forms.gle/QP9L16niPexozyhu5\" target=\"_blank\">https://forms.gle/QP9L16niPexozyhu5</a>.</p>\n<p>Thank you all,</p>\n<p>Regards,<br>\nChris</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1498125": "Hey Kagglers! ;-) Are you looking for a bunch of tfrecords that were stratified on the target values? Then these datasets might be helpful:\n\n- https://www.kaggle.com/allunia/1-g2net-data-balanced\n- https://www.kaggle.com/allunia/2-g2net-data-balanced\n- https://www.kaggle.com/allunia/3-g2net-data-balanced\n- https://www.kaggle.com/allunia/4-g2net-data-balanced\n\nAs this data is quite noisy I hope that this further balance is helpful to find a stable CV. I have split them as they took too much space to create all tfrecords at once. I'm not able to download the data at the moment and haven't found a way to combine all of them in one dataset only using the kaggle platform. If you know a way please explain it to me. ;-)",
    "1498408": "allunia I've read there are performance benefit from having multiple TFrecord files. This is mentioned in a few articles  but I can't find the original one I read which was very clear. \n\nFrom memory the benefit are\n* Very large files can be slower to read.\n* With multiple TFrecord files you can parallelize the reads \n\nFrom memory rule of thumb was to have the TFrecords under a under 500MB and have at least 10. \n\nI've used this approach in my work and get fantastic performance with the TPU. \n\nHowever I'm a newbie so take this as unverified.",
    "1498722": "Hi @kevinmcisaac , \n\nthank you for pointing it out. So far I haven't tried it to use even more tfrecords to improve my performance, I just split them such that they are somehow \"easy to use\" for cross validation. Do you have a resource where I can read more about it? As my TPU quota is already consumed and I don't want to use colab for this competition at the moment, do you have any further tips for improving the performance? I already used mixed precision. \n\nMy motivation for stratification on the target was simply that this data is extremely noisy. When looking at my model predictions I can really say that there are sometimes predictions like 0.89 with one model and 0.15 with another by only changing a bit of \"randomness\". They had the same architecture, training procedure etc.. This is really bad, the variance is really high. So my hope was that it can be helpful to find some balance in every step I can find. And this includes the stratification on the targets.",
    "1498735": "You can find my public data sets for these TFRs\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-part-ii\n- https://www.kaggle.com/kevinmcisaac/g2net-tfrecords-with-20500hz-bandpass-submissions\n\nMy set up is Keras/TF and I use\n* Continusious Wavelets Transforms (CTW) from [here](https://github.com/Kevin-McIsaac/cmorlet-tensorflow) \n* Efficientnet. Generally B0 for speed while testing things.\nWith a TPU and a CTW that is 128x128 an epoch takes about 3min.",
    "1559950": "Hey All,\n\nThank you all for taking part in our competition. The participation has been overwhelmingly positive. We are currently conducting a survey to gauge the demographic and outreach achieved. Kindly spare 2min and fill in this survey https://forms.gle/QP9L16niPexozyhu5.\n\nThank you all,\n\nRegards,\nChris"
  },
  "source": "meta"
}