{
  "id": 194150,
  "title": "TFRecord Versions of Eruption Data - With Helper Notebooks",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/194150",
  "author_name": "",
  "post_date": "2020-10-30T21:31:50.275205800Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all.</p>\n<p>I recently turned the data files for this competition into TFRecords and uploaded them as datasets.<br>\nAll TFRecord files contain 80 examples (except for the Validation TFRecord which contains 31 and the last test TFRecord which will contain whatever was leftover when split into chunks of 80).</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-training-tfrecords\" target=\"_blank\">Train Dataset</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-testing-tfrecords\" target=\"_blank\">Test Dataset</a></p>\n<p>Here are the corresponding links to notebooks showing how to access the TFRecord files and turn them into tf.data.Dataset objects.</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/how-to-use-train-tfrecords\" target=\"_blank\">Notebook for Reading Training Data</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/how-to-use-test-tfrecords\" target=\"_blank\">Notebook for Reading Testing Data</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords\" target=\"_blank\">Basic Conv1D SqueezeNet model Using TFRecords</a></p>\n<p><strong>The only preprocessing that was done prior to conversion was filling all NaN values as 0.</strong></p>\n<p>The notebooks also include an appendix cell showing how the TFRecord files were created.</p>\n<p>Let me know if you have any questions. This is my first time uploading a dataset for other people's usage and I hope I did everything right. If something is off or wrong, please let me know and I can change it ASAP.</p>\n<p>Warmest Regards,<br>\nDarien Schettler</p>",
  "messages": [
    {
      "id": "1065113",
      "postDate": "10/30/2020 21:31:50",
      "content": "<p>Hi all.</p>\n<p>I recently turned the data files for this competition into TFRecords and uploaded them as datasets.<br>\nAll TFRecord files contain 80 examples (except for the Validation TFRecord which contains 31 and the last test TFRecord which will contain whatever was leftover when split into chunks of 80).</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-training-tfrecords\" target=\"_blank\">Train Dataset</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-testing-tfrecords\" target=\"_blank\">Test Dataset</a></p>\n<p>Here are the corresponding links to notebooks showing how to access the TFRecord files and turn them into tf.data.Dataset objects.</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/how-to-use-train-tfrecords\" target=\"_blank\">Notebook for Reading Training Data</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/how-to-use-test-tfrecords\" target=\"_blank\">Notebook for Reading Testing Data</a><br>\n<a href=\"https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords\" target=\"_blank\">Basic Conv1D SqueezeNet model Using TFRecords</a></p>\n<p><strong>The only preprocessing that was done prior to conversion was filling all NaN values as 0.</strong></p>\n<p>The notebooks also include an appendix cell showing how the TFRecord files were created.</p>\n<p>Let me know if you have any questions. This is my first time uploading a dataset for other people's usage and I hope I did everything right. If something is off or wrong, please let me know and I can change it ASAP.</p>\n<p>Warmest Regards,<br>\nDarien Schettler</p>",
      "rawMarkdown": "Hi all.\n\nI recently turned the data files for this competition into TFRecords and uploaded them as datasets.\nAll TFRecord files contain 80 examples (except for the Validation TFRecord which contains 31 and the last test TFRecord which will contain whatever was leftover when split into chunks of 80).\n\n[Train Dataset](https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-training-tfrecords)\n[Test Dataset](https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-testing-tfrecords)\n\nHere are the corresponding links to notebooks showing how to access the TFRecord files and turn them into tf.data.Dataset objects.\n\n[Notebook for Reading Training Data](https://www.kaggle.com/dschettler8845/how-to-use-train-tfrecords)\n[Notebook for Reading Testing Data](https://www.kaggle.com/dschettler8845/how-to-use-test-tfrecords)\n[Basic Conv1D SqueezeNet model Using TFRecords](https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords)\n\n**The only preprocessing that was done prior to conversion was filling all NaN values as 0.**\n\nThe notebooks also include an appendix cell showing how the TFRecord files were created.\n\nLet me know if you have any questions. This is my first time uploading a dataset for other people's usage and I hope I did everything right. If something is off or wrong, please let me know and I can change it ASAP.\n\nWarmest Regards,\nDarien Schettler",
      "votes": null
    },
    {
      "id": "1065130",
      "postDate": "10/30/2020 22:04:09",
      "content": "<p>Update: There was an error in the Training TFRecords that is being corrected presently. It will be uploaded/fixed by 8PM EST October 30, 2020. </p>",
      "rawMarkdown": "Update: There was an error in the Training TFRecords that is being corrected presently. It will be uploaded/fixed by 8PM EST October 30, 2020.",
      "votes": null
    },
    {
      "id": "1065163",
      "postDate": "10/30/2020 23:41:51",
      "content": "<p>This is definitely useful.</p>\n<p>I'd tried creating a couple of TFRecords datasets with some spectrograms etc but unfortunately it didn't seem like my NN models were able to learn very much from them. This could just be a failure on my part.</p>\n<p>I certainly think that given the number of input feed files, a TFRecords-type approach will be really useful because it does currently take too long to generate features etc.</p>\n<p>If I manage to get a TF Records data feed/keras model which does appear to provide some useful outcome, will make it public.</p>",
      "rawMarkdown": "This is definitely useful.\n\nI'd tried creating a couple of TFRecords datasets with some spectrograms etc but unfortunately it didn't seem like my NN models were able to learn very much from them. This could just be a failure on my part.\n\nI certainly think that given the number of input feed files, a TFRecords-type approach will be really useful because it does currently take too long to generate features etc.\n\nIf I manage to get a TF Records data feed/keras model which does appear to provide some useful outcome, will make it public.",
      "votes": null
    },
    {
      "id": "1065646",
      "postDate": "10/31/2020 14:44:55",
      "content": "<p>New versions of the dataset are uploaded.</p>\n<p>I added the ID of the test examples in the place of the 'label' for the testing tfrecords so that you can align them with the required submission format.</p>\n<p>I also updated the notebooks.</p>\n<p>I haven't tested anything and will be attempting to make a simple SqueezeNet model later today.</p>\n<p><a href=\"https://www.kaggle.com/davidedwards1\" target=\"_blank\">@davidedwards1</a> – That sounds awesome. Keep me up-to-date!</p>",
      "rawMarkdown": "New versions of the dataset are uploaded.\n\nI added the ID of the test examples in the place of the 'label' for the testing tfrecords so that you can align them with the required submission format.\n\nI also updated the notebooks.\n\nI haven't tested anything and will be attempting to make a simple SqueezeNet model later today.\n\n@davidedwards1 – That sounds awesome. Keep me up-to-date!",
      "votes": null
    },
    {
      "id": "1065735",
      "postDate": "10/31/2020 17:04:38",
      "content": "<p>Update: There was an error in the Testing data with the first TFRecord being repeated over and over.</p>\n<p>I'll have it fixed and updated by this evening.</p>",
      "rawMarkdown": "Update: There was an error in the Testing data with the first TFRecord being repeated over and over.\n\nI'll have it fixed and updated by this evening.",
      "votes": null
    },
    {
      "id": "1068048",
      "postDate": "11/03/2020 02:53:41",
      "content": "<p>Thanks again. Have you had any luck with using TF/keras or other neural net approach on this? So far I've not been able to show any results comparable to result from using XGB, I think.</p>\n<p>Not sure if I'm doing something wrong. I feel like maybe part of the problem is that it's not that easy for the NN to pick out specific 'events', and some data samples which are close to the eruption don't have a strong / distinctive signal within the recording timeframe.</p>",
      "rawMarkdown": "Thanks again. Have you had any luck with using TF/keras or other neural net approach on this? So far I've not been able to show any results comparable to result from using XGB, I think.\n\nNot sure if I'm doing something wrong. I feel like maybe part of the problem is that it's not that easy for the NN to pick out specific 'events', and some data samples which are close to the eruption don't have a strong / distinctive signal within the recording timeframe.",
      "votes": null
    },
    {
      "id": "1069634",
      "postDate": "11/04/2020 17:41:08",
      "content": "<p>Thanks for sharing! Any luck using these with a TPU?</p>",
      "rawMarkdown": "Thanks for sharing! Any luck using these with a TPU?",
      "votes": null
    },
    {
      "id": "1070514",
      "postDate": "11/05/2020 21:03:00",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> – I have not tried using it with a TPU. But I don't see why it would have any problem.</p>\n<p><a href=\"https://www.kaggle.com/Dave\" target=\"_blank\">@Dave</a> E – I used a Conv 1D based Squeeze Net achieving around 7,500,000 MAE. Not ideal by any means. But this was a naive baseline with little preprocessing and little to no tuning. I believe if we did some feature engineering on the time-series to better capture 'event-like phenomena', and potentially added this to the model (i.e. attention), it would do very well. </p>\n<p>Normally I would spend much longer on the EDA portion of things, and I think this competition should be no exception. Without spending a lot of time doing analysis on the data, I know I made some gross errors/assumptions in modelling. </p>\n<p>Hopefully, someone can take the notebook I created and improve upon it. If I get some time in the following weeks I will return to this and attempt a better solution.</p>",
      "rawMarkdown": "sohier – I have not tried using it with a TPU. But I don't see why it would have any problem.\n\n@Dave E – I used a Conv 1D based Squeeze Net achieving around 7,500,000 MAE. Not ideal by any means. But this was a naive baseline with little preprocessing and little to no tuning. I believe if we did some feature engineering on the time-series to better capture 'event-like phenomena', and potentially added this to the model (i.e. attention), it would do very well. \n\nNormally I would spend much longer on the EDA portion of things, and I think this competition should be no exception. Without spending a lot of time doing analysis on the data, I know I made some gross errors/assumptions in modelling. \n\nHopefully, someone can take the notebook I created and improve upon it. If I get some time in the following weeks I will return to this and attempt a better solution.",
      "votes": null
    },
    {
      "id": "1070738",
      "postDate": "11/06/2020 06:02:17",
      "content": "<p>That's good to know! I'd tried LSTM and Conv1D without getting all that much better than a 'guessing' level error. And I did use TPU (just on my own data set), just didn't share results as they were rubbish…</p>\n<p>Aside from my own lack of knowledge I think one problem is that there are not always very distinctive 'events'. At least with the XGB feature importances, it looks like gradients and percentiles (e.g. 90th, 75th, 25th etc) are important and it feels like extracting these kind of features from the data has done better so far.</p>\n<p>I will try again if I get time as MAE = 7,500,000 feels like a starting point…</p>",
      "rawMarkdown": "That's good to know! I'd tried LSTM and Conv1D without getting all that much better than a 'guessing' level error. And I did use TPU (just on my own data set), just didn't share results as they were rubbish...\n\nAside from my own lack of knowledge I think one problem is that there are not always very distinctive 'events'. At least with the XGB feature importances, it looks like gradients and percentiles (e.g. 90th, 75th, 25th etc) are important and it feels like extracting these kind of features from the data has done better so far.\n\nI will try again if I get time as MAE = 7,500,000 feels like a starting point...",
      "votes": null
    },
    {
      "id": "1071371",
      "postDate": "11/06/2020 19:31:34",
      "content": "<p>Here's the version of the notebook I was working on. Hopefully, it helps!</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords\" target=\"_blank\">https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords</a></p>",
      "rawMarkdown": "Here's the version of the notebook I was working on. Hopefully, it helps!\n\nhttps://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1065130,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "10/30/2020 22:04:09",
      "content": "<p>Update: There was an error in the Training TFRecords that is being corrected presently. It will be uploaded/fixed by 8PM EST October 30, 2020. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065163,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "10/30/2020 23:41:51",
      "content": "<p>This is definitely useful.</p>\n<p>I'd tried creating a couple of TFRecords datasets with some spectrograms etc but unfortunately it didn't seem like my NN models were able to learn very much from them. This could just be a failure on my part.</p>\n<p>I certainly think that given the number of input feed files, a TFRecords-type approach will be really useful because it does currently take too long to generate features etc.</p>\n<p>If I manage to get a TF Records data feed/keras model which does appear to provide some useful outcome, will make it public.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065646,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "10/31/2020 14:44:55",
      "content": "<p>New versions of the dataset are uploaded.</p>\n<p>I added the ID of the test examples in the place of the 'label' for the testing tfrecords so that you can align them with the required submission format.</p>\n<p>I also updated the notebooks.</p>\n<p>I haven't tested anything and will be attempting to make a simple SqueezeNet model later today.</p>\n<p><a href=\"https://www.kaggle.com/davidedwards1\" target=\"_blank\">@davidedwards1</a> – That sounds awesome. Keep me up-to-date!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1065735,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "10/31/2020 17:04:38",
      "content": "<p>Update: There was an error in the Testing data with the first TFRecord being repeated over and over.</p>\n<p>I'll have it fixed and updated by this evening.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1068048,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "11/03/2020 02:53:41",
      "content": "<p>Thanks again. Have you had any luck with using TF/keras or other neural net approach on this? So far I've not been able to show any results comparable to result from using XGB, I think.</p>\n<p>Not sure if I'm doing something wrong. I feel like maybe part of the problem is that it's not that easy for the NN to pick out specific 'events', and some data samples which are close to the eruption don't have a strong / distinctive signal within the recording timeframe.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1069634,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "11/04/2020 17:41:08",
      "content": "<p>Thanks for sharing! Any luck using these with a TPU?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1070514,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "11/05/2020 21:03:00",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> – I have not tried using it with a TPU. But I don't see why it would have any problem.</p>\n<p><a href=\"https://www.kaggle.com/Dave\" target=\"_blank\">@Dave</a> E – I used a Conv 1D based Squeeze Net achieving around 7,500,000 MAE. Not ideal by any means. But this was a naive baseline with little preprocessing and little to no tuning. I believe if we did some feature engineering on the time-series to better capture 'event-like phenomena', and potentially added this to the model (i.e. attention), it would do very well. </p>\n<p>Normally I would spend much longer on the EDA portion of things, and I think this competition should be no exception. Without spending a lot of time doing analysis on the data, I know I made some gross errors/assumptions in modelling. </p>\n<p>Hopefully, someone can take the notebook I created and improve upon it. If I get some time in the following weeks I will return to this and attempt a better solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1070738,
          "author_name": "davidedwards1",
          "author_url": "",
          "post_date": "11/06/2020 06:02:17",
          "content": "<p>That's good to know! I'd tried LSTM and Conv1D without getting all that much better than a 'guessing' level error. And I did use TPU (just on my own data set), just didn't share results as they were rubbish…</p>\n<p>Aside from my own lack of knowledge I think one problem is that there are not always very distinctive 'events'. At least with the XGB feature importances, it looks like gradients and percentiles (e.g. 90th, 75th, 25th etc) are important and it feels like extracting these kind of features from the data has done better so far.</p>\n<p>I will try again if I get time as MAE = 7,500,000 feels like a starting point…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1071371,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "11/06/2020 19:31:34",
          "content": "<p>Here's the version of the notebook I was working on. Hopefully, it helps!</p>\n<p><a href=\"https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords\" target=\"_blank\">https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1065113": "Hi all.\n\nI recently turned the data files for this competition into TFRecords and uploaded them as datasets.\nAll TFRecord files contain 80 examples (except for the Validation TFRecord which contains 31 and the last test TFRecord which will contain whatever was leftover when split into chunks of 80).\n\n[Train Dataset](https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-training-tfrecords)\n[Test Dataset](https://www.kaggle.com/dschettler8845/ingv-volcanic-eruption-testing-tfrecords)\n\nHere are the corresponding links to notebooks showing how to access the TFRecord files and turn them into tf.data.Dataset objects.\n\n[Notebook for Reading Training Data](https://www.kaggle.com/dschettler8845/how-to-use-train-tfrecords)\n[Notebook for Reading Testing Data](https://www.kaggle.com/dschettler8845/how-to-use-test-tfrecords)\n[Basic Conv1D SqueezeNet model Using TFRecords](https://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords)\n\n**The only preprocessing that was done prior to conversion was filling all NaN values as 0.**\n\nThe notebooks also include an appendix cell showing how the TFRecord files were created.\n\nLet me know if you have any questions. This is my first time uploading a dataset for other people's usage and I hope I did everything right. If something is off or wrong, please let me know and I can change it ASAP.\n\nWarmest Regards,\nDarien Schettler",
    "1065130": "Update: There was an error in the Training TFRecords that is being corrected presently. It will be uploaded/fixed by 8PM EST October 30, 2020.",
    "1065163": "This is definitely useful.\n\nI'd tried creating a couple of TFRecords datasets with some spectrograms etc but unfortunately it didn't seem like my NN models were able to learn very much from them. This could just be a failure on my part.\n\nI certainly think that given the number of input feed files, a TFRecords-type approach will be really useful because it does currently take too long to generate features etc.\n\nIf I manage to get a TF Records data feed/keras model which does appear to provide some useful outcome, will make it public.",
    "1065646": "New versions of the dataset are uploaded.\n\nI added the ID of the test examples in the place of the 'label' for the testing tfrecords so that you can align them with the required submission format.\n\nI also updated the notebooks.\n\nI haven't tested anything and will be attempting to make a simple SqueezeNet model later today.\n\n@davidedwards1 – That sounds awesome. Keep me up-to-date!",
    "1065735": "Update: There was an error in the Testing data with the first TFRecord being repeated over and over.\n\nI'll have it fixed and updated by this evening.",
    "1068048": "Thanks again. Have you had any luck with using TF/keras or other neural net approach on this? So far I've not been able to show any results comparable to result from using XGB, I think.\n\nNot sure if I'm doing something wrong. I feel like maybe part of the problem is that it's not that easy for the NN to pick out specific 'events', and some data samples which are close to the eruption don't have a strong / distinctive signal within the recording timeframe.",
    "1069634": "Thanks for sharing! Any luck using these with a TPU?",
    "1070514": "sohier – I have not tried using it with a TPU. But I don't see why it would have any problem.\n\n@Dave E – I used a Conv 1D based Squeeze Net achieving around 7,500,000 MAE. Not ideal by any means. But this was a naive baseline with little preprocessing and little to no tuning. I believe if we did some feature engineering on the time-series to better capture 'event-like phenomena', and potentially added this to the model (i.e. attention), it would do very well. \n\nNormally I would spend much longer on the EDA portion of things, and I think this competition should be no exception. Without spending a lot of time doing analysis on the data, I know I made some gross errors/assumptions in modelling. \n\nHopefully, someone can take the notebook I created and improve upon it. If I get some time in the following weeks I will return to this and attempt a better solution.",
    "1070738": "That's good to know! I'd tried LSTM and Conv1D without getting all that much better than a 'guessing' level error. And I did use TPU (just on my own data set), just didn't share results as they were rubbish...\n\nAside from my own lack of knowledge I think one problem is that there are not always very distinctive 'events'. At least with the XGB feature importances, it looks like gradients and percentiles (e.g. 90th, 75th, 25th etc) are important and it feels like extracting these kind of features from the data has done better so far.\n\nI will try again if I get time as MAE = 7,500,000 feels like a starting point...",
    "1071371": "Here's the version of the notebook I was working on. Hopefully, it helps!\n\nhttps://www.kaggle.com/dschettler8845/basic-squeezenet-architecture-w-tfrecords"
  },
  "source": "meta"
}