{
  "id": 75259,
  "title": "Augmented training set",
  "url": "/competitions/PLAsTiCC-2018/discussion/75259",
  "author_name": "",
  "post_date": "2018-12-20T01:06:20.500596300Z",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The main \"trick\" in my model was to degrade the training set to make it look much more like the test set. For every object in the training set, I made up to 40 different versions of it. See my post here for details: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033</a></p>\n\n<p>If you are interested in trying to train your model on this augmented training set, I have attached it to this post. It is attached as an HDF file with both the flux data and meta data in the same file. To read it with pandas, unzip it and use:</p>\n\n<pre><code>flux_data = pd.read_hdf('./kyle_final_augment.h5', 'df')\nmeta_data = pd.read_hdf('./kyle_final_augment.h5', 'meta')\n</code></pre>\n\n<p>There are 271148 objects in the dataset. Make sure to keep objects that came from the same original object in the same fold for your classifier... there is a \"fold\" column added to the meta data that gives 5 preset folds to use. I would recommend not using ra, decl, gal_l, gal_b or mwebv as features because I didn't change them for the augmented data and your classifier will overfit on them.</p>\n\n<p>The object_ids of the generated data are the same as those of the original data with a random number between 0 and 1 added. If you take int(object_id), you'll get the original object_id that it came from. The original training data is still in the dataset with integer object_ids.</p>\n\n<p>Let me know if your models improve by training on this dataset!</p>",
  "messages": [
    {
      "id": "442443",
      "postDate": "12/20/2018 01:06:20",
      "content": "<p>The main \"trick\" in my model was to degrade the training set to make it look much more like the test set. For every object in the training set, I made up to 40 different versions of it. See my post here for details: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033</a></p>\n\n<p>If you are interested in trying to train your model on this augmented training set, I have attached it to this post. It is attached as an HDF file with both the flux data and meta data in the same file. To read it with pandas, unzip it and use:</p>\n\n<pre><code>flux_data = pd.read_hdf('./kyle_final_augment.h5', 'df')\nmeta_data = pd.read_hdf('./kyle_final_augment.h5', 'meta')\n</code></pre>\n\n<p>There are 271148 objects in the dataset. Make sure to keep objects that came from the same original object in the same fold for your classifier... there is a \"fold\" column added to the meta data that gives 5 preset folds to use. I would recommend not using ra, decl, gal_l, gal_b or mwebv as features because I didn't change them for the augmented data and your classifier will overfit on them.</p>\n\n<p>The object_ids of the generated data are the same as those of the original data with a random number between 0 and 1 added. If you take int(object_id), you'll get the original object_id that it came from. The original training data is still in the dataset with integer object_ids.</p>\n\n<p>Let me know if your models improve by training on this dataset!</p>",
      "rawMarkdown": "The main \"trick\" in my model was to degrade the training set to make it look much more like the test set. For every object in the training set, I made up to 40 different versions of it. See my post here for details: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033\n\nIf you are interested in trying to train your model on this augmented training set, I have attached it to this post. It is attached as an HDF file with both the flux data and meta data in the same file. To read it with pandas, unzip it and use:\n\n    flux_data = pd.read_hdf('./kyle_final_augment.h5', 'df')\n    meta_data = pd.read_hdf('./kyle_final_augment.h5', 'meta')\n\nThere are 271148 objects in the dataset. Make sure to keep objects that came from the same original object in the same fold for your classifier... there is a \"fold\" column added to the meta data that gives 5 preset folds to use. I would recommend not using ra, decl, gal\\_l, gal\\_b or mwebv as features because I didn't change them for the augmented data and your classifier will overfit on them.\n\nThe object_ids of the generated data are the same as those of the original data with a random number between 0 and 1 added. If you take int(object\\_id), you'll get the original object\\_id that it came from. The original training data is still in the dataset with integer object\\_ids.\n\nLet me know if your models improve by training on this dataset!",
      "votes": null
    },
    {
      "id": "442474",
      "postDate": "12/20/2018 02:19:40",
      "content": "<p>Wow, that was fast.  This and other sharing from other teams will make me work on this this week end I'm afraid... ;)  </p>\n\n<p>It is the first time I see so much great sharing from top teams and many teams in general.</p>",
      "rawMarkdown": "Wow, that was fast.  This and other sharing from other teams will make me work on this this week end I'm afraid... ;)  \n\nIt is the first time I see so much great sharing from top teams and many teams in general.",
      "votes": null
    },
    {
      "id": "442555",
      "postDate": "12/20/2018 05:44:05",
      "content": "<p>Kyle, Thanks for providing this data. \nIf possible, please share it as a Kaggle dataset. It will be easy to try it out in Kaggle Kernels. <a href=\"https://www.kaggle.com/datasets\">https://www.kaggle.com/datasets</a></p>",
      "rawMarkdown": "Kyle, Thanks for providing this data. \nIf possible, please share it as a Kaggle dataset. It will be easy to try it out in Kaggle Kernels. https://www.kaggle.com/datasets",
      "votes": null
    },
    {
      "id": "442880",
      "postDate": "12/20/2018 16:26:28",
      "content": "<p>Thanks for sharing, Kyle!</p>",
      "rawMarkdown": "Thanks for sharing, Kyle!",
      "votes": null
    },
    {
      "id": "443027",
      "postDate": "12/20/2018 22:51:47",
      "content": "<p>congratulations, thanks for sharing the motivation of your solution </p>",
      "rawMarkdown": "congratulations, thanks for sharing the motivation of your solution",
      "votes": null
    },
    {
      "id": "444487",
      "postDate": "12/24/2018 05:29:06",
      "content": "<p>Congratulations and thank you for sharing !</p>\n\n<p>I tried training my model (single lgb, no augmentation) on your augmented dataset (spend 3 days on my machine ...)\nand get improved score :\nPrivate LB +0.120, Public LB +0.115\nIt's really nice !</p>",
      "rawMarkdown": "Congratulations and thank you for sharing !\n\nI tried training my model (single lgb, no augmentation) on your augmented dataset (spend 3 days on my machine ...)\nand get improved score :\nPrivate LB +0.120, Public LB +0.115\nIt's really nice !",
      "votes": null
    },
    {
      "id": "445199",
      "postDate": "12/25/2018 23:17:06",
      "content": "<p>Thanks for sharing. I trained my best model on your augmented set and got no improvement (or barely). But this seems to fix my overfitting problem. I actually might be underfitting now.  So let's see what happens when I loosen up some of those regularization parameters.</p>",
      "rawMarkdown": "Thanks for sharing. I trained my best model on your augmented set and got no improvement (or barely). But this seems to fix my overfitting problem. I actually might be underfitting now.  So let's see what happens when I loosen up some of those regularization parameters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442474,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/20/2018 02:19:40",
      "content": "<p>Wow, that was fast.  This and other sharing from other teams will make me work on this this week end I'm afraid... ;)  </p>\n\n<p>It is the first time I see so much great sharing from top teams and many teams in general.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442555,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "12/20/2018 05:44:05",
      "content": "<p>Kyle, Thanks for providing this data. \nIf possible, please share it as a Kaggle dataset. It will be easy to try it out in Kaggle Kernels. <a href=\"https://www.kaggle.com/datasets\">https://www.kaggle.com/datasets</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 442880,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "12/20/2018 16:26:28",
      "content": "<p>Thanks for sharing, Kyle!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443027,
      "author_name": "ahmedbaz",
      "author_url": "",
      "post_date": "12/20/2018 22:51:47",
      "content": "<p>congratulations, thanks for sharing the motivation of your solution </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444487,
      "author_name": "tikutiku",
      "author_url": "",
      "post_date": "12/24/2018 05:29:06",
      "content": "<p>Congratulations and thank you for sharing !</p>\n\n<p>I tried training my model (single lgb, no augmentation) on your augmented dataset (spend 3 days on my machine ...)\nand get improved score :\nPrivate LB +0.120, Public LB +0.115\nIt's really nice !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445199,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "12/25/2018 23:17:06",
      "content": "<p>Thanks for sharing. I trained my best model on your augmented set and got no improvement (or barely). But this seems to fix my overfitting problem. I actually might be underfitting now.  So let's see what happens when I loosen up some of those regularization parameters.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "442443": "The main \"trick\" in my model was to degrade the training set to make it look much more like the test set. For every object in the training set, I made up to 40 different versions of it. See my post here for details: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75033\n\nIf you are interested in trying to train your model on this augmented training set, I have attached it to this post. It is attached as an HDF file with both the flux data and meta data in the same file. To read it with pandas, unzip it and use:\n\n    flux_data = pd.read_hdf('./kyle_final_augment.h5', 'df')\n    meta_data = pd.read_hdf('./kyle_final_augment.h5', 'meta')\n\nThere are 271148 objects in the dataset. Make sure to keep objects that came from the same original object in the same fold for your classifier... there is a \"fold\" column added to the meta data that gives 5 preset folds to use. I would recommend not using ra, decl, gal\\_l, gal\\_b or mwebv as features because I didn't change them for the augmented data and your classifier will overfit on them.\n\nThe object_ids of the generated data are the same as those of the original data with a random number between 0 and 1 added. If you take int(object\\_id), you'll get the original object\\_id that it came from. The original training data is still in the dataset with integer object\\_ids.\n\nLet me know if your models improve by training on this dataset!",
    "442474": "Wow, that was fast.  This and other sharing from other teams will make me work on this this week end I'm afraid... ;)  \n\nIt is the first time I see so much great sharing from top teams and many teams in general.",
    "442555": "Kyle, Thanks for providing this data. \nIf possible, please share it as a Kaggle dataset. It will be easy to try it out in Kaggle Kernels. https://www.kaggle.com/datasets",
    "442880": "Thanks for sharing, Kyle!",
    "443027": "congratulations, thanks for sharing the motivation of your solution",
    "444487": "Congratulations and thank you for sharing !\n\nI tried training my model (single lgb, no augmentation) on your augmented dataset (spend 3 days on my machine ...)\nand get improved score :\nPrivate LB +0.120, Public LB +0.115\nIt's really nice !",
    "445199": "Thanks for sharing. I trained my best model on your augmented set and got no improvement (or barely). But this seems to fix my overfitting problem. I actually might be underfitting now.  So let's see what happens when I loosen up some of those regularization parameters."
  },
  "source": "meta"
}