{
  "id": 85186,
  "title": "38th Solution & How I think I survived the shake-up",
  "url": "/competitions/vsb-power-line-fault-detection/writeups/theo-viel-38th-solution-how-i-think-i-survived-the",
  "author_name": "",
  "post_date": "2019-03-22T18:33:49.663Z",
  "votes": 16,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hello everybody!</p>\n\n<p><strong>The Start :</strong>\nI joined this comptetion approx. a month ago, at the start, my main idea was to use scipy.signal to extract features about the peaks, after applying wavelet denoising. I first went with LSTMs because of the hype around high scoring public kernels, but could not get close to 0.7 LB (public).</p>\n\n<p><strong>Downfalls about public Kernels :</strong>\nHere are a few things I believe were overused and caused the shaked_up.\n- Keras CuDNNLSTMs (because of the randomness), move to PyTorch !\n- Threshold selection for MCC optimization. It only caused overfitting on your oof data\n- Checkpoints while training your model. Again, you're overfitting on your oof data\n- Overall, I went with the idea that RNNs were predicting noise only. </p>\n\n<p>Not sure about the last point, because I've got one of my very first submissions that got 0.674 private with the ideas above.</p>\n\n<p><strong>To the magical submission : CV 0.659 - Public 0.652 - Private 0.667</strong></p>\n\n<p>Yes that's pretty stable, that's why I trusted this model more than others.</p>\n\n<p>Feature Engineering : \n- Wavelet Denoizing\n- \"Bucketted\" distribution features such as in the public kernels (exactly the same actually).\n- The three phases were fed together in the network, BUT as three phases of an image. My input shape was (3, 160, nb_features). I used the most common target among the three phases as the label.</p>\n\n<p>I tried all the features I came up with, these ones achieved the best CV, I don't know why.</p>\n\n<p><strong>Model :</strong> \nAs my inputs were similar to images now, I went with an architecture similar to AlexNet, but with less parameters. The architecture looks like this approx. (I changed it a bit, not sure if this is the one that got me the best score) :</p>\n\n<p>```\nclass Model(nn.Module):\n    def <strong>init</strong>(self):\n        super(Model, self).<strong>init</strong>()</p>\n\n<pre><code>    self.features = nn.Sequential(\n        nn.Conv2d(3, 16, kernel_size=5, padding=2),\n        nn.ReLU(inplace=True),\n        nn.MaxPool2d(kernel_size=3, stride=2),\n        nn.Conv2d(16, 32, kernel_size=3, padding=2),\n        nn.ReLU(inplace=True),\n        nn.MaxPool2d(kernel_size=3, stride=2))\n\n    self.avgpool = nn.AdaptiveAvgPool2d((6, 6))\n\n    self.classifier = nn.Sequential(\n        nn.Dropout(),\n        nn.Linear(32 * 6 * 6, 32),\n        nn.ReLU(inplace=True),\n        nn.Dropout(),\n        nn.Linear(32, 1))\n\ndef forward(self, x):\n    x = self.features(x)\n    x = self.avgpool(x)\n    x = x.view(x.size(0), 32 * 6 * 6)\n    x = self.classifier(x)\n    return x\n</code></pre>\n\n<p>```\n100 epochs, 5 folds CV, batch_size 128, Learning rate 0.001</p>\n\n<p>I worked on my gaming laptop which has 16Gb of ram and a 1060, so the 800000 long signals were quite a pain in the a** but I dealt with it.</p>\n\n<p><strong>Final Word</strong>\nThis competition was hard because you could not trust your LB, and you could only trust your CV if it achieved results similar to LB ones. Well that's what I understood at least.</p>\n\n<p>Thanks everyone for participating, I'm waiting to read about others' solutions.!</p>",
  "messages": [
    {
      "id": "496411",
      "postDate": "03/22/2019 06:53:41",
      "content": "<p>Hello everybody!</p>\n\n<p><strong>The Start :</strong>\nI joined this comptetion approx. a month ago, at the start, my main idea was to use scipy.signal to extract features about the peaks, after applying wavelet denoising. I first went with LSTMs because of the hype around high scoring public kernels, but could not get close to 0.7 LB (public).</p>\n\n<p><strong>Downfalls about public Kernels :</strong>\nHere are a few things I believe were overused and caused the shaked_up.\n- Keras CuDNNLSTMs (because of the randomness), move to PyTorch !\n- Threshold selection for MCC optimization. It only caused overfitting on your oof data\n- Checkpoints while training your model. Again, you're overfitting on your oof data\n- Overall, I went with the idea that RNNs were predicting noise only. </p>\n\n<p>Not sure about the last point, because I've got one of my very first submissions that got 0.674 private with the ideas above.</p>\n\n<p><strong>To the magical submission : CV 0.659 - Public 0.652 - Private 0.667</strong></p>\n\n<p>Yes that's pretty stable, that's why I trusted this model more than others.</p>\n\n<p>Feature Engineering : \n- Wavelet Denoizing\n- \"Bucketted\" distribution features such as in the public kernels (exactly the same actually).\n- The three phases were fed together in the network, BUT as three phases of an image. My input shape was (3, 160, nb_features). I used the most common target among the three phases as the label.</p>\n\n<p>I tried all the features I came up with, these ones achieved the best CV, I don't know why.</p>\n\n<p><strong>Model :</strong> \nAs my inputs were similar to images now, I went with an architecture similar to AlexNet, but with less parameters. The architecture looks like this approx. (I changed it a bit, not sure if this is the one that got me the best score) :</p>\n\n<p>```\nclass Model(nn.Module):\n    def <strong>init</strong>(self):\n        super(Model, self).<strong>init</strong>()</p>\n\n<pre><code>    self.features = nn.Sequential(\n        nn.Conv2d(3, 16, kernel_size=5, padding=2),\n        nn.ReLU(inplace=True),\n        nn.MaxPool2d(kernel_size=3, stride=2),\n        nn.Conv2d(16, 32, kernel_size=3, padding=2),\n        nn.ReLU(inplace=True),\n        nn.MaxPool2d(kernel_size=3, stride=2))\n\n    self.avgpool = nn.AdaptiveAvgPool2d((6, 6))\n\n    self.classifier = nn.Sequential(\n        nn.Dropout(),\n        nn.Linear(32 * 6 * 6, 32),\n        nn.ReLU(inplace=True),\n        nn.Dropout(),\n        nn.Linear(32, 1))\n\ndef forward(self, x):\n    x = self.features(x)\n    x = self.avgpool(x)\n    x = x.view(x.size(0), 32 * 6 * 6)\n    x = self.classifier(x)\n    return x\n</code></pre>\n\n<p>```\n100 epochs, 5 folds CV, batch_size 128, Learning rate 0.001</p>\n\n<p>I worked on my gaming laptop which has 16Gb of ram and a 1060, so the 800000 long signals were quite a pain in the a** but I dealt with it.</p>\n\n<p><strong>Final Word</strong>\nThis competition was hard because you could not trust your LB, and you could only trust your CV if it achieved results similar to LB ones. Well that's what I understood at least.</p>\n\n<p>Thanks everyone for participating, I'm waiting to read about others' solutions.!</p>",
      "rawMarkdown": "Hello everybody!\n\n**The Start :**\nI joined this comptetion approx. a month ago, at the start, my main idea was to use scipy.signal to extract features about the peaks, after applying wavelet denoising. I first went with LSTMs because of the hype around high scoring public kernels, but could not get close to 0.7 LB (public).\n\n**Downfalls about public Kernels :**\nHere are a few things I believe were overused and caused the shaked_up.\n- Keras CuDNNLSTMs (because of the randomness), move to PyTorch !\n- Threshold selection for MCC optimization. It only caused overfitting on your oof data\n- Checkpoints while training your model. Again, you're overfitting on your oof data\n- Overall, I went with the idea that RNNs were predicting noise only. \n\nNot sure about the last point, because I've got one of my very first submissions that got 0.674 private with the ideas above.\n\n**To the magical submission : CV 0.659 - Public 0.652 - Private 0.667**\n\nYes that's pretty stable, that's why I trusted this model more than others.\n\nFeature Engineering : \n- Wavelet Denoizing\n- \"Bucketted\" distribution features such as in the public kernels (exactly the same actually).\n- The three phases were fed together in the network, BUT as three phases of an image. My input shape was (3, 160, nb_features). I used the most common target among the three phases as the label.\n\nI tried all the features I came up with, these ones achieved the best CV, I don't know why.\n\n**Model :** \nAs my inputs were similar to images now, I went with an architecture similar to AlexNet, but with less parameters. The architecture looks like this approx. (I changed it a bit, not sure if this is the one that got me the best score) :\n\n```\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        \n        self.features = nn.Sequential(\n            nn.Conv2d(3, 16, kernel_size=5, padding=2),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=3, stride=2),\n            nn.Conv2d(16, 32, kernel_size=3, padding=2),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=3, stride=2))\n        \n        self.avgpool = nn.AdaptiveAvgPool2d((6, 6))\n        \n        self.classifier = nn.Sequential(\n            nn.Dropout(),\n            nn.Linear(32 * 6 * 6, 32),\n            nn.ReLU(inplace=True),\n            nn.Dropout(),\n            nn.Linear(32, 1))\n\n    def forward(self, x):\n        x = self.features(x)\n        x = self.avgpool(x)\n        x = x.view(x.size(0), 32 * 6 * 6)\n        x = self.classifier(x)\n        return x\n```\n100 epochs, 5 folds CV, batch_size 128, Learning rate 0.001\n\nI worked on my gaming laptop which has 16Gb of ram and a 1060, so the 800000 long signals were quite a pain in the a** but I dealt with it.\n\n**Final Word**\nThis competition was hard because you could not trust your LB, and you could only trust your CV if it achieved results similar to LB ones. Well that's what I understood at least.\n\nThanks everyone for participating, I'm waiting to read about others' solutions.!",
      "votes": null
    },
    {
      "id": "496431",
      "postDate": "03/22/2019 07:37:00",
      "content": "<p>Congrat Theoviel!  Thanks for sharing!</p>\n\n<p>It’s great to know that the image approach did work! I also tried using MobileNet however, in addition to features that you explained, I also included the spectrogram features (thinking that it’s more look like image haha) ....  And although it was able to fit well, it was not able to generalize to even the OOF.</p>\n\n<p>Would you mind publish your kernel if possible ? Did you just use 5 Folds CV as validation scheme ?</p>",
      "rawMarkdown": "Congrat Theoviel!  Thanks for sharing!\n\nIt’s great to know that the image approach did work! I also tried using MobileNet however, in addition to features that you explained, I also included the spectrogram features (thinking that it’s more look like image haha) ....  And although it was able to fit well, it was not able to generalize to even the OOF.\n\nWould you mind publish your kernel if possible ? Did you just use 5 Folds CV as validation scheme ?",
      "votes": null
    },
    {
      "id": "496459",
      "postDate": "03/22/2019 08:23:36",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "496464",
      "postDate": "03/22/2019 08:28:35",
      "content": "<p>I don't think I'll publish my work, mostly because it'll take me too much time to fit what I did in the kaggle kernels. I also tried to feed spectrograms to my nn, but it did not work at all.</p>\n\n<p>Yup my validation scheme is just a classic 5 folds CV</p>",
      "rawMarkdown": "I don't think I'll publish my work, mostly because it'll take me too much time to fit what I did in the kaggle kernels. I also tried to feed spectrograms to my nn, but it did not work at all.\n\nYup my validation scheme is just a classic 5 folds CV",
      "votes": null
    },
    {
      "id": "496476",
      "postDate": "03/22/2019 08:50:27",
      "content": "<p>Thanks for your answer! (just asked in case that you implemented in the kernel)</p>",
      "rawMarkdown": "Thanks for your answer! (just asked in case that you implemented in the kernel)",
      "votes": null
    },
    {
      "id": "496535",
      "postDate": "03/22/2019 09:59:18",
      "content": "<p>Congrats Theo, and thanks for sharing.\n \"Keras CuDNNLSTMs (because of the randomness), move to PyTorch\" : definitively !</p>",
      "rawMarkdown": "Congrats Theo, and thanks for sharing.\n \"Keras CuDNNLSTMs (because of the randomness), move to PyTorch\" : definitively !",
      "votes": null
    },
    {
      "id": "496604",
      "postDate": "03/22/2019 11:36:03",
      "content": "<p>Nice approach Theo!    Did you do any augmentations in your training dataloader?</p>",
      "rawMarkdown": "Nice approach Theo!    Did you do any augmentations in your training dataloader?",
      "votes": null
    },
    {
      "id": "496656",
      "postDate": "03/22/2019 12:48:34",
      "content": "<p>Thanks Russ!\nNot at all, but there sure are loads of possibilities here!</p>",
      "rawMarkdown": "Thanks Russ!\nNot at all, but there sure are loads of possibilities here!",
      "votes": null
    },
    {
      "id": "497057",
      "postDate": "03/22/2019 22:34:39",
      "content": "<p>Congrats <a href=\"/theoviel\">@theoviel</a> and thanks for sharing. Did you also have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224\">tree based model with a measly public score</a> that scored much better on the private LB than some of your NN models? </p>",
      "rawMarkdown": "Congrats @theoviel and thanks for sharing. Did you also have a [tree based model with a measly public score](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224) that scored much better on the private LB than some of your NN models?",
      "votes": null
    },
    {
      "id": "497102",
      "postDate": "03/23/2019 00:37:55",
      "content": "<p>Not really, mostly because I used my tree models for feature selection, and the few I submitted scored Not really, my tree based models scored so poorly on LB that I decided to keep my 2 daily submissions for more promising models (i.e. NNs). My best private LB is 0.674, and the submissions I picked scored 0.667 and 0.662 private. My best Lgbm model scored 0.60 private and 0.46 public. </p>",
      "rawMarkdown": "Not really, mostly because I used my tree models for feature selection, and the few I submitted scored Not really, my tree based models scored so poorly on LB that I decided to keep my 2 daily submissions for more promising models (i.e. NNs). My best private LB is 0.674, and the submissions I picked scored 0.667 and 0.662 private. My best Lgbm model scored 0.60 private and 0.46 public.",
      "votes": null
    },
    {
      "id": "497311",
      "postDate": "03/23/2019 10:59:58",
      "content": "<p>Congrats </p>",
      "rawMarkdown": "Congrats",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496431,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 07:37:00",
      "content": "<p>Congrat Theoviel!  Thanks for sharing!</p>\n\n<p>It’s great to know that the image approach did work! I also tried using MobileNet however, in addition to features that you explained, I also included the spectrogram features (thinking that it’s more look like image haha) ....  And although it was able to fit well, it was not able to generalize to even the OOF.</p>\n\n<p>Would you mind publish your kernel if possible ? Did you just use 5 Folds CV as validation scheme ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496464,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/22/2019 08:28:35",
          "content": "<p>I don't think I'll publish my work, mostly because it'll take me too much time to fit what I did in the kaggle kernels. I also tried to feed spectrograms to my nn, but it did not work at all.</p>\n\n<p>Yup my validation scheme is just a classic 5 folds CV</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496476,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 08:50:27",
          "content": "<p>Thanks for your answer! (just asked in case that you implemented in the kernel)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496459,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "03/22/2019 08:23:36",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496535,
      "author_name": "arthurllau",
      "author_url": "",
      "post_date": "03/22/2019 09:59:18",
      "content": "<p>Congrats Theo, and thanks for sharing.\n \"Keras CuDNNLSTMs (because of the randomness), move to PyTorch\" : definitively !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496604,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "03/22/2019 11:36:03",
      "content": "<p>Nice approach Theo!    Did you do any augmentations in your training dataloader?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496656,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/22/2019 12:48:34",
          "content": "<p>Thanks Russ!\nNot at all, but there sure are loads of possibilities here!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 497057,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2019 22:34:39",
      "content": "<p>Congrats <a href=\"/theoviel\">@theoviel</a> and thanks for sharing. Did you also have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224\">tree based model with a measly public score</a> that scored much better on the private LB than some of your NN models? </p>",
      "votes": null,
      "replies": [
        {
          "id": 497102,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/23/2019 00:37:55",
          "content": "<p>Not really, mostly because I used my tree models for feature selection, and the few I submitted scored Not really, my tree based models scored so poorly on LB that I decided to keep my 2 daily submissions for more promising models (i.e. NNs). My best private LB is 0.674, and the submissions I picked scored 0.667 and 0.662 private. My best Lgbm model scored 0.60 private and 0.46 public. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 497311,
      "author_name": "azizbenothman",
      "author_url": "",
      "post_date": "03/23/2019 10:59:58",
      "content": "<p>Congrats </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496411": "Hello everybody!\n\n**The Start :**\nI joined this comptetion approx. a month ago, at the start, my main idea was to use scipy.signal to extract features about the peaks, after applying wavelet denoising. I first went with LSTMs because of the hype around high scoring public kernels, but could not get close to 0.7 LB (public).\n\n**Downfalls about public Kernels :**\nHere are a few things I believe were overused and caused the shaked_up.\n- Keras CuDNNLSTMs (because of the randomness), move to PyTorch !\n- Threshold selection for MCC optimization. It only caused overfitting on your oof data\n- Checkpoints while training your model. Again, you're overfitting on your oof data\n- Overall, I went with the idea that RNNs were predicting noise only. \n\nNot sure about the last point, because I've got one of my very first submissions that got 0.674 private with the ideas above.\n\n**To the magical submission : CV 0.659 - Public 0.652 - Private 0.667**\n\nYes that's pretty stable, that's why I trusted this model more than others.\n\nFeature Engineering : \n- Wavelet Denoizing\n- \"Bucketted\" distribution features such as in the public kernels (exactly the same actually).\n- The three phases were fed together in the network, BUT as three phases of an image. My input shape was (3, 160, nb_features). I used the most common target among the three phases as the label.\n\nI tried all the features I came up with, these ones achieved the best CV, I don't know why.\n\n**Model :** \nAs my inputs were similar to images now, I went with an architecture similar to AlexNet, but with less parameters. The architecture looks like this approx. (I changed it a bit, not sure if this is the one that got me the best score) :\n\n```\nclass Model(nn.Module):\n    def __init__(self):\n        super(Model, self).__init__()\n        \n        self.features = nn.Sequential(\n            nn.Conv2d(3, 16, kernel_size=5, padding=2),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=3, stride=2),\n            nn.Conv2d(16, 32, kernel_size=3, padding=2),\n            nn.ReLU(inplace=True),\n            nn.MaxPool2d(kernel_size=3, stride=2))\n        \n        self.avgpool = nn.AdaptiveAvgPool2d((6, 6))\n        \n        self.classifier = nn.Sequential(\n            nn.Dropout(),\n            nn.Linear(32 * 6 * 6, 32),\n            nn.ReLU(inplace=True),\n            nn.Dropout(),\n            nn.Linear(32, 1))\n\n    def forward(self, x):\n        x = self.features(x)\n        x = self.avgpool(x)\n        x = x.view(x.size(0), 32 * 6 * 6)\n        x = self.classifier(x)\n        return x\n```\n100 epochs, 5 folds CV, batch_size 128, Learning rate 0.001\n\nI worked on my gaming laptop which has 16Gb of ram and a 1060, so the 800000 long signals were quite a pain in the a** but I dealt with it.\n\n**Final Word**\nThis competition was hard because you could not trust your LB, and you could only trust your CV if it achieved results similar to LB ones. Well that's what I understood at least.\n\nThanks everyone for participating, I'm waiting to read about others' solutions.!",
    "496431": "Congrat Theoviel!  Thanks for sharing!\n\nIt’s great to know that the image approach did work! I also tried using MobileNet however, in addition to features that you explained, I also included the spectrogram features (thinking that it’s more look like image haha) ....  And although it was able to fit well, it was not able to generalize to even the OOF.\n\nWould you mind publish your kernel if possible ? Did you just use 5 Folds CV as validation scheme ?",
    "496459": "Thanks for sharing!",
    "496464": "I don't think I'll publish my work, mostly because it'll take me too much time to fit what I did in the kaggle kernels. I also tried to feed spectrograms to my nn, but it did not work at all.\n\nYup my validation scheme is just a classic 5 folds CV",
    "496476": "Thanks for your answer! (just asked in case that you implemented in the kernel)",
    "496535": "Congrats Theo, and thanks for sharing.\n \"Keras CuDNNLSTMs (because of the randomness), move to PyTorch\" : definitively !",
    "496604": "Nice approach Theo!    Did you do any augmentations in your training dataloader?",
    "496656": "Thanks Russ!\nNot at all, but there sure are loads of possibilities here!",
    "497057": "Congrats @theoviel and thanks for sharing. Did you also have a [tree based model with a measly public score](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224) that scored much better on the private LB than some of your NN models?",
    "497102": "Not really, mostly because I used my tree models for feature selection, and the few I submitted scored Not really, my tree based models scored so poorly on LB that I decided to keep my 2 daily submissions for more promising models (i.e. NNs). My best private LB is 0.674, and the submissions I picked scored 0.667 and 0.662 private. My best Lgbm model scored 0.60 private and 0.46 public.",
    "497311": "Congrats"
  },
  "source": "meta"
}