{
  "id": 85163,
  "title": "44th solution... Thank you for your questions and advices",
  "url": "/competitions/vsb-power-line-fault-detection/writeups/little-snail-44th-solution-thank-you-for-your-ques",
  "author_name": "",
  "post_date": "2019-03-22T03:33:53.290Z",
  "votes": 19,
  "comment_count": 7,
  "views": 0,
  "content": "<p>My public lb ranks 107th and private lb ranks 44th.\nThe train set of this competition is very small. And the distribution of train set is very different from the distribution of LB dataset. So I don't think local cv can display your models' performance on private LB. Maybe this competition's key is fitting public LB without overfitting.</p>\n\n<p>My solution is as below:</p>\n\n<ul>\n<li><p>feature extraction\nMy FE is similar as this <a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694/data\">kenel</a>. But there is some difference as below. </p>\n\n<p>I exclude the trend of signal before feature extraction as below:\n<code>\nrollingdata = rawdata.rolling(20000, center=True, min_periods=1, axis=0)\ntrend = rollingdata.mean() <br>\nres = (rawdata - trend).values\n</code></p>\n\n<p>In order to mine more information from raw signal, I divide raw signal into 250 segments.\n<code>\nwhole_signal_len = 800000\nprocessed_len = 250\nbin_len = int(whole_signal_len/processed_len)\n</code></p>\n\n<p>And I don't use 25% and 75% quantile. Because I think trend is useless for PD detection, and some quantiles are redund ant.\n```\nmean_slice = signal_slice.mean(axis=0)\nstd_slice = signal_slice.std(axis=0)</p></li>\n</ul>\n\n<p>std_top = mean_slice + std_slice\nstd_bot = mean_slice - std_slice</p>\n\n<p>percentil_slice = np.percentile(signal_slice, [0, 1, 50, 99, 100], axis=0)\nmax_range = percentil_slice[-1] - percentil_slice[0]</p>\n\n<p>relative_percentile = percentil_slice - mean_slice</p>\n\n<p>out_slice = np.concatenate([\n                            percentil_slice.T,\n                            max_range.reshape(-1, 1),\n                            relative_percentile.T,\n                            std_top.reshape(-1, 1),\n                            std_bot.reshape(-1, 1),\n                            std_slice.reshape(-1, 1),\n                            mean_slice.reshape(-1, 1)\n                            ], axis=1)\n```</p>\n\n<ul>\n<li><p>model\nThere are two models in my solution.\n```\nclass LSTM_softmax(nn.Module, myBaseModule):\ndef <strong>init</strong>(self, random_seed,\n                   features_dims,\n                   seq_len,\n                   learning_rate,\n                   lstm_out_dim=80,\n                   lstm_layers=2,\n                   linearReduction_dim=64,\n                   env=None):\n    nn.Module.<strong>init</strong>(self)\n    myBaseModule.<strong>init</strong>(self, random_seed)</p>\n\n<pre><code>self.l2_weight = 0.0000\nself.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\nself.lstm_attention = Attention(lstm_out_dim*2, seq_len)\n\nself.linear = nn.Linear(lstm_out_dim*4, linearReduction_dim).cuda()\nself.relu = nn.ReLU()\nself.dropout = nn.Dropout(0.10)\nself.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\nself.sigmoid = nn.Sigmoid()\n\nself.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\nself.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\nself.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\nself.vis = env\n</code></pre>\n\n<p>def forward(self, x):\n    h_lstm, _ = self.lstm(x)\n    h_lstm_atten = self.lstm_attention(h_lstm)</p>\n\n<pre><code>lstm_max_pool, _ = torch.max(h_lstm, 1)\n\nnn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\nnn_out = self.dropout(nn_out)\nnn_out = self.relu(self.linear(nn_out))\nout = self.sigmoid(self.linear2(nn_out))\n\nreturn out\n</code></pre></li>\n</ul>\n\n<p><code>\n</code>\nclass LSTM_selfAttention_softmax(nn.Module, myBaseModule):\n    def <strong>init</strong>(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       selfAttention_dim=80,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.<strong>init</strong>(self)\n        myBaseModule.<strong>init</strong>(self, random_seed)</p>\n\n<pre><code>    self.l2_weight = 0.0000\n\n    self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n    self.selfAttention = selfAttention(selfAttention_dim, selfAttention_dim, 2*lstm_out_dim,dk=64)\n\n    self.lstm_attention = Attention(selfAttention_dim, seq_len)\n\n    self.linear = nn.Linear(2*(selfAttention_dim), linearReduction_dim).cuda()\n    self.relu = nn.ReLU()\n    self.dropout = nn.Dropout(0.10)\n    self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n    self.sigmoid = nn.Sigmoid()\n\n    self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n    self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n    self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n    self.vis = env\n\ndef forward(self, x):\n\n    h_lstm, _ = self.lstm(x)\n    h_lstm = self.selfAttention(h_lstm)\n\n    h_lstm_atten = self.lstm_attention(h_lstm)\n    lstm_max_pool, _ = torch.max(h_lstm, 1)\n    nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n\n\n    nn_out = self.dropout(nn_out)\n    nn_out = self.relu(self.linear(nn_out))\n    out = self.sigmoid(self.linear2(nn_out))\n    return out\n</code></pre>\n\n<p>```\nIn order to reduce model complexity, I train 6 models for 3 phase data. Each phase data has 2 models. The final result is a weighted average of predictions of two models.</p>\n\n<p><a href=\"https://github.com/qq563902455/VSB_Power_Line_Fault_Detection\">github link</a>\nIf you think this is helpful, please star my github project and upvote my discussion.Thank you!!!</p>",
  "messages": [
    {
      "id": "496252",
      "postDate": "03/22/2019 02:17:33",
      "content": "<p>My public lb ranks 107th and private lb ranks 44th.\nThe train set of this competition is very small. And the distribution of train set is very different from the distribution of LB dataset. So I don't think local cv can display your models' performance on private LB. Maybe this competition's key is fitting public LB without overfitting.</p>\n\n<p>My solution is as below:</p>\n\n<ul>\n<li><p>feature extraction\nMy FE is similar as this <a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694/data\">kenel</a>. But there is some difference as below. </p>\n\n<p>I exclude the trend of signal before feature extraction as below:\n<code>\nrollingdata = rawdata.rolling(20000, center=True, min_periods=1, axis=0)\ntrend = rollingdata.mean() <br>\nres = (rawdata - trend).values\n</code></p>\n\n<p>In order to mine more information from raw signal, I divide raw signal into 250 segments.\n<code>\nwhole_signal_len = 800000\nprocessed_len = 250\nbin_len = int(whole_signal_len/processed_len)\n</code></p>\n\n<p>And I don't use 25% and 75% quantile. Because I think trend is useless for PD detection, and some quantiles are redund ant.\n```\nmean_slice = signal_slice.mean(axis=0)\nstd_slice = signal_slice.std(axis=0)</p></li>\n</ul>\n\n<p>std_top = mean_slice + std_slice\nstd_bot = mean_slice - std_slice</p>\n\n<p>percentil_slice = np.percentile(signal_slice, [0, 1, 50, 99, 100], axis=0)\nmax_range = percentil_slice[-1] - percentil_slice[0]</p>\n\n<p>relative_percentile = percentil_slice - mean_slice</p>\n\n<p>out_slice = np.concatenate([\n                            percentil_slice.T,\n                            max_range.reshape(-1, 1),\n                            relative_percentile.T,\n                            std_top.reshape(-1, 1),\n                            std_bot.reshape(-1, 1),\n                            std_slice.reshape(-1, 1),\n                            mean_slice.reshape(-1, 1)\n                            ], axis=1)\n```</p>\n\n<ul>\n<li><p>model\nThere are two models in my solution.\n```\nclass LSTM_softmax(nn.Module, myBaseModule):\ndef <strong>init</strong>(self, random_seed,\n                   features_dims,\n                   seq_len,\n                   learning_rate,\n                   lstm_out_dim=80,\n                   lstm_layers=2,\n                   linearReduction_dim=64,\n                   env=None):\n    nn.Module.<strong>init</strong>(self)\n    myBaseModule.<strong>init</strong>(self, random_seed)</p>\n\n<pre><code>self.l2_weight = 0.0000\nself.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\nself.lstm_attention = Attention(lstm_out_dim*2, seq_len)\n\nself.linear = nn.Linear(lstm_out_dim*4, linearReduction_dim).cuda()\nself.relu = nn.ReLU()\nself.dropout = nn.Dropout(0.10)\nself.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\nself.sigmoid = nn.Sigmoid()\n\nself.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\nself.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\nself.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\nself.vis = env\n</code></pre>\n\n<p>def forward(self, x):\n    h_lstm, _ = self.lstm(x)\n    h_lstm_atten = self.lstm_attention(h_lstm)</p>\n\n<pre><code>lstm_max_pool, _ = torch.max(h_lstm, 1)\n\nnn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\nnn_out = self.dropout(nn_out)\nnn_out = self.relu(self.linear(nn_out))\nout = self.sigmoid(self.linear2(nn_out))\n\nreturn out\n</code></pre></li>\n</ul>\n\n<p><code>\n</code>\nclass LSTM_selfAttention_softmax(nn.Module, myBaseModule):\n    def <strong>init</strong>(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       selfAttention_dim=80,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.<strong>init</strong>(self)\n        myBaseModule.<strong>init</strong>(self, random_seed)</p>\n\n<pre><code>    self.l2_weight = 0.0000\n\n    self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n    self.selfAttention = selfAttention(selfAttention_dim, selfAttention_dim, 2*lstm_out_dim,dk=64)\n\n    self.lstm_attention = Attention(selfAttention_dim, seq_len)\n\n    self.linear = nn.Linear(2*(selfAttention_dim), linearReduction_dim).cuda()\n    self.relu = nn.ReLU()\n    self.dropout = nn.Dropout(0.10)\n    self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n    self.sigmoid = nn.Sigmoid()\n\n    self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n    self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n    self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n    self.vis = env\n\ndef forward(self, x):\n\n    h_lstm, _ = self.lstm(x)\n    h_lstm = self.selfAttention(h_lstm)\n\n    h_lstm_atten = self.lstm_attention(h_lstm)\n    lstm_max_pool, _ = torch.max(h_lstm, 1)\n    nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n\n\n    nn_out = self.dropout(nn_out)\n    nn_out = self.relu(self.linear(nn_out))\n    out = self.sigmoid(self.linear2(nn_out))\n    return out\n</code></pre>\n\n<p>```\nIn order to reduce model complexity, I train 6 models for 3 phase data. Each phase data has 2 models. The final result is a weighted average of predictions of two models.</p>\n\n<p><a href=\"https://github.com/qq563902455/VSB_Power_Line_Fault_Detection\">github link</a>\nIf you think this is helpful, please star my github project and upvote my discussion.Thank you!!!</p>",
      "rawMarkdown": "My public lb ranks 107th and private lb ranks 44th.\nThe train set of this competition is very small. And the distribution of train set is very different from the distribution of LB dataset. So I don't think local cv can display your models' performance on private LB. Maybe this competition's key is fitting public LB without overfitting.\n\nMy solution is as below:\n\n- feature extraction\n  My FE is similar as this [kenel](https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694/data). But there is some difference as below. \n\n  I exclude the trend of signal before feature extraction as below:\n``` \nrollingdata = rawdata.rolling(20000, center=True, min_periods=1, axis=0)\ntrend = rollingdata.mean()  \nres = (rawdata - trend).values\n```\n\n  In order to mine more information from raw signal, I divide raw signal into 250 segments.\n```\n    whole_signal_len = 800000\n    processed_len = 250\n    bin_len = int(whole_signal_len/processed_len)\n```\n\n  And I don't use 25% and 75% quantile. Because I think trend is useless for PD detection, and some quantiles are redund ant.\n```\nmean_slice = signal_slice.mean(axis=0)\nstd_slice = signal_slice.std(axis=0)\n\nstd_top = mean_slice + std_slice\nstd_bot = mean_slice - std_slice\n\npercentil_slice = np.percentile(signal_slice, [0, 1, 50, 99, 100], axis=0)\nmax_range = percentil_slice[-1] - percentil_slice[0]\n\nrelative_percentile = percentil_slice - mean_slice\n\n\nout_slice = np.concatenate([\n                            percentil_slice.T,\n                            max_range.reshape(-1, 1),\n                            relative_percentile.T,\n                            std_top.reshape(-1, 1),\n                            std_bot.reshape(-1, 1),\n                            std_slice.reshape(-1, 1),\n                            mean_slice.reshape(-1, 1)\n                            ], axis=1)\n```\n\n- model\nThere are two models in my solution.\n```\nclass LSTM_softmax(nn.Module, myBaseModule):\n    def __init__(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.__init__(self)\n        myBaseModule.__init__(self, random_seed)\n\n        self.l2_weight = 0.0000\n        self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n        self.lstm_attention = Attention(lstm_out_dim*2, seq_len)\n\n        self.linear = nn.Linear(lstm_out_dim*4, linearReduction_dim).cuda()\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.10)\n        self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n        self.sigmoid = nn.Sigmoid()\n\n        self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n        self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n        self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n        self.vis = env\n\n    def forward(self, x):\n        h_lstm, _ = self.lstm(x)\n        h_lstm_atten = self.lstm_attention(h_lstm)\n\n        lstm_max_pool, _ = torch.max(h_lstm, 1)\n\n        nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n        nn_out = self.dropout(nn_out)\n        nn_out = self.relu(self.linear(nn_out))\n        out = self.sigmoid(self.linear2(nn_out))\n\n        return out\n\n```\n```\nclass LSTM_selfAttention_softmax(nn.Module, myBaseModule):\n    def __init__(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       selfAttention_dim=80,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.__init__(self)\n        myBaseModule.__init__(self, random_seed)\n\n        self.l2_weight = 0.0000\n\n        self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n        self.selfAttention = selfAttention(selfAttention_dim, selfAttention_dim, 2*lstm_out_dim,dk=64)\n\n        self.lstm_attention = Attention(selfAttention_dim, seq_len)\n\n        self.linear = nn.Linear(2*(selfAttention_dim), linearReduction_dim).cuda()\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.10)\n        self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n        self.sigmoid = nn.Sigmoid()\n\n        self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n        self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n        self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n        self.vis = env\n\n    def forward(self, x):\n\n        h_lstm, _ = self.lstm(x)\n        h_lstm = self.selfAttention(h_lstm)\n\n        h_lstm_atten = self.lstm_attention(h_lstm)\n        lstm_max_pool, _ = torch.max(h_lstm, 1)\n        nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n\n\n        nn_out = self.dropout(nn_out)\n        nn_out = self.relu(self.linear(nn_out))\n        out = self.sigmoid(self.linear2(nn_out))\n        return out\n```\nIn order to reduce model complexity, I train 6 models for 3 phase data. Each phase data has 2 models. The final result is a weighted average of predictions of two models.\n\n\n\n[github link](https://github.com/qq563902455/VSB_Power_Line_Fault_Detection)\nIf you think this is helpful, please star my github project and upvote my discussion.Thank you!!!",
      "votes": null
    },
    {
      "id": "496301",
      "postDate": "03/22/2019 03:10:09",
      "content": "<p>excellent！</p>",
      "rawMarkdown": "excellent！",
      "votes": null
    },
    {
      "id": "496302",
      "postDate": "03/22/2019 03:15:31",
      "content": "<p>Thank you so much <a href=\"/blackboards\">@blackboards</a> , to share your model and insight!!</p>\n\n<p>You spoke my mind on the issue CV vs. LB by saying that we should trust Public LB without trying to overfit it.</p>\n\n<p>Would you mind elaborate on how to do this, i.e. how were you able to make a progress on improving your model performance without too much overfit the LB ? How do you know that you were not too much overfith the LB?</p>",
      "rawMarkdown": "Thank you so much @blackboards , to share your model and insight!!\n\nYou spoke my mind on the issue CV vs. LB by saying that we should trust Public LB without trying to overfit it.\n\nWould you mind elaborate on how to do this, i.e. how were you able to make a progress on improving your model performance without too much overfit the LB ? How do you know that you were not too much overfith the LB?",
      "votes": null
    },
    {
      "id": "496312",
      "postDate": "03/22/2019 03:30:38",
      "content": "<p>I think the key of avoiding overfiting is feature extraction.\nSo we need to extract features based on 'Prior Knowledge', such as <a href=\"https://ieeexplore.ieee.org/document/7909221/\">paper</a></p>\n\n<p>Specifically, in my solution, I exclude trend from raw data.\nSecondly, I think simple models have better generalization capabilities. So it may be a wrong decision to use complex model to improve public LB score.</p>",
      "rawMarkdown": "I think the key of avoiding overfiting is feature extraction.\nSo we need to extract features based on 'Prior Knowledge', such as [paper](https://ieeexplore.ieee.org/document/7909221/)\n\nSpecifically, in my solution, I exclude trend from raw data.\nSecondly, I think simple models have better generalization capabilities. So it may be a wrong decision to use complex model to improve public LB score.",
      "votes": null
    },
    {
      "id": "496319",
      "postDate": "03/22/2019 03:37:06",
      "content": "<p>Thanks again for sharing!</p>",
      "rawMarkdown": "Thanks again for sharing!",
      "votes": null
    },
    {
      "id": "496325",
      "postDate": "03/22/2019 03:46:19",
      "content": "<p>Congrats <a href=\"/blackboards\">@blackboards</a> and thanks for sharing your code.</p>\n\n<p>On the CV strategy, I have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">discussion post here</a> if you want to add your thoughts.</p>",
      "rawMarkdown": "Congrats @blackboards and thanks for sharing your code.\n\nOn the CV strategy, I have a [discussion post here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143) if you want to add your thoughts.",
      "votes": null
    },
    {
      "id": "496387",
      "postDate": "03/22/2019 05:44:12",
      "content": "<p>Nice! Thanks for sharing.</p>",
      "rawMarkdown": "Nice! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "496462",
      "postDate": "03/22/2019 08:26:32",
      "content": "<p>Thanks for the writeup!</p>",
      "rawMarkdown": "Thanks for the writeup!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496301,
      "author_name": "xu627586640",
      "author_url": "",
      "post_date": "03/22/2019 03:10:09",
      "content": "<p>excellent！</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496302,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 03:15:31",
      "content": "<p>Thank you so much <a href=\"/blackboards\">@blackboards</a> , to share your model and insight!!</p>\n\n<p>You spoke my mind on the issue CV vs. LB by saying that we should trust Public LB without trying to overfit it.</p>\n\n<p>Would you mind elaborate on how to do this, i.e. how were you able to make a progress on improving your model performance without too much overfit the LB ? How do you know that you were not too much overfith the LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496312,
          "author_name": "blackboards",
          "author_url": "",
          "post_date": "03/22/2019 03:30:38",
          "content": "<p>I think the key of avoiding overfiting is feature extraction.\nSo we need to extract features based on 'Prior Knowledge', such as <a href=\"https://ieeexplore.ieee.org/document/7909221/\">paper</a></p>\n\n<p>Specifically, in my solution, I exclude trend from raw data.\nSecondly, I think simple models have better generalization capabilities. So it may be a wrong decision to use complex model to improve public LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496319,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 03:37:06",
          "content": "<p>Thanks again for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496325,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2019 03:46:19",
      "content": "<p>Congrats <a href=\"/blackboards\">@blackboards</a> and thanks for sharing your code.</p>\n\n<p>On the CV strategy, I have a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">discussion post here</a> if you want to add your thoughts.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496387,
      "author_name": "dmitrsl",
      "author_url": "",
      "post_date": "03/22/2019 05:44:12",
      "content": "<p>Nice! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496462,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "03/22/2019 08:26:32",
      "content": "<p>Thanks for the writeup!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496252": "My public lb ranks 107th and private lb ranks 44th.\nThe train set of this competition is very small. And the distribution of train set is very different from the distribution of LB dataset. So I don't think local cv can display your models' performance on private LB. Maybe this competition's key is fitting public LB without overfitting.\n\nMy solution is as below:\n\n- feature extraction\n  My FE is similar as this [kenel](https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694/data). But there is some difference as below. \n\n  I exclude the trend of signal before feature extraction as below:\n``` \nrollingdata = rawdata.rolling(20000, center=True, min_periods=1, axis=0)\ntrend = rollingdata.mean()  \nres = (rawdata - trend).values\n```\n\n  In order to mine more information from raw signal, I divide raw signal into 250 segments.\n```\n    whole_signal_len = 800000\n    processed_len = 250\n    bin_len = int(whole_signal_len/processed_len)\n```\n\n  And I don't use 25% and 75% quantile. Because I think trend is useless for PD detection, and some quantiles are redund ant.\n```\nmean_slice = signal_slice.mean(axis=0)\nstd_slice = signal_slice.std(axis=0)\n\nstd_top = mean_slice + std_slice\nstd_bot = mean_slice - std_slice\n\npercentil_slice = np.percentile(signal_slice, [0, 1, 50, 99, 100], axis=0)\nmax_range = percentil_slice[-1] - percentil_slice[0]\n\nrelative_percentile = percentil_slice - mean_slice\n\n\nout_slice = np.concatenate([\n                            percentil_slice.T,\n                            max_range.reshape(-1, 1),\n                            relative_percentile.T,\n                            std_top.reshape(-1, 1),\n                            std_bot.reshape(-1, 1),\n                            std_slice.reshape(-1, 1),\n                            mean_slice.reshape(-1, 1)\n                            ], axis=1)\n```\n\n- model\nThere are two models in my solution.\n```\nclass LSTM_softmax(nn.Module, myBaseModule):\n    def __init__(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.__init__(self)\n        myBaseModule.__init__(self, random_seed)\n\n        self.l2_weight = 0.0000\n        self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n        self.lstm_attention = Attention(lstm_out_dim*2, seq_len)\n\n        self.linear = nn.Linear(lstm_out_dim*4, linearReduction_dim).cuda()\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.10)\n        self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n        self.sigmoid = nn.Sigmoid()\n\n        self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n        self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n        self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n        self.vis = env\n\n    def forward(self, x):\n        h_lstm, _ = self.lstm(x)\n        h_lstm_atten = self.lstm_attention(h_lstm)\n\n        lstm_max_pool, _ = torch.max(h_lstm, 1)\n\n        nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n        nn_out = self.dropout(nn_out)\n        nn_out = self.relu(self.linear(nn_out))\n        out = self.sigmoid(self.linear2(nn_out))\n\n        return out\n\n```\n```\nclass LSTM_selfAttention_softmax(nn.Module, myBaseModule):\n    def __init__(self, random_seed,\n                       features_dims,\n                       seq_len,\n                       learning_rate,\n                       lstm_out_dim=80,\n                       lstm_layers=2,\n                       selfAttention_dim=80,\n                       linearReduction_dim=64,\n                       env=None):\n        nn.Module.__init__(self)\n        myBaseModule.__init__(self, random_seed)\n\n        self.l2_weight = 0.0000\n\n        self.lstm = nn.LSTM(features_dims, lstm_out_dim, lstm_layers, bidirectional=True, batch_first=True).cuda()\n        self.selfAttention = selfAttention(selfAttention_dim, selfAttention_dim, 2*lstm_out_dim,dk=64)\n\n        self.lstm_attention = Attention(selfAttention_dim, seq_len)\n\n        self.linear = nn.Linear(2*(selfAttention_dim), linearReduction_dim).cuda()\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.10)\n        self.linear2 = nn.Linear(linearReduction_dim, 1).cuda()\n        self.sigmoid = nn.Sigmoid()\n\n        self.optimizer = torch.optim.Adam(self.parameters(), lr=learning_rate)\n        self.scheduler = torch.optim.lr_scheduler.LambdaLR(self.optimizer, lambda epoch_num: 1/math.sqrt(epoch_num+1))\n        self.loss_fn = torch.nn.BCELoss(reduction=\"sum\")\n\n        self.vis = env\n\n    def forward(self, x):\n\n        h_lstm, _ = self.lstm(x)\n        h_lstm = self.selfAttention(h_lstm)\n\n        h_lstm_atten = self.lstm_attention(h_lstm)\n        lstm_max_pool, _ = torch.max(h_lstm, 1)\n        nn_out = torch.cat([lstm_max_pool, h_lstm_atten], 1)\n\n\n        nn_out = self.dropout(nn_out)\n        nn_out = self.relu(self.linear(nn_out))\n        out = self.sigmoid(self.linear2(nn_out))\n        return out\n```\nIn order to reduce model complexity, I train 6 models for 3 phase data. Each phase data has 2 models. The final result is a weighted average of predictions of two models.\n\n\n\n[github link](https://github.com/qq563902455/VSB_Power_Line_Fault_Detection)\nIf you think this is helpful, please star my github project and upvote my discussion.Thank you!!!",
    "496301": "excellent！",
    "496302": "Thank you so much @blackboards , to share your model and insight!!\n\nYou spoke my mind on the issue CV vs. LB by saying that we should trust Public LB without trying to overfit it.\n\nWould you mind elaborate on how to do this, i.e. how were you able to make a progress on improving your model performance without too much overfit the LB ? How do you know that you were not too much overfith the LB?",
    "496312": "I think the key of avoiding overfiting is feature extraction.\nSo we need to extract features based on 'Prior Knowledge', such as [paper](https://ieeexplore.ieee.org/document/7909221/)\n\nSpecifically, in my solution, I exclude trend from raw data.\nSecondly, I think simple models have better generalization capabilities. So it may be a wrong decision to use complex model to improve public LB score.",
    "496319": "Thanks again for sharing!",
    "496325": "Congrats @blackboards and thanks for sharing your code.\n\nOn the CV strategy, I have a [discussion post here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143) if you want to add your thoughts.",
    "496387": "Nice! Thanks for sharing.",
    "496462": "Thanks for the writeup!"
  },
  "source": "meta"
}