{
  "id": 85301,
  "title": "26th place solution",
  "url": "/competitions/vsb-power-line-fault-detection/writeups/michael-kazachok-26th-place-solution",
  "author_name": "",
  "post_date": "2019-03-22T20:49:30.497Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>My solution based on kernel <a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features\">VSB Competition : Base Neural Network</a>  by <a href=\"https://www.kaggle.com/tarunpaparaju\">Tarun Sriranga Paparaju</a>.  But I'm added some modification:\n- I used denoising algorithm based on wavelet(bior1.3)\n- I changed neural network architecture\n```\ndef model_gru(input_shape, feat_shape, alpha):\n    inp = Input(shape=(input_shape[1], input_shape[2],))\n    ifeat = Input(shape=(feat_shape[1],))\n    feat = Dense(8)(ifeat)\n    feat = BatchNormalization()(feat)\n    feat = Activation('tanh')(feat)</p>\n\n<pre><code>x = SpatialDropout1D(0.2)(inp)\nx = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\nx = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n\nattention = Attention(input_shape[1])(x)\n\nx = concatenate([attention, feat], axis=1)\n\nx = Dropout(0.8)(x)\n\nx = Dense(64)(x)\nx = BatchNormalization()(x)\nx = Activation('tanh')(x)\n\nx = Dense(1, activation=\"sigmoid\")(x)\n\nmodel = Model(inputs=[inp, ifeat], outputs=x)\nmodel.compile(loss=focal_loss(gamma=2, alpha=alpha), optimizer=Nadam(lr=0.003), \n</code></pre>\n\n<p>metrics=[matthews_correlation])</p>\n\n<pre><code>return model\n</code></pre>\n\n<p><code>\n-  For prevent overfitting I used strong dropout (SpatialDropout 0.2 before GRU and   dropout 0.8 before fully connected)\n-  I used focal loss with gamma = 2 and alpha calculated based on data\n</code>\nalpha = np.sqrt((1-sum(train_y)/len(train_y))*0.8))\n<code>\n- Since the signal is periodic, I used a generator with a simple augmentation.\n</code>\n    def cyclic_shift(x, alpha=0.5):\n        s = np.random.uniform(0, alpha)\n        part = int(len(x)*s)\n        x_ = x[:part, :]\n        _x = x[-len(x)+part:, :]\n        return np.concatenate([<em>x, x</em>], axis=0)\n```\n- Model was trained on StragtifiedKfolds(10 folds) used Nadam and CLR with triangle mode (clr 300 steps lr= (1e-4, 6*1e-6)) on  batch size 256 and 50 steps per epoch and early stoping with loss monitoring. </p>\n\n<p>Public score: 0.59499\nPrivate score: 0.67159</p>",
  "messages": [
    {
      "id": "496986",
      "postDate": "03/22/2019 19:56:22",
      "content": "<p>My solution based on kernel <a href=\"https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features\">VSB Competition : Base Neural Network</a>  by <a href=\"https://www.kaggle.com/tarunpaparaju\">Tarun Sriranga Paparaju</a>.  But I'm added some modification:\n- I used denoising algorithm based on wavelet(bior1.3)\n- I changed neural network architecture\n```\ndef model_gru(input_shape, feat_shape, alpha):\n    inp = Input(shape=(input_shape[1], input_shape[2],))\n    ifeat = Input(shape=(feat_shape[1],))\n    feat = Dense(8)(ifeat)\n    feat = BatchNormalization()(feat)\n    feat = Activation('tanh')(feat)</p>\n\n<pre><code>x = SpatialDropout1D(0.2)(inp)\nx = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\nx = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n\nattention = Attention(input_shape[1])(x)\n\nx = concatenate([attention, feat], axis=1)\n\nx = Dropout(0.8)(x)\n\nx = Dense(64)(x)\nx = BatchNormalization()(x)\nx = Activation('tanh')(x)\n\nx = Dense(1, activation=\"sigmoid\")(x)\n\nmodel = Model(inputs=[inp, ifeat], outputs=x)\nmodel.compile(loss=focal_loss(gamma=2, alpha=alpha), optimizer=Nadam(lr=0.003), \n</code></pre>\n\n<p>metrics=[matthews_correlation])</p>\n\n<pre><code>return model\n</code></pre>\n\n<p><code>\n-  For prevent overfitting I used strong dropout (SpatialDropout 0.2 before GRU and   dropout 0.8 before fully connected)\n-  I used focal loss with gamma = 2 and alpha calculated based on data\n</code>\nalpha = np.sqrt((1-sum(train_y)/len(train_y))*0.8))\n<code>\n- Since the signal is periodic, I used a generator with a simple augmentation.\n</code>\n    def cyclic_shift(x, alpha=0.5):\n        s = np.random.uniform(0, alpha)\n        part = int(len(x)*s)\n        x_ = x[:part, :]\n        _x = x[-len(x)+part:, :]\n        return np.concatenate([<em>x, x</em>], axis=0)\n```\n- Model was trained on StragtifiedKfolds(10 folds) used Nadam and CLR with triangle mode (clr 300 steps lr= (1e-4, 6*1e-6)) on  batch size 256 and 50 steps per epoch and early stoping with loss monitoring. </p>\n\n<p>Public score: 0.59499\nPrivate score: 0.67159</p>",
      "rawMarkdown": "My solution based on kernel [VSB Competition : Base Neural Network](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features)  by [Tarun Sriranga Paparaju](https://www.kaggle.com/tarunpaparaju).  But I'm added some modification:\n- I used denoising algorithm based on wavelet(bior1.3)\n- I changed neural network architecture\n```\ndef model_gru(input_shape, feat_shape, alpha):\n    inp = Input(shape=(input_shape[1], input_shape[2],))\n    ifeat = Input(shape=(feat_shape[1],))\n    feat = Dense(8)(ifeat)\n    feat = BatchNormalization()(feat)\n    feat = Activation('tanh')(feat)\n    \n    x = SpatialDropout1D(0.2)(inp)\n    x = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n    x = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n    \n    attention = Attention(input_shape[1])(x)\n    \n    x = concatenate([attention, feat], axis=1)\n    \n    x = Dropout(0.8)(x)\n    \n    x = Dense(64)(x)\n    x = BatchNormalization()(x)\n    x = Activation('tanh')(x)\n    \n    x = Dense(1, activation=\"sigmoid\")(x)\n    \n    model = Model(inputs=[inp, ifeat], outputs=x)\n    model.compile(loss=focal_loss(gamma=2, alpha=alpha), optimizer=Nadam(lr=0.003), \nmetrics=[matthews_correlation])\n    \n    return model\n```\n-  For prevent overfitting I used strong dropout (SpatialDropout 0.2 before GRU and   dropout 0.8 before fully connected)\n-  I used focal loss with gamma = 2 and alpha calculated based on data\n```\nalpha = np.sqrt((1-sum(train_y)/len(train_y))*0.8))\n```\n- Since the signal is periodic, I used a generator with a simple augmentation.\n```\n    def cyclic_shift(x, alpha=0.5):\n        s = np.random.uniform(0, alpha)\n        part = int(len(x)*s)\n        x_ = x[:part, :]\n        _x = x[-len(x)+part:, :]\n        return np.concatenate([_x, x_], axis=0)\n```\n- Model was trained on StragtifiedKfolds(10 folds) used Nadam and CLR with triangle mode (clr 300 steps lr= (1e-4, 6*1e-6)) on  batch size 256 and 50 steps per epoch and early stoping with loss monitoring. \n\nPublic score: 0.59499\nPrivate score: 0.67159",
      "votes": null
    },
    {
      "id": "496995",
      "postDate": "03/22/2019 20:05:27",
      "content": "<p>Thanks for sharing, interesting value of dropout. I'm surprised how the model learned despite the high value of the dropout. </p>",
      "rawMarkdown": "Thanks for sharing, interesting value of dropout. I'm surprised how the model learned despite the high value of the dropout.",
      "votes": null
    },
    {
      "id": "497049",
      "postDate": "03/22/2019 22:24:16",
      "content": "<p>Congrats <a href=\"/miklgr500\">@miklgr500</a> and thanks for sharing. </p>\n\n<p><a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224</a></p>",
      "rawMarkdown": "Congrats @miklgr500 and thanks for sharing. \n\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496995,
      "author_name": "azizbenothman",
      "author_url": "",
      "post_date": "03/22/2019 20:05:27",
      "content": "<p>Thanks for sharing, interesting value of dropout. I'm surprised how the model learned despite the high value of the dropout. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 497049,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2019 22:24:16",
      "content": "<p>Congrats <a href=\"/miklgr500\">@miklgr500</a> and thanks for sharing. </p>\n\n<p><a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496986": "My solution based on kernel [VSB Competition : Base Neural Network](https://www.kaggle.com/tarunpaparaju/vsb-competition-attention-bilstm-with-features)  by [Tarun Sriranga Paparaju](https://www.kaggle.com/tarunpaparaju).  But I'm added some modification:\n- I used denoising algorithm based on wavelet(bior1.3)\n- I changed neural network architecture\n```\ndef model_gru(input_shape, feat_shape, alpha):\n    inp = Input(shape=(input_shape[1], input_shape[2],))\n    ifeat = Input(shape=(feat_shape[1],))\n    feat = Dense(8)(ifeat)\n    feat = BatchNormalization()(feat)\n    feat = Activation('tanh')(feat)\n    \n    x = SpatialDropout1D(0.2)(inp)\n    x = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n    x = Bidirectional(CuDNNGRU(160, return_sequences=True))(x)\n    \n    attention = Attention(input_shape[1])(x)\n    \n    x = concatenate([attention, feat], axis=1)\n    \n    x = Dropout(0.8)(x)\n    \n    x = Dense(64)(x)\n    x = BatchNormalization()(x)\n    x = Activation('tanh')(x)\n    \n    x = Dense(1, activation=\"sigmoid\")(x)\n    \n    model = Model(inputs=[inp, ifeat], outputs=x)\n    model.compile(loss=focal_loss(gamma=2, alpha=alpha), optimizer=Nadam(lr=0.003), \nmetrics=[matthews_correlation])\n    \n    return model\n```\n-  For prevent overfitting I used strong dropout (SpatialDropout 0.2 before GRU and   dropout 0.8 before fully connected)\n-  I used focal loss with gamma = 2 and alpha calculated based on data\n```\nalpha = np.sqrt((1-sum(train_y)/len(train_y))*0.8))\n```\n- Since the signal is periodic, I used a generator with a simple augmentation.\n```\n    def cyclic_shift(x, alpha=0.5):\n        s = np.random.uniform(0, alpha)\n        part = int(len(x)*s)\n        x_ = x[:part, :]\n        _x = x[-len(x)+part:, :]\n        return np.concatenate([_x, x_], axis=0)\n```\n- Model was trained on StragtifiedKfolds(10 folds) used Nadam and CLR with triangle mode (clr 300 steps lr= (1e-4, 6*1e-6)) on  batch size 256 and 50 steps per epoch and early stoping with loss monitoring. \n\nPublic score: 0.59499\nPrivate score: 0.67159",
    "496995": "Thanks for sharing, interesting value of dropout. I'm surprised how the model learned despite the high value of the dropout.",
    "497049": "Congrats @miklgr500 and thanks for sharing. \n\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146#496224"
  },
  "source": "meta"
}