{
  "id": 85183,
  "title": "We had a 4th position submission and we didn't know it",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85183",
  "author_name": "",
  "post_date": "2019-03-22T06:32:13.151273900Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We still have to make head or tails of this competition. We had a 4th place submission but we couldn't recognize it among the others (which scored equally good), no matter how much we kept trace of the local cv. Moreover it was powered by a dnn which we considered an overfitter given its complexity:</p>\n\n<p><code>\ndef model_lstm_A2(input_shape, ex_shape):\n   inp = Input(shape=(input_shape[1], input_shape[2],))\n   ex  = Input(shape=(ex_shape[1],))\n   x = Bidirectional(CuDNNLSTM(128, return_sequences=True))(inp)\n   max_pool1 = GlobalMaxPooling1D()(x)\n   avg_pool1 = GlobalAveragePooling1D()(x)\n   x = Bidirectional(CuDNNLSTM(64, return_sequences=True))(x)\n   attention = Attention(input_shape[1])(x)\n   conv = Conv1D(filters=16, kernel_size=4,\n                 padding='valid', kernel_initializer='he_uniform')(x)\n   avg_pool2 = GlobalAveragePooling1D()(conv)\n   max_pool2 = GlobalMaxPooling1D()(conv)\n   capsule = Capsule(num_capsule=10, dim_capsule=10, routings=4)(x)\n   capsule = Flatten()(capsule)\n   x = Concatenate()([max_pool1, avg_pool1, max_pool2, avg_pool2, attention, capsule, ex])\n   x = BatchNormalization()(x)\n   x = Dense(64, activation=\"relu\")(x)\n   x = Dropout(0.3)(x)\n   x = Dense(1, activation=\"sigmoid\")(x)\n   model = Model(inputs=[inp, ex], outputs=x)\n   model.compile(loss='binary_crossentropy',\n                 optimizer=Adam(lr=LR),\n                 metrics=[matthews_correlation])\n   return model\n</code></p>\n\n<p>Anyway we did lear a lot in this competition, really. I even implemented my own attention layer because I didn't trust the one on the kernels :-) (actually the one on the public kernels is ok ;-)).</p>\n\n<p>A big thank you to @pietromarinelli and @deveaup for their great collaborations and endless ideas!</p>",
  "messages": [
    {
      "id": "496404",
      "postDate": "03/22/2019 06:32:13",
      "content": "<p>We still have to make head or tails of this competition. We had a 4th place submission but we couldn't recognize it among the others (which scored equally good), no matter how much we kept trace of the local cv. Moreover it was powered by a dnn which we considered an overfitter given its complexity:</p>\n\n<p><code>\ndef model_lstm_A2(input_shape, ex_shape):\n   inp = Input(shape=(input_shape[1], input_shape[2],))\n   ex  = Input(shape=(ex_shape[1],))\n   x = Bidirectional(CuDNNLSTM(128, return_sequences=True))(inp)\n   max_pool1 = GlobalMaxPooling1D()(x)\n   avg_pool1 = GlobalAveragePooling1D()(x)\n   x = Bidirectional(CuDNNLSTM(64, return_sequences=True))(x)\n   attention = Attention(input_shape[1])(x)\n   conv = Conv1D(filters=16, kernel_size=4,\n                 padding='valid', kernel_initializer='he_uniform')(x)\n   avg_pool2 = GlobalAveragePooling1D()(conv)\n   max_pool2 = GlobalMaxPooling1D()(conv)\n   capsule = Capsule(num_capsule=10, dim_capsule=10, routings=4)(x)\n   capsule = Flatten()(capsule)\n   x = Concatenate()([max_pool1, avg_pool1, max_pool2, avg_pool2, attention, capsule, ex])\n   x = BatchNormalization()(x)\n   x = Dense(64, activation=\"relu\")(x)\n   x = Dropout(0.3)(x)\n   x = Dense(1, activation=\"sigmoid\")(x)\n   model = Model(inputs=[inp, ex], outputs=x)\n   model.compile(loss='binary_crossentropy',\n                 optimizer=Adam(lr=LR),\n                 metrics=[matthews_correlation])\n   return model\n</code></p>\n\n<p>Anyway we did lear a lot in this competition, really. I even implemented my own attention layer because I didn't trust the one on the kernels :-) (actually the one on the public kernels is ok ;-)).</p>\n\n<p>A big thank you to @pietromarinelli and @deveaup for their great collaborations and endless ideas!</p>",
      "rawMarkdown": "We still have to make head or tails of this competition. We had a 4th place submission but we couldn't recognize it among the others (which scored equally good), no matter how much we kept trace of the local cv. Moreover it was powered by a dnn which we considered an overfitter given its complexity:\n\n```\ndef model_lstm_A2(input_shape, ex_shape):\n   inp = Input(shape=(input_shape[1], input_shape[2],))\n   ex  = Input(shape=(ex_shape[1],))\n   x = Bidirectional(CuDNNLSTM(128, return_sequences=True))(inp)\n   max_pool1 = GlobalMaxPooling1D()(x)\n   avg_pool1 = GlobalAveragePooling1D()(x)\n   x = Bidirectional(CuDNNLSTM(64, return_sequences=True))(x)\n   attention = Attention(input_shape[1])(x)\n   conv = Conv1D(filters=16, kernel_size=4,\n                 padding='valid', kernel_initializer='he_uniform')(x)\n   avg_pool2 = GlobalAveragePooling1D()(conv)\n   max_pool2 = GlobalMaxPooling1D()(conv)\n   capsule = Capsule(num_capsule=10, dim_capsule=10, routings=4)(x)\n   capsule = Flatten()(capsule)\n   x = Concatenate()([max_pool1, avg_pool1, max_pool2, avg_pool2, attention, capsule, ex])\n   x = BatchNormalization()(x)\n   x = Dense(64, activation=\"relu\")(x)\n   x = Dropout(0.3)(x)\n   x = Dense(1, activation=\"sigmoid\")(x)\n   model = Model(inputs=[inp, ex], outputs=x)\n   model.compile(loss='binary_crossentropy',\n                 optimizer=Adam(lr=LR),\n                 metrics=[matthews_correlation])\n   return model\n```\n\nAnyway we did lear a lot in this competition, really. I even implemented my own attention layer because I didn't trust the one on the kernels :-) (actually the one on the public kernels is ok ;-)).\n\nA big thank you to @pietromarinelli and @deveaup for their great collaborations and endless ideas!",
      "votes": null
    },
    {
      "id": "496430",
      "postDate": "03/22/2019 07:32:23",
      "content": "<p>Thanks <a href=\"/lucamassaron\">@lucamassaron</a> for sharing. Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features ?)</p>\n\n<p>More question, how did you implement your CV strategy, did it correlate OK with public/private ?</p>",
      "rawMarkdown": "Thanks @lucamassaron for sharing. Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features ?)\n\nMore question, how did you implement your CV strategy, did it correlate OK with public/private ?",
      "votes": null
    },
    {
      "id": "496461",
      "postDate": "03/22/2019 08:25:42",
      "content": "<p>Thanks for sharing!\nI tried to play with pooling layers too but in th end i rejected this idea...\nAlso had a 20th place submission but didn't use it as final</p>",
      "rawMarkdown": "Thanks for sharing!\nI tried to play with pooling layers too but in th end i rejected this idea...\nAlso had a 20th place submission but didn't use it as final",
      "votes": null
    },
    {
      "id": "496585",
      "postDate": "03/22/2019 11:13:37",
      "content": "<p>Thanks for sharing Luca.   I'm wondering if there may be some hidden location structure behind the training/public/private splits.   Interesting idea to use a Capsule layer.   Did you pull it from <a href=\"https://github.com/XifengGuo/CapsNet-Keras\">https://github.com/XifengGuo/CapsNet-Keras</a>?</p>",
      "rawMarkdown": "Thanks for sharing Luca.   I'm wondering if there may be some hidden location structure behind the training/public/private splits.   Interesting idea to use a Capsule layer.   Did you pull it from https://github.com/XifengGuo/CapsNet-Keras?",
      "votes": null
    },
    {
      "id": "496603",
      "postDate": "03/22/2019 11:35:59",
      "content": "<p><a href=\"/lucamassaron\">@lucamassaron</a>, thanks for sharing. Though not as high as yours,  I missed a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">76th place catboost submission</a> myself and there is no way I would have chosen that because both CV and public LB scores were really low.</p>\n\n<p>If you have any insights about a better CV setup for this competition, would you care to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">share it here</a>?</p>",
      "rawMarkdown": "lucamassaron, thanks for sharing. Though not as high as yours,  I missed a [76th place catboost submission](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) myself and there is no way I would have chosen that because both CV and public LB scores were really low.\n\nIf you have any insights about a better CV setup for this competition, would you care to [share it here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143)?",
      "votes": null
    },
    {
      "id": "496681",
      "postDate": "03/22/2019 13:16:25",
      "content": "<p>Hi Russ,\nwe exactly used Xifeng Guo's implementation :-) In our opinion, anyway, there is some trouble in the test set. It is problematic that a sub-sample of it provides so different results than the complete set.</p>",
      "rawMarkdown": "Hi Russ,\nwe exactly used Xifeng Guo's implementation :-) In our opinion, anyway, there is some trouble in the test set. It is problematic that a sub-sample of it provides so different results than the complete set.",
      "votes": null
    },
    {
      "id": "496688",
      "postDate": "03/22/2019 13:23:50",
      "content": "<p>We actually noticed little correlation between our submissions and the LB, so we couldn't make the right decisions at the end of the competition.</p>",
      "rawMarkdown": "We actually noticed little correlation between our submissions and the LB, so we couldn't make the right decisions at the end of the competition.",
      "votes": null
    },
    {
      "id": "496806",
      "postDate": "03/22/2019 15:53:01",
      "content": "<p>Thanks Luca <a href=\"/lucamassaron\">@lucamassaron</a>!  Would you mind answering this question?\n&gt; Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features = 2904 x 160 x 57 ?  )</p>",
      "rawMarkdown": "Thanks Luca @lucamassaron!  Would you mind answering this question?\n&gt; Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features = 2904 x 160 x 57 ?  )",
      "votes": null
    },
    {
      "id": "496978",
      "postDate": "03/22/2019 19:50:34",
      "content": "<p>I think that <a href=\"/pietromarinelli\">@pietromarinelli</a> and <a href=\"/deveaup\">@deveaup</a> could answer you better than I can. In the team I focused on testing dnn architectures and fixing the pipelines, but it has been Pietro and Paul who did the data processing, finding new ideas developing features.</p>",
      "rawMarkdown": "I think that @pietromarinelli and @deveaup could answer you better than I can. In the team I focused on testing dnn architectures and fixing the pipelines, but it has been Pietro and Paul who did the data processing, finding new ideas developing features.",
      "votes": null
    },
    {
      "id": "497112",
      "postDate": "03/23/2019 01:00:36",
      "content": "<p>Alright, thanks Luca.</p>",
      "rawMarkdown": "Alright, thanks Luca.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496430,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 07:32:23",
      "content": "<p>Thanks <a href=\"/lucamassaron\">@lucamassaron</a> for sharing. Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features ?)</p>\n\n<p>More question, how did you implement your CV strategy, did it correlate OK with public/private ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496688,
          "author_name": "lucamassaron",
          "author_url": "",
          "post_date": "03/22/2019 13:23:50",
          "content": "<p>We actually noticed little correlation between our submissions and the LB, so we couldn't make the right decisions at the end of the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496806,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 15:53:01",
          "content": "<p>Thanks Luca <a href=\"/lucamassaron\">@lucamassaron</a>!  Would you mind answering this question?\n&gt; Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features = 2904 x 160 x 57 ?  )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496978,
          "author_name": "lucamassaron",
          "author_url": "",
          "post_date": "03/22/2019 19:50:34",
          "content": "<p>I think that <a href=\"/pietromarinelli\">@pietromarinelli</a> and <a href=\"/deveaup\">@deveaup</a> could answer you better than I can. In the team I focused on testing dnn architectures and fixing the pipelines, but it has been Pietro and Paul who did the data processing, finding new ideas developing features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497112,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/23/2019 01:00:36",
          "content": "<p>Alright, thanks Luca.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496461,
      "author_name": "stanislavblinov",
      "author_url": "",
      "post_date": "03/22/2019 08:25:42",
      "content": "<p>Thanks for sharing!\nI tried to play with pooling layers too but in th end i rejected this idea...\nAlso had a 20th place submission but didn't use it as final</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496585,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "03/22/2019 11:13:37",
      "content": "<p>Thanks for sharing Luca.   I'm wondering if there may be some hidden location structure behind the training/public/private splits.   Interesting idea to use a Capsule layer.   Did you pull it from <a href=\"https://github.com/XifengGuo/CapsNet-Keras\">https://github.com/XifengGuo/CapsNet-Keras</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496681,
          "author_name": "lucamassaron",
          "author_url": "",
          "post_date": "03/22/2019 13:16:25",
          "content": "<p>Hi Russ,\nwe exactly used Xifeng Guo's implementation :-) In our opinion, anyway, there is some trouble in the test set. It is problematic that a sub-sample of it provides so different results than the complete set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496603,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/22/2019 11:35:59",
      "content": "<p><a href=\"/lucamassaron\">@lucamassaron</a>, thanks for sharing. Though not as high as yours,  I missed a <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146\">76th place catboost submission</a> myself and there is no way I would have chosen that because both CV and public LB scores were really low.</p>\n\n<p>If you have any insights about a better CV setup for this competition, would you care to <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143\">share it here</a>?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496404": "We still have to make head or tails of this competition. We had a 4th place submission but we couldn't recognize it among the others (which scored equally good), no matter how much we kept trace of the local cv. Moreover it was powered by a dnn which we considered an overfitter given its complexity:\n\n```\ndef model_lstm_A2(input_shape, ex_shape):\n   inp = Input(shape=(input_shape[1], input_shape[2],))\n   ex  = Input(shape=(ex_shape[1],))\n   x = Bidirectional(CuDNNLSTM(128, return_sequences=True))(inp)\n   max_pool1 = GlobalMaxPooling1D()(x)\n   avg_pool1 = GlobalAveragePooling1D()(x)\n   x = Bidirectional(CuDNNLSTM(64, return_sequences=True))(x)\n   attention = Attention(input_shape[1])(x)\n   conv = Conv1D(filters=16, kernel_size=4,\n                 padding='valid', kernel_initializer='he_uniform')(x)\n   avg_pool2 = GlobalAveragePooling1D()(conv)\n   max_pool2 = GlobalMaxPooling1D()(conv)\n   capsule = Capsule(num_capsule=10, dim_capsule=10, routings=4)(x)\n   capsule = Flatten()(capsule)\n   x = Concatenate()([max_pool1, avg_pool1, max_pool2, avg_pool2, attention, capsule, ex])\n   x = BatchNormalization()(x)\n   x = Dense(64, activation=\"relu\")(x)\n   x = Dropout(0.3)(x)\n   x = Dense(1, activation=\"sigmoid\")(x)\n   model = Model(inputs=[inp, ex], outputs=x)\n   model.compile(loss='binary_crossentropy',\n                 optimizer=Adam(lr=LR),\n                 metrics=[matthews_correlation])\n   return model\n```\n\nAnyway we did lear a lot in this competition, really. I even implemented my own attention layer because I didn't trust the one on the kernels :-) (actually the one on the public kernels is ok ;-)).\n\nA big thank you to @pietromarinelli and @deveaup for their great collaborations and endless ideas!",
    "496430": "Thanks @lucamassaron for sharing. Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features ?)\n\nMore question, how did you implement your CV strategy, did it correlate OK with public/private ?",
    "496461": "Thanks for sharing!\nI tried to play with pooling layers too but in th end i rejected this idea...\nAlso had a 20th place submission but didn't use it as final",
    "496585": "Thanks for sharing Luca.   I'm wondering if there may be some hidden location structure behind the training/public/private splits.   Interesting idea to use a Capsule layer.   Did you pull it from https://github.com/XifengGuo/CapsNet-Keras?",
    "496603": "lucamassaron, thanks for sharing. Though not as high as yours,  I missed a [76th place catboost submission](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85146) myself and there is no way I would have chosen that because both CV and public LB scores were really low.\n\nIf you have any insights about a better CV setup for this competition, would you care to [share it here](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85143)?",
    "496681": "Hi Russ,\nwe exactly used Xifeng Guo's implementation :-) In our opinion, anyway, there is some trouble in the test set. It is problematic that a sub-sample of it provides so different results than the complete set.",
    "496688": "We actually noticed little correlation between our submissions and the LB, so we couldn't make the right decisions at the end of the competition.",
    "496806": "Thanks Luca @lucamassaron!  Would you mind answering this question?\n&gt; Regarding this model which gets such a high score, how was your processed input? Did the input process in the same way as the Bruno’s public kernel on 5Folds-LSTM-Attention? (i.e. 160 time steps with mean/std/percentile features = 2904 x 160 x 57 ?  )",
    "496978": "I think that @pietromarinelli and @deveaup could answer you better than I can. In the team I focused on testing dnn architectures and fixing the pipelines, but it has been Pietro and Paul who did the data processing, finding new ideas developing features.",
    "497112": "Alright, thanks Luca."
  },
  "source": "meta"
}