{
  "id": 84982,
  "title": "Mind Sharing your observations after competition ended?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/84982",
  "author_name": "",
  "post_date": "2019-03-21T00:31:15.554251Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Do you mind sharing unique observations you made over this dataset once the competition is finished? </p>\n\n<p>For me here are couple of the observations:</p>\n\n<p>1- I received validation better results on phase 0 signals comparing to phase 1 and phase 2. </p>\n\n<p>2- Based on some previous research, <a href=\"https://www.sciencedirect.com/science/article/pii/S1474667017666345/pdf?md5=e2cc84b71912792d89b492f8ba272013&amp;pid=1-s2.0-S1474667017666345-main.pdf&amp;_valck=1\">park vector</a> seemed to be a good option. Although it could have generated results per measurement as it combines the signals. In my trials, I realized replacing signal features with park vector would not be very helpful in generating better outcome. </p>\n\n<p>3- After extracting a vast set of features, it seemed simpler models with carefully selected simpler features worked better. </p>",
  "messages": [
    {
      "id": "495312",
      "postDate": "03/21/2019 00:31:15",
      "content": "<p>Do you mind sharing unique observations you made over this dataset once the competition is finished? </p>\n\n<p>For me here are couple of the observations:</p>\n\n<p>1- I received validation better results on phase 0 signals comparing to phase 1 and phase 2. </p>\n\n<p>2- Based on some previous research, <a href=\"https://www.sciencedirect.com/science/article/pii/S1474667017666345/pdf?md5=e2cc84b71912792d89b492f8ba272013&amp;pid=1-s2.0-S1474667017666345-main.pdf&amp;_valck=1\">park vector</a> seemed to be a good option. Although it could have generated results per measurement as it combines the signals. In my trials, I realized replacing signal features with park vector would not be very helpful in generating better outcome. </p>\n\n<p>3- After extracting a vast set of features, it seemed simpler models with carefully selected simpler features worked better. </p>",
      "rawMarkdown": "Do you mind sharing unique observations you made over this dataset once the competition is finished? \n\nFor me here are couple of the observations:\n\n1- I received validation better results on phase 0 signals comparing to phase 1 and phase 2. \n\n2- Based on some previous research, [park vector](https://www.sciencedirect.com/science/article/pii/S1474667017666345/pdf?md5=e2cc84b71912792d89b492f8ba272013&amp;pid=1-s2.0-S1474667017666345-main.pdf&amp;_valck=1) seemed to be a good option. Although it could have generated results per measurement as it combines the signals. In my trials, I realized replacing signal features with park vector would not be very helpful in generating better outcome. \n\n3- After extracting a vast set of features, it seemed simpler models with carefully selected simpler features worked better.",
      "votes": null
    },
    {
      "id": "495406",
      "postDate": "03/21/2019 03:40:45",
      "content": "<p>Thanks $B^2$ for initiate this topic. I observed one phenomenon from my experiments that may somehow counter to point 3 above : </p>\n\n<p>Based on the framework of Bruno's kernel\n<a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\">https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694</a></p>\n\n<p>I try to reduce the number of time step (<code>n_dim</code>) from 160 to 16 or even 4. And also simplify X's features from 57 to 27. So the resulted data is much more simpler than the kernel. And by using much simpler version of the same 2-layers BI-LSTM model, but reduce the number of hidden node from 128/64 to 32/16, this simpler model can still fit the training data almost perfectly (have to tune a bit on LR). </p>\n\n<p>Nevertheless, one counter-intuitive evidence is that this simpler model with simpler data overfits much more than the complex model, i.e. <code>validation_mcc</code> is much worse even though <code>train_mcc</code> gets almost perfect score.</p>\n\n<p>On this last day, I try to reverse the direction; increase complexity of both data &amp; model, and am waiting to see the result.</p>",
      "rawMarkdown": "Thanks $B^2$ for initiate this topic. I observed one phenomenon from my experiments that may somehow counter to point 3 above : \n\nBased on the framework of Bruno's kernel\nhttps://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\n\nI try to reduce the number of time step (`n_dim`) from 160 to 16 or even 4. And also simplify X's features from 57 to 27. So the resulted data is much more simpler than the kernel. And by using much simpler version of the same 2-layers BI-LSTM model, but reduce the number of hidden node from 128/64 to 32/16, this simpler model can still fit the training data almost perfectly (have to tune a bit on LR). \n\nNevertheless, one counter-intuitive evidence is that this simpler model with simpler data overfits much more than the complex model, i.e. `validation_mcc` is much worse even though `train_mcc` gets almost perfect score.\n\nOn this last day, I try to reverse the direction; increase complexity of both data &amp; model, and am waiting to see the result.",
      "votes": null
    },
    {
      "id": "495686",
      "postDate": "03/21/2019 12:13:09",
      "content": "<p>I tried more or less the same thing as you, decreased complexity, and overfitting went up! I am still at a loss how that may be..</p>",
      "rawMarkdown": "I tried more or less the same thing as you, decreased complexity, and overfitting went up! I am still at a loss how that may be..",
      "votes": null
    },
    {
      "id": "496203",
      "postDate": "03/22/2019 01:12:11",
      "content": "<ul>\n<li><p>I found out that some measurements are out of order. E.g., one would expect relative angles of phases 0, 1, 2 to follow the (0deg, 120 deg, -120 deg) pattern, but sometimes they appear in opposite direction (0deg, -120 deg, 120 deg). </p></li>\n<li><p>I did no feature extraction at all. I.e, fed pure 3-phase signals \"as is\" (except for phase re-ordering, which is relevant for my architecture). No LSTM :-)</p></li>\n<li><p>This approach yielded .629 on public LB and up to .667 on private LB, (.64.. .68 local CV). I guess with more patience for traning and better hyperparameter tuning it is capable to reach .7 mark.</p></li>\n</ul>",
      "rawMarkdown": "I found out that some measurements are out of order. E.g., one would expect relative angles of phases 0, 1, 2 to follow the (0deg, 120 deg, -120 deg) pattern, but sometimes they appear in opposite direction (0deg, -120 deg, 120 deg). \n\n- I did no feature extraction at all. I.e, fed pure 3-phase signals \"as is\" (except for phase re-ordering, which is relevant for my architecture). No LSTM :-)\n\n- This approach yielded .629 on public LB and up to .667 on private LB, (.64.. .68 local CV). I guess with more patience for traning and better hyperparameter tuning it is capable to reach .7 mark.",
      "votes": null
    },
    {
      "id": "496204",
      "postDate": "03/22/2019 01:14:00",
      "content": "<p>Hi Joop, <a href=\"/proef2\">@proef2</a> , it looks like you did very well! congratulation!</p>",
      "rawMarkdown": "Hi Joop, @proef2 , it looks like you did very well! congratulation!",
      "votes": null
    },
    {
      "id": "497075",
      "postDate": "03/22/2019 23:23:42",
      "content": "<p>Congratulations <a href=\"/proef2\">@proef2</a> :) great achievement. </p>",
      "rawMarkdown": "Congratulations @proef2 :) great achievement.",
      "votes": null
    },
    {
      "id": "497078",
      "postDate": "03/22/2019 23:25:24",
      "content": "<p>I didn't pay attention to the degrees but I noticed classifying phase 0 was easier for the models comparing to phase 1 and 2. If the out of order issue was only for phase 1 and 2, this could potentially describe the reason. </p>",
      "rawMarkdown": "I didn't pay attention to the degrees but I noticed classifying phase 0 was easier for the models comparing to phase 1 and 2. If the out of order issue was only for phase 1 and 2, this could potentially describe the reason.",
      "votes": null
    },
    {
      "id": "497079",
      "postDate": "03/22/2019 23:27:08",
      "content": "<p>Interesting enough, the Park Vectors that I mentioned was able to generate very good results in Private leader board. Unfortunately, I didn't select the submission that was purely by Park Vector for final results (as the public score was low), but it could have achieved a silver medal. </p>",
      "rawMarkdown": "Interesting enough, the Park Vectors that I mentioned was able to generate very good results in Private leader board. Unfortunately, I didn't select the submission that was purely by Park Vector for final results (as the public score was low), but it could have achieved a silver medal.",
      "votes": null
    },
    {
      "id": "497081",
      "postDate": "03/22/2019 23:29:05",
      "content": "<p>Turned out one of my versions of LSTM was the best, I need to look into the specific version of the kernel to remember what features and what setting I was using for that. I'll lookup and post here. </p>",
      "rawMarkdown": "Turned out one of my versions of LSTM was the best, I need to look into the specific version of the kernel to remember what features and what setting I was using for that. I'll lookup and post here.",
      "votes": null
    },
    {
      "id": "497291",
      "postDate": "03/23/2019 09:36:54",
      "content": "<p>Thank you! That was a surprise :-)  . My submission was from three weeks ago that got a .66 or so on the public LB. Just a single model simple four layer CNN, with I believe three or four simple features for 800 parts.\nMy friend did much better than me with some handcrafted smart features on a random forest with standard settings from scikit-learn :-)</p>",
      "rawMarkdown": "Thank you! That was a surprise :-)  . My submission was from three weeks ago that got a .66 or so on the public LB. Just a single model simple four layer CNN, with I believe three or four simple features for 800 parts.\nMy friend did much better than me with some handcrafted smart features on a random forest with standard settings from scikit-learn :-)",
      "votes": null
    },
    {
      "id": "497375",
      "postDate": "03/23/2019 13:12:12",
      "content": "<p><a href=\"/proef2\">@proef2</a> did you denoise/detrend the data?</p>",
      "rawMarkdown": "proef2 did you denoise/detrend the data?",
      "votes": null
    },
    {
      "id": "497503",
      "postDate": "03/23/2019 16:17:23",
      "content": "<p>Yes, I detrended for convenience for some features by taking the absolute first difference of the raw data like so:</p>\n\n<p><code>segment_diff = np.absolute(np.diff(segment))</code>\n<code>max = np.log1p(segment_diff.max())</code></p>\n\n<p>This was the feature that seemed to help the most in the local crossval.</p>",
      "rawMarkdown": "Yes, I detrended for convenience for some features by taking the absolute first difference of the raw data like so:\n\n`segment_diff = np.absolute(np.diff(segment))`\n`max = np.log1p(segment_diff.max())`\n\nThis was the feature that seemed to help the most in the local crossval.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 495406,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/21/2019 03:40:45",
      "content": "<p>Thanks $B^2$ for initiate this topic. I observed one phenomenon from my experiments that may somehow counter to point 3 above : </p>\n\n<p>Based on the framework of Bruno's kernel\n<a href=\"https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\">https://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694</a></p>\n\n<p>I try to reduce the number of time step (<code>n_dim</code>) from 160 to 16 or even 4. And also simplify X's features from 57 to 27. So the resulted data is much more simpler than the kernel. And by using much simpler version of the same 2-layers BI-LSTM model, but reduce the number of hidden node from 128/64 to 32/16, this simpler model can still fit the training data almost perfectly (have to tune a bit on LR). </p>\n\n<p>Nevertheless, one counter-intuitive evidence is that this simpler model with simpler data overfits much more than the complex model, i.e. <code>validation_mcc</code> is much worse even though <code>train_mcc</code> gets almost perfect score.</p>\n\n<p>On this last day, I try to reverse the direction; increase complexity of both data &amp; model, and am waiting to see the result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 495686,
          "author_name": "proef2",
          "author_url": "",
          "post_date": "03/21/2019 12:13:09",
          "content": "<p>I tried more or less the same thing as you, decreased complexity, and overfitting went up! I am still at a loss how that may be..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496204,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 01:14:00",
          "content": "<p>Hi Joop, <a href=\"/proef2\">@proef2</a> , it looks like you did very well! congratulation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497075,
          "author_name": "behrad3d",
          "author_url": "",
          "post_date": "03/22/2019 23:23:42",
          "content": "<p>Congratulations <a href=\"/proef2\">@proef2</a> :) great achievement. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497291,
          "author_name": "proef2",
          "author_url": "",
          "post_date": "03/23/2019 09:36:54",
          "content": "<p>Thank you! That was a surprise :-)  . My submission was from three weeks ago that got a .66 or so on the public LB. Just a single model simple four layer CNN, with I believe three or four simple features for 800 parts.\nMy friend did much better than me with some handcrafted smart features on a random forest with standard settings from scikit-learn :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497375,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/23/2019 13:12:12",
          "content": "<p><a href=\"/proef2\">@proef2</a> did you denoise/detrend the data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497503,
          "author_name": "proef2",
          "author_url": "",
          "post_date": "03/23/2019 16:17:23",
          "content": "<p>Yes, I detrended for convenience for some features by taking the absolute first difference of the raw data like so:</p>\n\n<p><code>segment_diff = np.absolute(np.diff(segment))</code>\n<code>max = np.log1p(segment_diff.max())</code></p>\n\n<p>This was the feature that seemed to help the most in the local crossval.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496203,
      "author_name": "m0nzderr",
      "author_url": "",
      "post_date": "03/22/2019 01:12:11",
      "content": "<ul>\n<li><p>I found out that some measurements are out of order. E.g., one would expect relative angles of phases 0, 1, 2 to follow the (0deg, 120 deg, -120 deg) pattern, but sometimes they appear in opposite direction (0deg, -120 deg, 120 deg). </p></li>\n<li><p>I did no feature extraction at all. I.e, fed pure 3-phase signals \"as is\" (except for phase re-ordering, which is relevant for my architecture). No LSTM :-)</p></li>\n<li><p>This approach yielded .629 on public LB and up to .667 on private LB, (.64.. .68 local CV). I guess with more patience for traning and better hyperparameter tuning it is capable to reach .7 mark.</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 497078,
          "author_name": "behrad3d",
          "author_url": "",
          "post_date": "03/22/2019 23:25:24",
          "content": "<p>I didn't pay attention to the degrees but I noticed classifying phase 0 was easier for the models comparing to phase 1 and 2. If the out of order issue was only for phase 1 and 2, this could potentially describe the reason. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 497081,
          "author_name": "behrad3d",
          "author_url": "",
          "post_date": "03/22/2019 23:29:05",
          "content": "<p>Turned out one of my versions of LSTM was the best, I need to look into the specific version of the kernel to remember what features and what setting I was using for that. I'll lookup and post here. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 497079,
      "author_name": "behrad3d",
      "author_url": "",
      "post_date": "03/22/2019 23:27:08",
      "content": "<p>Interesting enough, the Park Vectors that I mentioned was able to generate very good results in Private leader board. Unfortunately, I didn't select the submission that was purely by Park Vector for final results (as the public score was low), but it could have achieved a silver medal. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "495312": "Do you mind sharing unique observations you made over this dataset once the competition is finished? \n\nFor me here are couple of the observations:\n\n1- I received validation better results on phase 0 signals comparing to phase 1 and phase 2. \n\n2- Based on some previous research, [park vector](https://www.sciencedirect.com/science/article/pii/S1474667017666345/pdf?md5=e2cc84b71912792d89b492f8ba272013&amp;pid=1-s2.0-S1474667017666345-main.pdf&amp;_valck=1) seemed to be a good option. Although it could have generated results per measurement as it combines the signals. In my trials, I realized replacing signal features with park vector would not be very helpful in generating better outcome. \n\n3- After extracting a vast set of features, it seemed simpler models with carefully selected simpler features worked better.",
    "495406": "Thanks $B^2$ for initiate this topic. I observed one phenomenon from my experiments that may somehow counter to point 3 above : \n\nBased on the framework of Bruno's kernel\nhttps://www.kaggle.com/braquino/5-fold-lstm-attention-fully-commented-0-694\n\nI try to reduce the number of time step (`n_dim`) from 160 to 16 or even 4. And also simplify X's features from 57 to 27. So the resulted data is much more simpler than the kernel. And by using much simpler version of the same 2-layers BI-LSTM model, but reduce the number of hidden node from 128/64 to 32/16, this simpler model can still fit the training data almost perfectly (have to tune a bit on LR). \n\nNevertheless, one counter-intuitive evidence is that this simpler model with simpler data overfits much more than the complex model, i.e. `validation_mcc` is much worse even though `train_mcc` gets almost perfect score.\n\nOn this last day, I try to reverse the direction; increase complexity of both data &amp; model, and am waiting to see the result.",
    "495686": "I tried more or less the same thing as you, decreased complexity, and overfitting went up! I am still at a loss how that may be..",
    "496203": "I found out that some measurements are out of order. E.g., one would expect relative angles of phases 0, 1, 2 to follow the (0deg, 120 deg, -120 deg) pattern, but sometimes they appear in opposite direction (0deg, -120 deg, 120 deg). \n\n- I did no feature extraction at all. I.e, fed pure 3-phase signals \"as is\" (except for phase re-ordering, which is relevant for my architecture). No LSTM :-)\n\n- This approach yielded .629 on public LB and up to .667 on private LB, (.64.. .68 local CV). I guess with more patience for traning and better hyperparameter tuning it is capable to reach .7 mark.",
    "496204": "Hi Joop, @proef2 , it looks like you did very well! congratulation!",
    "497075": "Congratulations @proef2 :) great achievement.",
    "497078": "I didn't pay attention to the degrees but I noticed classifying phase 0 was easier for the models comparing to phase 1 and 2. If the out of order issue was only for phase 1 and 2, this could potentially describe the reason.",
    "497079": "Interesting enough, the Park Vectors that I mentioned was able to generate very good results in Private leader board. Unfortunately, I didn't select the submission that was purely by Park Vector for final results (as the public score was low), but it could have achieved a silver medal.",
    "497081": "Turned out one of my versions of LSTM was the best, I need to look into the specific version of the kernel to remember what features and what setting I was using for that. I'll lookup and post here.",
    "497291": "Thank you! That was a surprise :-)  . My submission was from three weeks ago that got a .66 or so on the public LB. Just a single model simple four layer CNN, with I believe three or four simple features for 800 parts.\nMy friend did much better than me with some handcrafted smart features on a random forest with standard settings from scikit-learn :-)",
    "497375": "proef2 did you denoise/detrend the data?",
    "497503": "Yes, I detrended for convenience for some features by taking the absolute first difference of the raw data like so:\n\n`segment_diff = np.absolute(np.diff(segment))`\n`max = np.log1p(segment_diff.max())`\n\nThis was the feature that seemed to help the most in the local crossval."
  },
  "source": "meta"
}