{
  "id": 81317,
  "title": "Adversarial Validation example for VSB Power Line Fault Detection",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/81317",
  "author_name": "",
  "post_date": "2019-02-20T17:46:42.064731200Z",
  "votes": 30,
  "comment_count": 15,
  "views": 0,
  "content": "<p>An Adversarial Validation approach may be useful when the TEST set may be very different from the TRAINING set. Simply choosing a subset of the training set to perform validation may not yield the best results. In the case of the VSB Power Line Fault Detection, this indeed seems to be the situation. </p>\n\n<p>The example here includes code to create such an Adversarial Validation Set for this competition, and hopefully you can make good use of it in your software!</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12\">https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12</a> </p>\n\n<p>What Is Adversarial Validation ?</p>\n\n<p>As you likely already know, traditional methods for creation of a validation set include stratified k-fold, stratified percentage split, and a simple percentage split (as included in the \"fit\" method's \"validation_split\" argument), as well as others.</p>\n\n<p>The example Python Jupyter Notebook demonstrates a different kind of validation split, popular in several Kaggle competitions, called the Adversarial Validation approach. In this approach, we create a machine learning algorithm to distinguish between the training set and the testing set. We then use that algorithm to find those training set examples that \"most resemble\" testing set examples, an we use those as our validation set. We then train our recognition algorithm as we normally would.</p>\n\n<p>The example linked to above uses a Peak to Peak (PTP) feature from three phases of non-overlapping windows of data, and a Convolutional Neural Network (CNN) to both create the Adversarial Validation Set as well as to recognize when a fault has taken place in the VSB Power competition.</p>\n\n<p>You can naturally use your favorite feature extractions, and learning algorithm, and the example is written to make that easy for you to make those changes.</p>\n\n<p>Enjoy and good luck!</p>",
  "messages": [
    {
      "id": "475393",
      "postDate": "02/20/2019 17:46:42",
      "content": "<p>An Adversarial Validation approach may be useful when the TEST set may be very different from the TRAINING set. Simply choosing a subset of the training set to perform validation may not yield the best results. In the case of the VSB Power Line Fault Detection, this indeed seems to be the situation. </p>\n\n<p>The example here includes code to create such an Adversarial Validation Set for this competition, and hopefully you can make good use of it in your software!</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12\">https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12</a> </p>\n\n<p>What Is Adversarial Validation ?</p>\n\n<p>As you likely already know, traditional methods for creation of a validation set include stratified k-fold, stratified percentage split, and a simple percentage split (as included in the \"fit\" method's \"validation_split\" argument), as well as others.</p>\n\n<p>The example Python Jupyter Notebook demonstrates a different kind of validation split, popular in several Kaggle competitions, called the Adversarial Validation approach. In this approach, we create a machine learning algorithm to distinguish between the training set and the testing set. We then use that algorithm to find those training set examples that \"most resemble\" testing set examples, an we use those as our validation set. We then train our recognition algorithm as we normally would.</p>\n\n<p>The example linked to above uses a Peak to Peak (PTP) feature from three phases of non-overlapping windows of data, and a Convolutional Neural Network (CNN) to both create the Adversarial Validation Set as well as to recognize when a fault has taken place in the VSB Power competition.</p>\n\n<p>You can naturally use your favorite feature extractions, and learning algorithm, and the example is written to make that easy for you to make those changes.</p>\n\n<p>Enjoy and good luck!</p>",
      "rawMarkdown": "An Adversarial Validation approach may be useful when the TEST set may be very different from the TRAINING set. Simply choosing a subset of the training set to perform validation may not yield the best results. In the case of the VSB Power Line Fault Detection, this indeed seems to be the situation. \n\nThe example here includes code to create such an Adversarial Validation Set for this competition, and hopefully you can make good use of it in your software!\n\nhttps://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12 \n\nWhat Is Adversarial Validation ?\n\nAs you likely already know, traditional methods for creation of a validation set include stratified k-fold, stratified percentage split, and a simple percentage split (as included in the \"fit\" method's \"validation_split\" argument), as well as others.\n\nThe example Python Jupyter Notebook demonstrates a different kind of validation split, popular in several Kaggle competitions, called the Adversarial Validation approach. In this approach, we create a machine learning algorithm to distinguish between the training set and the testing set. We then use that algorithm to find those training set examples that \"most resemble\" testing set examples, an we use those as our validation set. We then train our recognition algorithm as we normally would.\n\nThe example linked to above uses a Peak to Peak (PTP) feature from three phases of non-overlapping windows of data, and a Convolutional Neural Network (CNN) to both create the Adversarial Validation Set as well as to recognize when a fault has taken place in the VSB Power competition.\n\nYou can naturally use your favorite feature extractions, and learning algorithm, and the example is written to make that easy for you to make those changes.\n\nEnjoy and good luck!",
      "votes": null
    },
    {
      "id": "475863",
      "postDate": "02/21/2019 09:54:17",
      "content": "<p>Thank you. I like all your Kaggle kernels for this competition.</p>",
      "rawMarkdown": "Thank you. I like all your Kaggle kernels for this competition.",
      "votes": null
    },
    {
      "id": "476029",
      "postDate": "02/21/2019 14:15:18",
      "content": "<p>Thanks Subrahmanyam - I'm trying to be a good \"Kernel Contributor\" so it pleases me to know when folks find them useful. Feel free to up-vote the kernel / discussion, and more importantly - have fun!</p>",
      "rawMarkdown": "Thanks Subrahmanyam - I'm trying to be a good \"Kernel Contributor\" so it pleases me to know when folks find them useful. Feel free to up-vote the kernel / discussion, and more importantly - have fun!",
      "votes": null
    },
    {
      "id": "476037",
      "postDate": "02/21/2019 14:29:28",
      "content": "<p>Thanks for the useful kernel. </p>\n\n<p>BTW, why is the approach called \"adversarial validation\"? What is adversarial in it? :)</p>",
      "rawMarkdown": "Thanks for the useful kernel. \n\nBTW, why is the approach called \"adversarial validation\"? What is adversarial in it? :)",
      "votes": null
    },
    {
      "id": "476104",
      "postDate": "02/21/2019 16:04:57",
      "content": "<p>If I'm not mistaken, it comes from the \"Generative Adversarial Network (GAN)\" technique popular in creating realistic video game textures and other AI created art. In that technique, there are two networks that are trained simultaneously - one \"generative\" to create the artwork and one \"discriminative\" to decide if it looks realistic or not. They work like enemies, in an adversarial fashion, one trying to fool the other. </p>\n\n<p>Similarly, in the case of the Adversarial Validation kernel example, we are trying to create a validation set that is the most realistic - or the most similar to the test data set.</p>",
      "rawMarkdown": "If I'm not mistaken, it comes from the \"Generative Adversarial Network (GAN)\" technique popular in creating realistic video game textures and other AI created art. In that technique, there are two networks that are trained simultaneously - one \"generative\" to create the artwork and one \"discriminative\" to decide if it looks realistic or not. They work like enemies, in an adversarial fashion, one trying to fool the other. \n\nSimilarly, in the case of the Adversarial Validation kernel example, we are trying to create a validation set that is the most realistic - or the most similar to the test data set.",
      "votes": null
    },
    {
      "id": "476363",
      "postDate": "02/22/2019 03:07:05",
      "content": "<p>With such a \"adversarial validation\", do we risk over-fitting the validation set by only using examples that \"most resemble\" the test set, compared with the traditional k-fold validation?</p>\n\n<p>Besides, I think the adversarial validation set is just like a hold-out validation set, though this set is chosen elaborately.</p>",
      "rawMarkdown": "With such a \"adversarial validation\", do we risk over-fitting the validation set by only using examples that \"most resemble\" the test set, compared with the traditional k-fold validation?\n\nBesides, I think the adversarial validation set is just like a hold-out validation set, though this set is chosen elaborately.",
      "votes": null
    },
    {
      "id": "477096",
      "postDate": "02/23/2019 21:23:52",
      "content": "<p>Definitely! This is only useful when we feel the training data set has some significant differences from the testing data set. </p>\n\n<p>Summary: (for those not as familiar with the techniques) With k-fold, generally, we sequentially hold out a different portion of the training set for validation and use the rest for training. Each time in the sequence (or \"fold\"), we use the trained network to evaluate the test set, and hopefully the average result is better than if we only did that segmentation once. </p>\n\n<p>With Adversarial Validation, we do not feel the average of the k-folds will be better than the hand-picked fold. We justify this hypothesis based on a big difference between the training and testing sets that we observe. </p>\n\n<p>It's really a judgement call, and then check the validity of your judgement with the LB numbers.</p>\n\n<p>Next steps: Speaking of \"chosen elaborately\" - I think I need to make an improvement to the kernel I posted (any recommendations are welcome) <a href=\"https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12\">https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12</a> . The Adversarial selected Validation set is not stratified currently. What I mean to say is that the percentage of \"true\" examples in the validation set is not necessarily equal to the percentage of \"true\" examples in the entire training set - skewing a Bayesian training algorithm. I need to find a way to fix that...</p>",
      "rawMarkdown": "Definitely! This is only useful when we feel the training data set has some significant differences from the testing data set. \n\nSummary: (for those not as familiar with the techniques) With k-fold, generally, we sequentially hold out a different portion of the training set for validation and use the rest for training. Each time in the sequence (or \"fold\"), we use the trained network to evaluate the test set, and hopefully the average result is better than if we only did that segmentation once. \n\nWith Adversarial Validation, we do not feel the average of the k-folds will be better than the hand-picked fold. We justify this hypothesis based on a big difference between the training and testing sets that we observe. \n\nIt's really a judgement call, and then check the validity of your judgement with the LB numbers.\n\nNext steps: Speaking of \"chosen elaborately\" - I think I need to make an improvement to the kernel I posted (any recommendations are welcome) https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12 . The Adversarial selected Validation set is not stratified currently. What I mean to say is that the percentage of \"true\" examples in the validation set is not necessarily equal to the percentage of \"true\" examples in the entire training set - skewing a Bayesian training algorithm. I need to find a way to fix that...",
      "votes": null
    },
    {
      "id": "477106",
      "postDate": "02/23/2019 22:41:27",
      "content": "<p>So far, adversarial validation helped me for this competition. </p>",
      "rawMarkdown": "So far, adversarial validation helped me for this competition.",
      "votes": null
    },
    {
      "id": "478843",
      "postDate": "02/26/2019 17:39:29",
      "content": "<p>That’s great! It didn’t help me so far. May be i am doing something wrong. Will give it another shot.</p>",
      "rawMarkdown": "That’s great! It didn’t help me so far. May be i am doing something wrong. Will give it another shot.",
      "votes": null
    },
    {
      "id": "478884",
      "postDate": "02/26/2019 18:38:55",
      "content": "<p>I used it to create a better stratification for my CV. For instance, cut adversarial probabilities (for train) in 5 buckets and mix with target (0/1) and you have a way to balance data per target and test dataset features similarities. It's not magic but it gives better stability for me.</p>\n\n<p>It looks there is something better to do when I see top LB ... but there is no much sharing about what works and what does not. Currently I'm trying multi-labels as measurement target, did you try it?</p>",
      "rawMarkdown": "I used it to create a better stratification for my CV. For instance, cut adversarial probabilities (for train) in 5 buckets and mix with target (0/1) and you have a way to balance data per target and test dataset features similarities. It's not magic but it gives better stability for me.\n\nIt looks there is something better to do when I see top LB ... but there is no much sharing about what works and what does not. Currently I'm trying multi-labels as measurement target, did you try it?",
      "votes": null
    },
    {
      "id": "478889",
      "postDate": "02/26/2019 18:46:13",
      "content": "<p>Thanks! I will try that. \nI just started working on multi-label work as mentioned by Max <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py</a>. \nI am redoing it again. Building features again and getting some help from here : <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a>. \nUp to now, i was using the randomness of LSTM model on denoised data plus blending to get the lb score of .761. </p>",
      "rawMarkdown": "Thanks! I will try that. \nI just started working on multi-label work as mentioned by Max https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py. \nI am redoing it again. Building features again and getting some help from here : https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples. \nUp to now, i was using the randomness of LSTM model on denoised data plus blending to get the lb score of .761.",
      "votes": null
    },
    {
      "id": "479947",
      "postDate": "02/27/2019 15:51:59",
      "content": "<p><a href=\"/harshit92\">@harshit92</a> I'm surprized denoising the signal works in your case. I tried the filtering based on Fourrier Transform to filter out frequencies higher than 1e7 . And also the high pass filter + wavelet denoising with 15 Level coef to leave most of the noise in the signal, but both gave me worst results.</p>",
      "rawMarkdown": "harshit92 I'm surprized denoising the signal works in your case. I tried the filtering based on Fourrier Transform to filter out frequencies higher than 1e7 . And also the high pass filter + wavelet denoising with 15 Level coef to leave most of the noise in the signal, but both gave me worst results.",
      "votes": null
    },
    {
      "id": "480036",
      "postDate": "02/27/2019 17:53:44",
      "content": "<p>Hi Antoine, For denoising i mostly used the paper mentioned here <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236</a> and this kernel <a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">https://www.kaggle.com/jackvial/dwt-signal-denoising</a>.</p>\n\n<p>Just by varying the seeds, denoised model (BI_LSTM +attention) lb varies from .644 to .724 and LB of model without denoising varies from .653 to .707. My model is highly unstable. I am switching back to decision tree, hopefully that will help me out. </p>",
      "rawMarkdown": "Hi Antoine, For denoising i mostly used the paper mentioned here https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236 and this kernel https://www.kaggle.com/jackvial/dwt-signal-denoising.\n\nJust by varying the seeds, denoised model (BI_LSTM +attention) lb varies from .644 to .724 and LB of model without denoising varies from .653 to .707. My model is highly unstable. I am switching back to decision tree, hopefully that will help me out.",
      "votes": null
    },
    {
      "id": "480179",
      "postDate": "02/27/2019 21:43:03",
      "content": "<p>@HarshitMehta and which target are you using? Is it per measurement with target=1 if any of the three phase has a fault? Or all? Or at least 2?</p>\n\n<p>BTW: I'm testing multi-labels and I don't think the MCC keras code available in kernels can work because it's only for binary classification. For multi-labels, I've wrapped the scikit learn one with:</p>\n\n<pre><code># Scikit Learn wrapper for MCC multi-labels\ndef mcc_np(y_true, y_pred):\n  y_true = np.argmax(y_true, axis=1) # if using np_utils.to_categorical on y_train\n  y_pred = np.argmax(y_pred, axis=1) # if using np_utils.to_categorical on y_train\n  m = matthews_corrcoef(y_true, y_pred)\n  return m\n\ndef mcc_tf(y_true, y_pred):\n    m = tf.py_func(mcc_np, [y_true, y_pred], tf.double)\n    return m\n</code></pre>\n\n<p>And so far (with 2 models), only 2 classes (all zeros or all ones) are predicted on test dataset. And only 3 classes on CV (around 0.70). All other classes have too low probabilities.</p>\n\n<p><strong>Update</strong>: Tried snapshots ensemble (x18) with multi-labels target (8 classes) with BiLSTM/Attention and got CV 0.70 and LB 0.60. Again, what is interesting is that only 000 or 111 are predicted even in CV. Model is not able to distinguish individual phase failure (001, 010, 100) or both (110, 011, 101).</p>",
      "rawMarkdown": "HarshitMehta and which target are you using? Is it per measurement with target=1 if any of the three phase has a fault? Or all? Or at least 2?\n\nBTW: I'm testing multi-labels and I don't think the MCC keras code available in kernels can work because it's only for binary classification. For multi-labels, I've wrapped the scikit learn one with:\n\n\n    # Scikit Learn wrapper for MCC multi-labels\n    def mcc_np(y_true, y_pred):\n      y_true = np.argmax(y_true, axis=1) # if using np_utils.to_categorical on y_train\n      y_pred = np.argmax(y_pred, axis=1) # if using np_utils.to_categorical on y_train\n      m = matthews_corrcoef(y_true, y_pred)\n      return m\n    \n    def mcc_tf(y_true, y_pred):\n        m = tf.py_func(mcc_np, [y_true, y_pred], tf.double)\n        return m\n\n\nAnd so far (with 2 models), only 2 classes (all zeros or all ones) are predicted on test dataset. And only 3 classes on CV (around 0.70). All other classes have too low probabilities.\n\n**Update**: Tried snapshots ensemble (x18) with multi-labels target (8 classes) with BiLSTM/Attention and got CV 0.70 and LB 0.60. Again, what is interesting is that only 000 or 111 are predicted even in CV. Model is not able to distinguish individual phase failure (001, 010, 100) or both (110, 011, 101).",
      "votes": null
    },
    {
      "id": "480204",
      "postDate": "02/27/2019 22:58:46",
      "content": "<p>Upto now, I treated target as 1, if any phase for a given id_measurement has fault, else 0. Now, i am switching that to condition when all the phases has fault for a given id_measurement. Thank you for bringing this up, i didn't consider this so far. I will work on this. Are you getting stable CV with fixed seed ?\nAssuming, you are using LSTM stuff.\nPS: I haven't tested multiclass so far, i am creating features and that script is still running for test dataset, after that i can try that. I will first try decision tree for MultiClass Stuff, then may move to LSTM.</p>",
      "rawMarkdown": "Upto now, I treated target as 1, if any phase for a given id_measurement has fault, else 0. Now, i am switching that to condition when all the phases has fault for a given id_measurement. Thank you for bringing this up, i didn't consider this so far. I will work on this. Are you getting stable CV with fixed seed ?\nAssuming, you are using LSTM stuff.\nPS: I haven't tested multiclass so far, i am creating features and that script is still running for test dataset, after that i can try that. I will first try decision tree for MultiClass Stuff, then may move to LSTM.",
      "votes": null
    },
    {
      "id": "484333",
      "postDate": "03/05/2019 20:59:01",
      "content": "<p>I get better results with target = 1 when idmeasurement has 2 or 3 faults. I gave up on multi-labels approach.</p>",
      "rawMarkdown": "I get better results with target = 1 when idmeasurement has 2 or 3 faults. I gave up on multi-labels approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 475863,
      "author_name": "subrahmanyamv",
      "author_url": "",
      "post_date": "02/21/2019 09:54:17",
      "content": "<p>Thank you. I like all your Kaggle kernels for this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 476029,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "02/21/2019 14:15:18",
          "content": "<p>Thanks Subrahmanyam - I'm trying to be a good \"Kernel Contributor\" so it pleases me to know when folks find them useful. Feel free to up-vote the kernel / discussion, and more importantly - have fun!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 476037,
      "author_name": "donkeys",
      "author_url": "",
      "post_date": "02/21/2019 14:29:28",
      "content": "<p>Thanks for the useful kernel. </p>\n\n<p>BTW, why is the approach called \"adversarial validation\"? What is adversarial in it? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 476104,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "02/21/2019 16:04:57",
          "content": "<p>If I'm not mistaken, it comes from the \"Generative Adversarial Network (GAN)\" technique popular in creating realistic video game textures and other AI created art. In that technique, there are two networks that are trained simultaneously - one \"generative\" to create the artwork and one \"discriminative\" to decide if it looks realistic or not. They work like enemies, in an adversarial fashion, one trying to fool the other. </p>\n\n<p>Similarly, in the case of the Adversarial Validation kernel example, we are trying to create a validation set that is the most realistic - or the most similar to the test data set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 476363,
      "author_name": "shuhuagao",
      "author_url": "",
      "post_date": "02/22/2019 03:07:05",
      "content": "<p>With such a \"adversarial validation\", do we risk over-fitting the validation set by only using examples that \"most resemble\" the test set, compared with the traditional k-fold validation?</p>\n\n<p>Besides, I think the adversarial validation set is just like a hold-out validation set, though this set is chosen elaborately.</p>",
      "votes": null,
      "replies": [
        {
          "id": 477096,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "02/23/2019 21:23:52",
          "content": "<p>Definitely! This is only useful when we feel the training data set has some significant differences from the testing data set. </p>\n\n<p>Summary: (for those not as familiar with the techniques) With k-fold, generally, we sequentially hold out a different portion of the training set for validation and use the rest for training. Each time in the sequence (or \"fold\"), we use the trained network to evaluate the test set, and hopefully the average result is better than if we only did that segmentation once. </p>\n\n<p>With Adversarial Validation, we do not feel the average of the k-folds will be better than the hand-picked fold. We justify this hypothesis based on a big difference between the training and testing sets that we observe. </p>\n\n<p>It's really a judgement call, and then check the validity of your judgement with the LB numbers.</p>\n\n<p>Next steps: Speaking of \"chosen elaborately\" - I think I need to make an improvement to the kernel I posted (any recommendations are welcome) <a href=\"https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12\">https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12</a> . The Adversarial selected Validation set is not stratified currently. What I mean to say is that the percentage of \"true\" examples in the validation set is not necessarily equal to the percentage of \"true\" examples in the entire training set - skewing a Bayesian training algorithm. I need to find a way to fix that...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477106,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/23/2019 22:41:27",
          "content": "<p>So far, adversarial validation helped me for this competition. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 478843,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/26/2019 17:39:29",
          "content": "<p>That’s great! It didn’t help me so far. May be i am doing something wrong. Will give it another shot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 478884,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/26/2019 18:38:55",
          "content": "<p>I used it to create a better stratification for my CV. For instance, cut adversarial probabilities (for train) in 5 buckets and mix with target (0/1) and you have a way to balance data per target and test dataset features similarities. It's not magic but it gives better stability for me.</p>\n\n<p>It looks there is something better to do when I see top LB ... but there is no much sharing about what works and what does not. Currently I'm trying multi-labels as measurement target, did you try it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 478889,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/26/2019 18:46:13",
          "content": "<p>Thanks! I will try that. \nI just started working on multi-label work as mentioned by Max <a href=\"https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py\">https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py</a>. \nI am redoing it again. Building features again and getting some help from here : <a href=\"https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples\">https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples</a>. \nUp to now, i was using the randomness of LSTM model on denoised data plus blending to get the lb score of .761. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 479947,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "02/27/2019 15:51:59",
          "content": "<p><a href=\"/harshit92\">@harshit92</a> I'm surprized denoising the signal works in your case. I tried the filtering based on Fourrier Transform to filter out frequencies higher than 1e7 . And also the high pass filter + wavelet denoising with 15 Level coef to leave most of the noise in the signal, but both gave me worst results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480036,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/27/2019 17:53:44",
          "content": "<p>Hi Antoine, For denoising i mostly used the paper mentioned here <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236</a> and this kernel <a href=\"https://www.kaggle.com/jackvial/dwt-signal-denoising\">https://www.kaggle.com/jackvial/dwt-signal-denoising</a>.</p>\n\n<p>Just by varying the seeds, denoised model (BI_LSTM +attention) lb varies from .644 to .724 and LB of model without denoising varies from .653 to .707. My model is highly unstable. I am switching back to decision tree, hopefully that will help me out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480179,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/27/2019 21:43:03",
          "content": "<p>@HarshitMehta and which target are you using? Is it per measurement with target=1 if any of the three phase has a fault? Or all? Or at least 2?</p>\n\n<p>BTW: I'm testing multi-labels and I don't think the MCC keras code available in kernels can work because it's only for binary classification. For multi-labels, I've wrapped the scikit learn one with:</p>\n\n<pre><code># Scikit Learn wrapper for MCC multi-labels\ndef mcc_np(y_true, y_pred):\n  y_true = np.argmax(y_true, axis=1) # if using np_utils.to_categorical on y_train\n  y_pred = np.argmax(y_pred, axis=1) # if using np_utils.to_categorical on y_train\n  m = matthews_corrcoef(y_true, y_pred)\n  return m\n\ndef mcc_tf(y_true, y_pred):\n    m = tf.py_func(mcc_np, [y_true, y_pred], tf.double)\n    return m\n</code></pre>\n\n<p>And so far (with 2 models), only 2 classes (all zeros or all ones) are predicted on test dataset. And only 3 classes on CV (around 0.70). All other classes have too low probabilities.</p>\n\n<p><strong>Update</strong>: Tried snapshots ensemble (x18) with multi-labels target (8 classes) with BiLSTM/Attention and got CV 0.70 and LB 0.60. Again, what is interesting is that only 000 or 111 are predicted even in CV. Model is not able to distinguish individual phase failure (001, 010, 100) or both (110, 011, 101).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 480204,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/27/2019 22:58:46",
          "content": "<p>Upto now, I treated target as 1, if any phase for a given id_measurement has fault, else 0. Now, i am switching that to condition when all the phases has fault for a given id_measurement. Thank you for bringing this up, i didn't consider this so far. I will work on this. Are you getting stable CV with fixed seed ?\nAssuming, you are using LSTM stuff.\nPS: I haven't tested multiclass so far, i am creating features and that script is still running for test dataset, after that i can try that. I will first try decision tree for MultiClass Stuff, then may move to LSTM.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 484333,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/05/2019 20:59:01",
          "content": "<p>I get better results with target = 1 when idmeasurement has 2 or 3 faults. I gave up on multi-labels approach.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "475393": "An Adversarial Validation approach may be useful when the TEST set may be very different from the TRAINING set. Simply choosing a subset of the training set to perform validation may not yield the best results. In the case of the VSB Power Line Fault Detection, this indeed seems to be the situation. \n\nThe example here includes code to create such an Adversarial Validation Set for this competition, and hopefully you can make good use of it in your software!\n\nhttps://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12 \n\nWhat Is Adversarial Validation ?\n\nAs you likely already know, traditional methods for creation of a validation set include stratified k-fold, stratified percentage split, and a simple percentage split (as included in the \"fit\" method's \"validation_split\" argument), as well as others.\n\nThe example Python Jupyter Notebook demonstrates a different kind of validation split, popular in several Kaggle competitions, called the Adversarial Validation approach. In this approach, we create a machine learning algorithm to distinguish between the training set and the testing set. We then use that algorithm to find those training set examples that \"most resemble\" testing set examples, an we use those as our validation set. We then train our recognition algorithm as we normally would.\n\nThe example linked to above uses a Peak to Peak (PTP) feature from three phases of non-overlapping windows of data, and a Convolutional Neural Network (CNN) to both create the Adversarial Validation Set as well as to recognize when a fault has taken place in the VSB Power competition.\n\nYou can naturally use your favorite feature extractions, and learning algorithm, and the example is written to make that easy for you to make those changes.\n\nEnjoy and good luck!",
    "475863": "Thank you. I like all your Kaggle kernels for this competition.",
    "476029": "Thanks Subrahmanyam - I'm trying to be a good \"Kernel Contributor\" so it pleases me to know when folks find them useful. Feel free to up-vote the kernel / discussion, and more importantly - have fun!",
    "476037": "Thanks for the useful kernel. \n\nBTW, why is the approach called \"adversarial validation\"? What is adversarial in it? :)",
    "476104": "If I'm not mistaken, it comes from the \"Generative Adversarial Network (GAN)\" technique popular in creating realistic video game textures and other AI created art. In that technique, there are two networks that are trained simultaneously - one \"generative\" to create the artwork and one \"discriminative\" to decide if it looks realistic or not. They work like enemies, in an adversarial fashion, one trying to fool the other. \n\nSimilarly, in the case of the Adversarial Validation kernel example, we are trying to create a validation set that is the most realistic - or the most similar to the test data set.",
    "476363": "With such a \"adversarial validation\", do we risk over-fitting the validation set by only using examples that \"most resemble\" the test set, compared with the traditional k-fold validation?\n\nBesides, I think the adversarial validation set is just like a hold-out validation set, though this set is chosen elaborately.",
    "477096": "Definitely! This is only useful when we feel the training data set has some significant differences from the testing data set. \n\nSummary: (for those not as familiar with the techniques) With k-fold, generally, we sequentially hold out a different portion of the training set for validation and use the rest for training. Each time in the sequence (or \"fold\"), we use the trained network to evaluate the test set, and hopefully the average result is better than if we only did that segmentation once. \n\nWith Adversarial Validation, we do not feel the average of the k-folds will be better than the hand-picked fold. We justify this hypothesis based on a big difference between the training and testing sets that we observe. \n\nIt's really a judgement call, and then check the validity of your judgement with the LB numbers.\n\nNext steps: Speaking of \"chosen elaborately\" - I think I need to make an improvement to the kernel I posted (any recommendations are welcome) https://www.kaggle.com/pnussbaum/adversarial-cnn-of-ptp-for-vsb-power-v12 . The Adversarial selected Validation set is not stratified currently. What I mean to say is that the percentage of \"true\" examples in the validation set is not necessarily equal to the percentage of \"true\" examples in the entire training set - skewing a Bayesian training algorithm. I need to find a way to fix that...",
    "477106": "So far, adversarial validation helped me for this competition.",
    "478843": "That’s great! It didn’t help me so far. May be i am doing something wrong. Will give it another shot.",
    "478884": "I used it to create a better stratification for my CV. For instance, cut adversarial probabilities (for train) in 5 buckets and mix with target (0/1) and you have a way to balance data per target and test dataset features similarities. It's not magic but it gives better stability for me.\n\nIt looks there is something better to do when I see top LB ... but there is no much sharing about what works and what does not. Currently I'm trying multi-labels as measurement target, did you try it?",
    "478889": "Thanks! I will try that. \nI just started working on multi-label work as mentioned by Max https://github.com/MaxHalford/kaggle-vsb-power/blob/master/scripts/extract_solo_features.py. \nI am redoing it again. Building features again and getting some help from here : https://www.kaggle.com/artgor/earthquakes-fe-more-features-and-samples. \nUp to now, i was using the randomness of LSTM model on denoised data plus blending to get the lb score of .761.",
    "479947": "harshit92 I'm surprized denoising the signal works in your case. I tried the filtering based on Fourrier Transform to filter out frequencies higher than 1e7 . And also the high pass filter + wavelet denoising with 15 Level coef to leave most of the noise in the signal, but both gave me worst results.",
    "480036": "Hi Antoine, For denoising i mostly used the paper mentioned here https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771#447236 and this kernel https://www.kaggle.com/jackvial/dwt-signal-denoising.\n\nJust by varying the seeds, denoised model (BI_LSTM +attention) lb varies from .644 to .724 and LB of model without denoising varies from .653 to .707. My model is highly unstable. I am switching back to decision tree, hopefully that will help me out.",
    "480179": "HarshitMehta and which target are you using? Is it per measurement with target=1 if any of the three phase has a fault? Or all? Or at least 2?\n\nBTW: I'm testing multi-labels and I don't think the MCC keras code available in kernels can work because it's only for binary classification. For multi-labels, I've wrapped the scikit learn one with:\n\n\n    # Scikit Learn wrapper for MCC multi-labels\n    def mcc_np(y_true, y_pred):\n      y_true = np.argmax(y_true, axis=1) # if using np_utils.to_categorical on y_train\n      y_pred = np.argmax(y_pred, axis=1) # if using np_utils.to_categorical on y_train\n      m = matthews_corrcoef(y_true, y_pred)\n      return m\n    \n    def mcc_tf(y_true, y_pred):\n        m = tf.py_func(mcc_np, [y_true, y_pred], tf.double)\n        return m\n\n\nAnd so far (with 2 models), only 2 classes (all zeros or all ones) are predicted on test dataset. And only 3 classes on CV (around 0.70). All other classes have too low probabilities.\n\n**Update**: Tried snapshots ensemble (x18) with multi-labels target (8 classes) with BiLSTM/Attention and got CV 0.70 and LB 0.60. Again, what is interesting is that only 000 or 111 are predicted even in CV. Model is not able to distinguish individual phase failure (001, 010, 100) or both (110, 011, 101).",
    "480204": "Upto now, I treated target as 1, if any phase for a given id_measurement has fault, else 0. Now, i am switching that to condition when all the phases has fault for a given id_measurement. Thank you for bringing this up, i didn't consider this so far. I will work on this. Are you getting stable CV with fixed seed ?\nAssuming, you are using LSTM stuff.\nPS: I haven't tested multiclass so far, i am creating features and that script is still running for test dataset, after that i can try that. I will first try decision tree for MultiClass Stuff, then may move to LSTM.",
    "484333": "I get better results with target = 1 when idmeasurement has 2 or 3 faults. I gave up on multi-labels approach."
  },
  "source": "meta"
}