{
  "id": 86960,
  "title": "73rd place solution overview",
  "url": "/competitions/vsb-power-line-fault-detection/writeups/anna-lorenz-73rd-place-solution-overview",
  "author_name": "",
  "post_date": "2019-03-27T21:27:15.047706900Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>Now that this exciting competition is over, it is interesting to read about other people's solution and I though I could also share mine.\nThis is the first Kaggle competition I ever took part in and I'm very excited about my 73rd rank. \nI missed the silver medal by only one place! Maybe I was just lucky in the view of the major shakeup we observed here. Nevertheless, I'd like to share the key aspects of my solution with you. </p>\n\n<p>From what I read in the discussions, many people used DWT denoising and RNN / LSTM architectures to tackle this problem. My appraoch is a bit different. I used a very simple curve-fitting based preprocessing routine and a CNN type of architecture.</p>\n\n<h2>Preprocessing</h2>\n\n<p>I fitted a sine curve with a fixed frequency of 50 Hz (corresponding to the frequency of the underlying electric grid) to each signal. Then I subtracted the fitted sine curve from the original signal. I used the calculated phase shift of the fitted curve to eliminate the phase shift in the original data. I did this, because I read in a paper that the occurrence of the PD pattern was correlated to the phase of the sine wave.</p>\n\n<p>Here is the code to preprocess one signal <code>sig</code>:\n```</p>\n\n<h1>define the parameterized model for the fitted sine curve</h1>\n\n<p>def model(t, alpha, phi):\n    return alpha * np.sin(2 * math.pi * 50 * t + phi)</p>\n\n<p>tdata = np.linspace(0, 0.02, 800000)</p>\n\n<h1>fit the sine curve</h1>\n\n<p>popt, pcov = curve_fit(model, tdata, sig, p0=[30, 0])</p>\n\n<h1>subtract the sine wave from the original signal</h1>\n\n<p>sine = model(tdata, *popt)\nsig_without_sine = sig - sine</p>\n\n<h1>eliminate the phase shift</h1>\n\n<p>phi = popt[1]\nsign = np.sign(popt[0])\nif (sign &lt; 0):\n    shift = int((math.pi + phi) * 800000.0 // (2*math.pi))\nelse :\n    shift = int(phi * 800000.0 // (2*math.pi))</p>\n\n<p>sig_without_sine_and_shifted = np.roll(sig_without_sine, shift)\n```</p>\n\n<h2>Model architecture</h2>\n\n<p>I utilized a neural net with convolutional layers, max pooling and a dense layer as the output layer.\nThe model takes the preprocessed signals as inputs. Some people wrote that I would not be possible to train a conv net accepting this input size, but it worked quite well for me.  </p>\n\n<p><code>\nmodel = Sequential()\nmodel.add(Conv1D(16, kernel_size=11, activation='relu', input_shape=(800000, 1)))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(16, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Flatten())\nmodel.add(Dense(1, activation='sigmoid')) \n</code></p>\n\n<h2>Training</h2>\n\n<p>I used mean squared error as the loss function with a tuned learning rate. I tried binary cross entropy first, but it gave me worse results. I trained the model on two epochs of the data with class weights 20:80 (result of manual parameter tuning).</p>\n\n<p>In fact, I did not train only one model, but 25 of them, on different splits of the training data. I used 80% of the data for training and 20% for validation. None of the 25 models showed clear signs of overfitting, so in the end I kept all of them. </p>\n\n<p><code>\nmodel.compile(loss='mean_squared_error', optimizer=optimizers.SGD(lr=0.005))\nmodel.fit_generator(train_generator, epochs=2, class_weight = {0: 20., 1: 80.})\n</code></p>\n\n<h2>Stacking</h2>\n\n<p>I used a weighted voting scheme to combine the predictions of the 25 models into one final output. I trained the weights on the complete training set using a small NN with basically only one dense layer taking 25 inputs and producing one output. I achieved an MCC of 0.70 on the training set. </p>\n\n<h2>A very simple, yet effective trick</h2>\n\n<p>I observed that in almost all cases, all three signals taken from the same measurement had the same label in the training data, so I assumed that this would also be the case in the test data. Therefore, I applyed a final correction step where I inspected the predictions of three corresponding signals and if one of the signals would get a different label than the other two, I changed that label to match the majority vote. Therefore, all three signals in one measurement would always get the same prediction. This increased my final score on the training set from 0.70 to 0.72. </p>\n\n<h2>Leaderboard scores</h2>\n\n<p>My final model achieved 0.72498 on the training set, 0.52441 on the public leaderboard and 0.65595 on the private leaderboard. </p>\n\n<p>This was not my best scoring model on the public leaderboard.\nMy best scoring model on the public leaderboard was a simpler ensemble made up of only 5 CNNs with a slightly different architecture and slightly different preprocessing routine. This model scored 0.62432 on the private leaderboard (also not too bad). </p>\n\n<p>Nevertheless, I also selected the ensemble of 25, maily because I had put quite some time into tuning and training it ;-) Turned out to be a good decision and got me my first bronze medal.</p>\n\n<p>Please let me know what you think of my approach or if you have any questions. </p>\n\n<p>And thank you all for taking part in this very interesting competition and sharing your thoughts. </p>",
  "messages": [
    {
      "id": "501835",
      "postDate": "03/27/2019 21:27:15",
      "content": "<p>Hi everyone!</p>\n\n<p>Now that this exciting competition is over, it is interesting to read about other people's solution and I though I could also share mine.\nThis is the first Kaggle competition I ever took part in and I'm very excited about my 73rd rank. \nI missed the silver medal by only one place! Maybe I was just lucky in the view of the major shakeup we observed here. Nevertheless, I'd like to share the key aspects of my solution with you. </p>\n\n<p>From what I read in the discussions, many people used DWT denoising and RNN / LSTM architectures to tackle this problem. My appraoch is a bit different. I used a very simple curve-fitting based preprocessing routine and a CNN type of architecture.</p>\n\n<h2>Preprocessing</h2>\n\n<p>I fitted a sine curve with a fixed frequency of 50 Hz (corresponding to the frequency of the underlying electric grid) to each signal. Then I subtracted the fitted sine curve from the original signal. I used the calculated phase shift of the fitted curve to eliminate the phase shift in the original data. I did this, because I read in a paper that the occurrence of the PD pattern was correlated to the phase of the sine wave.</p>\n\n<p>Here is the code to preprocess one signal <code>sig</code>:\n```</p>\n\n<h1>define the parameterized model for the fitted sine curve</h1>\n\n<p>def model(t, alpha, phi):\n    return alpha * np.sin(2 * math.pi * 50 * t + phi)</p>\n\n<p>tdata = np.linspace(0, 0.02, 800000)</p>\n\n<h1>fit the sine curve</h1>\n\n<p>popt, pcov = curve_fit(model, tdata, sig, p0=[30, 0])</p>\n\n<h1>subtract the sine wave from the original signal</h1>\n\n<p>sine = model(tdata, *popt)\nsig_without_sine = sig - sine</p>\n\n<h1>eliminate the phase shift</h1>\n\n<p>phi = popt[1]\nsign = np.sign(popt[0])\nif (sign &lt; 0):\n    shift = int((math.pi + phi) * 800000.0 // (2*math.pi))\nelse :\n    shift = int(phi * 800000.0 // (2*math.pi))</p>\n\n<p>sig_without_sine_and_shifted = np.roll(sig_without_sine, shift)\n```</p>\n\n<h2>Model architecture</h2>\n\n<p>I utilized a neural net with convolutional layers, max pooling and a dense layer as the output layer.\nThe model takes the preprocessed signals as inputs. Some people wrote that I would not be possible to train a conv net accepting this input size, but it worked quite well for me.  </p>\n\n<p><code>\nmodel = Sequential()\nmodel.add(Conv1D(16, kernel_size=11, activation='relu', input_shape=(800000, 1)))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(16, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Flatten())\nmodel.add(Dense(1, activation='sigmoid')) \n</code></p>\n\n<h2>Training</h2>\n\n<p>I used mean squared error as the loss function with a tuned learning rate. I tried binary cross entropy first, but it gave me worse results. I trained the model on two epochs of the data with class weights 20:80 (result of manual parameter tuning).</p>\n\n<p>In fact, I did not train only one model, but 25 of them, on different splits of the training data. I used 80% of the data for training and 20% for validation. None of the 25 models showed clear signs of overfitting, so in the end I kept all of them. </p>\n\n<p><code>\nmodel.compile(loss='mean_squared_error', optimizer=optimizers.SGD(lr=0.005))\nmodel.fit_generator(train_generator, epochs=2, class_weight = {0: 20., 1: 80.})\n</code></p>\n\n<h2>Stacking</h2>\n\n<p>I used a weighted voting scheme to combine the predictions of the 25 models into one final output. I trained the weights on the complete training set using a small NN with basically only one dense layer taking 25 inputs and producing one output. I achieved an MCC of 0.70 on the training set. </p>\n\n<h2>A very simple, yet effective trick</h2>\n\n<p>I observed that in almost all cases, all three signals taken from the same measurement had the same label in the training data, so I assumed that this would also be the case in the test data. Therefore, I applyed a final correction step where I inspected the predictions of three corresponding signals and if one of the signals would get a different label than the other two, I changed that label to match the majority vote. Therefore, all three signals in one measurement would always get the same prediction. This increased my final score on the training set from 0.70 to 0.72. </p>\n\n<h2>Leaderboard scores</h2>\n\n<p>My final model achieved 0.72498 on the training set, 0.52441 on the public leaderboard and 0.65595 on the private leaderboard. </p>\n\n<p>This was not my best scoring model on the public leaderboard.\nMy best scoring model on the public leaderboard was a simpler ensemble made up of only 5 CNNs with a slightly different architecture and slightly different preprocessing routine. This model scored 0.62432 on the private leaderboard (also not too bad). </p>\n\n<p>Nevertheless, I also selected the ensemble of 25, maily because I had put quite some time into tuning and training it ;-) Turned out to be a good decision and got me my first bronze medal.</p>\n\n<p>Please let me know what you think of my approach or if you have any questions. </p>\n\n<p>And thank you all for taking part in this very interesting competition and sharing your thoughts. </p>",
      "rawMarkdown": "Hi everyone!\n\nNow that this exciting competition is over, it is interesting to read about other people's solution and I though I could also share mine.\nThis is the first Kaggle competition I ever took part in and I'm very excited about my 73rd rank. \nI missed the silver medal by only one place! Maybe I was just lucky in the view of the major shakeup we observed here. Nevertheless, I'd like to share the key aspects of my solution with you. \n\nFrom what I read in the discussions, many people used DWT denoising and RNN / LSTM architectures to tackle this problem. My appraoch is a bit different. I used a very simple curve-fitting based preprocessing routine and a CNN type of architecture.\n\n## Preprocessing\nI fitted a sine curve with a fixed frequency of 50 Hz (corresponding to the frequency of the underlying electric grid) to each signal. Then I subtracted the fitted sine curve from the original signal. I used the calculated phase shift of the fitted curve to eliminate the phase shift in the original data. I did this, because I read in a paper that the occurrence of the PD pattern was correlated to the phase of the sine wave.\n\nHere is the code to preprocess one signal `sig`:\n```\n# define the parameterized model for the fitted sine curve\ndef model(t, alpha, phi):\n    return alpha * np.sin(2 * math.pi * 50 * t + phi)\n\ntdata = np.linspace(0, 0.02, 800000)\n\n# fit the sine curve\npopt, pcov = curve_fit(model, tdata, sig, p0=[30, 0])\n\n# subtract the sine wave from the original signal\nsine = model(tdata, *popt)\nsig_without_sine = sig - sine\n\n# eliminate the phase shift\nphi = popt[1]\nsign = np.sign(popt[0])\nif (sign &lt; 0):\n    shift = int((math.pi + phi) * 800000.0 // (2*math.pi))\nelse :\n    shift = int(phi * 800000.0 // (2*math.pi))\n    \nsig_without_sine_and_shifted = np.roll(sig_without_sine, shift)\n```\n## Model architecture\nI utilized a neural net with convolutional layers, max pooling and a dense layer as the output layer.\nThe model takes the preprocessed signals as inputs. Some people wrote that I would not be possible to train a conv net accepting this input size, but it worked quite well for me.  \n\n```\nmodel = Sequential()\nmodel.add(Conv1D(16, kernel_size=11, activation='relu', input_shape=(800000, 1)))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(16, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Flatten())\nmodel.add(Dense(1, activation='sigmoid')) \n```\n\n## Training\nI used mean squared error as the loss function with a tuned learning rate. I tried binary cross entropy first, but it gave me worse results. I trained the model on two epochs of the data with class weights 20:80 (result of manual parameter tuning).\n\nIn fact, I did not train only one model, but 25 of them, on different splits of the training data. I used 80% of the data for training and 20% for validation. None of the 25 models showed clear signs of overfitting, so in the end I kept all of them. \n\n```\nmodel.compile(loss='mean_squared_error', optimizer=optimizers.SGD(lr=0.005))\nmodel.fit_generator(train_generator, epochs=2, class_weight = {0: 20., 1: 80.})\n```\n\n## Stacking\nI used a weighted voting scheme to combine the predictions of the 25 models into one final output. I trained the weights on the complete training set using a small NN with basically only one dense layer taking 25 inputs and producing one output. I achieved an MCC of 0.70 on the training set. \n\n## A very simple, yet effective trick\nI observed that in almost all cases, all three signals taken from the same measurement had the same label in the training data, so I assumed that this would also be the case in the test data. Therefore, I applyed a final correction step where I inspected the predictions of three corresponding signals and if one of the signals would get a different label than the other two, I changed that label to match the majority vote. Therefore, all three signals in one measurement would always get the same prediction. This increased my final score on the training set from 0.70 to 0.72. \n\n## Leaderboard scores\nMy final model achieved 0.72498 on the training set, 0.52441 on the public leaderboard and 0.65595 on the private leaderboard. \n\nThis was not my best scoring model on the public leaderboard.\nMy best scoring model on the public leaderboard was a simpler ensemble made up of only 5 CNNs with a slightly different architecture and slightly different preprocessing routine. This model scored 0.62432 on the private leaderboard (also not too bad). \n\nNevertheless, I also selected the ensemble of 25, maily because I had put quite some time into tuning and training it ;-) Turned out to be a good decision and got me my first bronze medal.\n\nPlease let me know what you think of my approach or if you have any questions. \n\nAnd thank you all for taking part in this very interesting competition and sharing your thoughts.",
      "votes": null
    },
    {
      "id": "513748",
      "postDate": "04/11/2019 08:49:12",
      "content": "<p>Congrats <a href=\"/annalorenz\">@annalorenz</a> and thanks for sharing. The sine curve fitting is a clever approach.</p>",
      "rawMarkdown": "Congrats @annalorenz and thanks for sharing. The sine curve fitting is a clever approach.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 513748,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "04/11/2019 08:49:12",
      "content": "<p>Congrats <a href=\"/annalorenz\">@annalorenz</a> and thanks for sharing. The sine curve fitting is a clever approach.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "501835": "Hi everyone!\n\nNow that this exciting competition is over, it is interesting to read about other people's solution and I though I could also share mine.\nThis is the first Kaggle competition I ever took part in and I'm very excited about my 73rd rank. \nI missed the silver medal by only one place! Maybe I was just lucky in the view of the major shakeup we observed here. Nevertheless, I'd like to share the key aspects of my solution with you. \n\nFrom what I read in the discussions, many people used DWT denoising and RNN / LSTM architectures to tackle this problem. My appraoch is a bit different. I used a very simple curve-fitting based preprocessing routine and a CNN type of architecture.\n\n## Preprocessing\nI fitted a sine curve with a fixed frequency of 50 Hz (corresponding to the frequency of the underlying electric grid) to each signal. Then I subtracted the fitted sine curve from the original signal. I used the calculated phase shift of the fitted curve to eliminate the phase shift in the original data. I did this, because I read in a paper that the occurrence of the PD pattern was correlated to the phase of the sine wave.\n\nHere is the code to preprocess one signal `sig`:\n```\n# define the parameterized model for the fitted sine curve\ndef model(t, alpha, phi):\n    return alpha * np.sin(2 * math.pi * 50 * t + phi)\n\ntdata = np.linspace(0, 0.02, 800000)\n\n# fit the sine curve\npopt, pcov = curve_fit(model, tdata, sig, p0=[30, 0])\n\n# subtract the sine wave from the original signal\nsine = model(tdata, *popt)\nsig_without_sine = sig - sine\n\n# eliminate the phase shift\nphi = popt[1]\nsign = np.sign(popt[0])\nif (sign &lt; 0):\n    shift = int((math.pi + phi) * 800000.0 // (2*math.pi))\nelse :\n    shift = int(phi * 800000.0 // (2*math.pi))\n    \nsig_without_sine_and_shifted = np.roll(sig_without_sine, shift)\n```\n## Model architecture\nI utilized a neural net with convolutional layers, max pooling and a dense layer as the output layer.\nThe model takes the preprocessed signals as inputs. Some people wrote that I would not be possible to train a conv net accepting this input size, but it worked quite well for me.  \n\n```\nmodel = Sequential()\nmodel.add(Conv1D(16, kernel_size=11, activation='relu', input_shape=(800000, 1)))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(16, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(32, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Conv1D(64, kernel_size=11, activation='relu'))\nmodel.add(MaxPooling1D(4))\nmodel.add(Flatten())\nmodel.add(Dense(1, activation='sigmoid')) \n```\n\n## Training\nI used mean squared error as the loss function with a tuned learning rate. I tried binary cross entropy first, but it gave me worse results. I trained the model on two epochs of the data with class weights 20:80 (result of manual parameter tuning).\n\nIn fact, I did not train only one model, but 25 of them, on different splits of the training data. I used 80% of the data for training and 20% for validation. None of the 25 models showed clear signs of overfitting, so in the end I kept all of them. \n\n```\nmodel.compile(loss='mean_squared_error', optimizer=optimizers.SGD(lr=0.005))\nmodel.fit_generator(train_generator, epochs=2, class_weight = {0: 20., 1: 80.})\n```\n\n## Stacking\nI used a weighted voting scheme to combine the predictions of the 25 models into one final output. I trained the weights on the complete training set using a small NN with basically only one dense layer taking 25 inputs and producing one output. I achieved an MCC of 0.70 on the training set. \n\n## A very simple, yet effective trick\nI observed that in almost all cases, all three signals taken from the same measurement had the same label in the training data, so I assumed that this would also be the case in the test data. Therefore, I applyed a final correction step where I inspected the predictions of three corresponding signals and if one of the signals would get a different label than the other two, I changed that label to match the majority vote. Therefore, all three signals in one measurement would always get the same prediction. This increased my final score on the training set from 0.70 to 0.72. \n\n## Leaderboard scores\nMy final model achieved 0.72498 on the training set, 0.52441 on the public leaderboard and 0.65595 on the private leaderboard. \n\nThis was not my best scoring model on the public leaderboard.\nMy best scoring model on the public leaderboard was a simpler ensemble made up of only 5 CNNs with a slightly different architecture and slightly different preprocessing routine. This model scored 0.62432 on the private leaderboard (also not too bad). \n\nNevertheless, I also selected the ensemble of 25, maily because I had put quite some time into tuning and training it ;-) Turned out to be a good decision and got me my first bronze medal.\n\nPlease let me know what you think of my approach or if you have any questions. \n\nAnd thank you all for taking part in this very interesting competition and sharing your thoughts.",
    "513748": "Congrats @annalorenz and thanks for sharing. The sine curve fitting is a clever approach."
  },
  "source": "meta"
}