{
  "id": 93908,
  "title": "No Feature Extraction Needed, Discrete Wavelet Transf. via CNN",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93908",
  "author_name": "",
  "post_date": "2019-05-31T02:23:15.962423500Z",
  "votes": 5,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Apologies for the late posting. </p>\n\n<p>I finished a Kernel for you to look over. Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>This Kernel demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>Enjoy, and good luck!</p>",
  "messages": [
    {
      "id": "540083",
      "postDate": "05/31/2019 02:23:15",
      "content": "<p>Apologies for the late posting. </p>\n\n<p>I finished a Kernel for you to look over. Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.</p>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01</a></p>\n\n<p>This Kernel demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)</p>\n\n<p>This Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.</p>\n\n<p>Enjoy, and good luck!</p>",
      "rawMarkdown": "Apologies for the late posting. \n\nI finished a Kernel for you to look over. Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nThis Kernel demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nEnjoy, and good luck!",
      "votes": null
    },
    {
      "id": "540092",
      "postDate": "05/31/2019 02:37:18",
      "content": "<p>good job, thank you for sharing </p>",
      "rawMarkdown": "good job, thank you for sharing",
      "votes": null
    },
    {
      "id": "540131",
      "postDate": "05/31/2019 04:37:33",
      "content": "<p>Good work. But next time please don't spam your comments in all other topics. One is enough.</p>",
      "rawMarkdown": "Good work. But next time please don't spam your comments in all other topics. One is enough.",
      "votes": null
    },
    {
      "id": "540297",
      "postDate": "05/31/2019 10:08:44",
      "content": "<p>What is k-means cross validation?</p>\n\n<p>Edit: after reading, it is just k folds cv, a typo repeated consistently.  </p>\n\n<p>Thanks for sharing.</p>",
      "rawMarkdown": "What is k-means cross validation?\n\nEdit: after reading, it is just k folds cv, a typo repeated consistently.  \n\nThanks for sharing.",
      "votes": null
    },
    {
      "id": "540337",
      "postDate": "05/31/2019 11:20:39",
      "content": "<p>K-fold, yes you are correct. I guess I was calling it the wrong thing.</p>",
      "rawMarkdown": "K-fold, yes you are correct. I guess I was calling it the wrong thing.",
      "votes": null
    },
    {
      "id": "540338",
      "postDate": "05/31/2019 11:21:06",
      "content": "<p>OK, will do...</p>",
      "rawMarkdown": "OK, will do...",
      "votes": null
    },
    {
      "id": "540339",
      "postDate": "05/31/2019 11:21:54",
      "content": "<p>You’re welcome. I’m still in the learning curve so all advice is appreciated.</p>",
      "rawMarkdown": "You’re welcome. I’m still in the learning curve so all advice is appreciated.",
      "votes": null
    },
    {
      "id": "541162",
      "postDate": "06/01/2019 23:26:25",
      "content": "<p>Interesting idea - still digesting it. </p>\n\n<p>It looks like you get a bunch of different wavelet choices at each approximation step (in the channels of the cnn layer). How does this multitude of wavelets square with the orthonormality property of the discrete wavelet transform?And the similarity across scales?</p>",
      "rawMarkdown": "Interesting idea - still digesting it. \n\nIt looks like you get a bunch of different wavelet choices at each approximation step (in the channels of the cnn layer). How does this multitude of wavelets square with the orthonormality property of the discrete wavelet transform?And the similarity across scales?",
      "votes": null
    },
    {
      "id": "542328",
      "postDate": "06/03/2019 17:55:08",
      "content": "<p>That's a great question. You actually can get not half-bad results by changing \"scale\" to = 1. This is equivalent to changing the \"filters\" value in the Conv1D Keras function to 1. So that means, you're only allowing the CNN to learn one best wavelet for each decomposition layer (similar to techniques used in hand-crated feature extraction discussed in literature). </p>\n\n<p>The reason why yours is a great question is that I can't seem to figure out how to observe these wavelet shapes during training. This might be an artifact of spreading the epochs/work over several CPU/GPUs thanks to TensorFlow.</p>\n\n<p>I do hope to make a new Kernel (probably for a different competition) that allows visualization of the trained wavelet shapes - but that means giving each layer a unique name, and then loading only some of the trained layers into a smaller network whose purpose it is for us to look inside the Robot's mind (the subject of my research  <a href=\"https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/\">https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/</a> ).</p>",
      "rawMarkdown": "That's a great question. You actually can get not half-bad results by changing \"scale\" to = 1. This is equivalent to changing the \"filters\" value in the Conv1D Keras function to 1. So that means, you're only allowing the CNN to learn one best wavelet for each decomposition layer (similar to techniques used in hand-crated feature extraction discussed in literature). \n\nThe reason why yours is a great question is that I can't seem to figure out how to observe these wavelet shapes during training. This might be an artifact of spreading the epochs/work over several CPU/GPUs thanks to TensorFlow.\n\nI do hope to make a new Kernel (probably for a different competition) that allows visualization of the trained wavelet shapes - but that means giving each layer a unique name, and then loading only some of the trained layers into a smaller network whose purpose it is for us to look inside the Robot's mind (the subject of my research  https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/ ).",
      "votes": null
    },
    {
      "id": "542344",
      "postDate": "06/03/2019 18:16:13",
      "content": "<p>I tweaked the Kernel a bit.</p>\n\n<ul>\n<li>Increased the number of trained parameters to approx. 90,000.</li>\n<li>Corrected \"k-fold\" terminology as pointed out by <a href=\"/cpmpml\">@cpmpml</a> .</li>\n<li>Explained the \"Van Halen\" activation function (I guess I'm revealing how old I am).</li>\n<li>Got a marginally better LB score.</li>\n<li>Only posting it here (as per <a href=\"/khahuras\">@khahuras</a> )</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301</a></p>",
      "rawMarkdown": "I tweaked the Kernel a bit.\n\n - Increased the number of trained parameters to approx. 90,000.\n - Corrected \"k-fold\" terminology as pointed out by @cpmpml .\n - Explained the \"Van Halen\" activation function (I guess I'm revealing how old I am).\n - Got a marginally better LB score.\n - Only posting it here (as per @khahuras )\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301",
      "votes": null
    },
    {
      "id": "542442",
      "postDate": "06/03/2019 23:50:57",
      "content": "<p>So, if you would set number of filters to 1, then each conv output would be a scaled version of a “Mother” wavelet? Although I think that this is a very nice way to look at the problem, I am skeptical of the proposition as there seems to be no force in the optimization moving toward orthonormal. Perhaps something in the loss function could do this? Also, maybe we would get some kind of continuous wavelet transform?</p>\n\n<p>p.s. It should be possible to look at the kernel, even in keras. It’s possible in pytorch.</p>",
      "rawMarkdown": "So, if you would set number of filters to 1, then each conv output would be a scaled version of a “Mother” wavelet? Although I think that this is a very nice way to look at the problem, I am skeptical of the proposition as there seems to be no force in the optimization moving toward orthonormal. Perhaps something in the loss function could do this? Also, maybe we would get some kind of continuous wavelet transform?\n\np.s. It should be possible to look at the kernel, even in keras. It’s possible in pytorch.",
      "votes": null
    },
    {
      "id": "543237",
      "postDate": "06/04/2019 12:02:54",
      "content": "<p>Let me rephrase your question to see if I understand it. At each layer n of the DWT, you are asking how we make sure that g(n) is orthonormal with h(n). </p>\n\n<p>In the proposed solution, we do not actually find h(n) (I will look into determining the kernel during training) but instead create a pseudo-residual (pseudo-detail coefficients) that approximates the result of convolving h(n) with the signal. These convolutions (the signal with g(n) and the signal with h(n)) are by design of the \"Van Halen\" activation function, maximum when the other is zero, and vice versa. Convolved with each other, the result should tend towards 0 (orthogonal), although individually, they are not scaled to a magnitude of 1 (normalized). </p>\n\n<p>So g(n) and h(n) are \"pseudo-orthogonal\" but not \"normalized\" in the current Kernel.</p>\n\n<p>Did I understand your question correctly?</p>",
      "rawMarkdown": "Let me rephrase your question to see if I understand it. At each layer n of the DWT, you are asking how we make sure that g(n) is orthonormal with h(n). \n\nIn the proposed solution, we do not actually find h(n) (I will look into determining the kernel during training) but instead create a pseudo-residual (pseudo-detail coefficients) that approximates the result of convolving h(n) with the signal. These convolutions (the signal with g(n) and the signal with h(n)) are by design of the \"Van Halen\" activation function, maximum when the other is zero, and vice versa. Convolved with each other, the result should tend towards 0 (orthogonal), although individually, they are not scaled to a magnitude of 1 (normalized). \n\nSo g(n) and h(n) are \"pseudo-orthogonal\" but not \"normalized\" in the current Kernel.\n\nDid I understand your question correctly?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 540092,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "05/31/2019 02:37:18",
      "content": "<p>good job, thank you for sharing </p>",
      "votes": null,
      "replies": [
        {
          "id": 540339,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/31/2019 11:21:54",
          "content": "<p>You’re welcome. I’m still in the learning curve so all advice is appreciated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540131,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "05/31/2019 04:37:33",
      "content": "<p>Good work. But next time please don't spam your comments in all other topics. One is enough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540338,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/31/2019 11:21:06",
          "content": "<p>OK, will do...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540297,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/31/2019 10:08:44",
      "content": "<p>What is k-means cross validation?</p>\n\n<p>Edit: after reading, it is just k folds cv, a typo repeated consistently.  </p>\n\n<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540337,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "05/31/2019 11:20:39",
          "content": "<p>K-fold, yes you are correct. I guess I was calling it the wrong thing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 541162,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "06/01/2019 23:26:25",
      "content": "<p>Interesting idea - still digesting it. </p>\n\n<p>It looks like you get a bunch of different wavelet choices at each approximation step (in the channels of the cnn layer). How does this multitude of wavelets square with the orthonormality property of the discrete wavelet transform?And the similarity across scales?</p>",
      "votes": null,
      "replies": [
        {
          "id": 542328,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "06/03/2019 17:55:08",
          "content": "<p>That's a great question. You actually can get not half-bad results by changing \"scale\" to = 1. This is equivalent to changing the \"filters\" value in the Conv1D Keras function to 1. So that means, you're only allowing the CNN to learn one best wavelet for each decomposition layer (similar to techniques used in hand-crated feature extraction discussed in literature). </p>\n\n<p>The reason why yours is a great question is that I can't seem to figure out how to observe these wavelet shapes during training. This might be an artifact of spreading the epochs/work over several CPU/GPUs thanks to TensorFlow.</p>\n\n<p>I do hope to make a new Kernel (probably for a different competition) that allows visualization of the trained wavelet shapes - but that means giving each layer a unique name, and then loading only some of the trained layers into a smaller network whose purpose it is for us to look inside the Robot's mind (the subject of my research  <a href=\"https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/\">https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/</a> ).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 542442,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "06/03/2019 23:50:57",
          "content": "<p>So, if you would set number of filters to 1, then each conv output would be a scaled version of a “Mother” wavelet? Although I think that this is a very nice way to look at the problem, I am skeptical of the proposition as there seems to be no force in the optimization moving toward orthonormal. Perhaps something in the loss function could do this? Also, maybe we would get some kind of continuous wavelet transform?</p>\n\n<p>p.s. It should be possible to look at the kernel, even in keras. It’s possible in pytorch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 543237,
          "author_name": "pnussbaum",
          "author_url": "",
          "post_date": "06/04/2019 12:02:54",
          "content": "<p>Let me rephrase your question to see if I understand it. At each layer n of the DWT, you are asking how we make sure that g(n) is orthonormal with h(n). </p>\n\n<p>In the proposed solution, we do not actually find h(n) (I will look into determining the kernel during training) but instead create a pseudo-residual (pseudo-detail coefficients) that approximates the result of convolving h(n) with the signal. These convolutions (the signal with g(n) and the signal with h(n)) are by design of the \"Van Halen\" activation function, maximum when the other is zero, and vice versa. Convolved with each other, the result should tend towards 0 (orthogonal), although individually, they are not scaled to a magnitude of 1 (normalized). </p>\n\n<p>So g(n) and h(n) are \"pseudo-orthogonal\" but not \"normalized\" in the current Kernel.</p>\n\n<p>Did I understand your question correctly?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542344,
      "author_name": "pnussbaum",
      "author_url": "",
      "post_date": "06/03/2019 18:16:13",
      "content": "<p>I tweaked the Kernel a bit.</p>\n\n<ul>\n<li>Increased the number of trained parameters to approx. 90,000.</li>\n<li>Corrected \"k-fold\" terminology as pointed out by <a href=\"/cpmpml\">@cpmpml</a> .</li>\n<li>Explained the \"Van Halen\" activation function (I guess I'm revealing how old I am).</li>\n<li>Got a marginally better LB score.</li>\n<li>Only posting it here (as per <a href=\"/khahuras\">@khahuras</a> )</li>\n</ul>\n\n<p><a href=\"https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301\">https://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "540083": "Apologies for the late posting. \n\nI finished a Kernel for you to look over. Gets a pretty good LB score, with no feature extraction. The CNN does all the feature extraction itself, using Discrete Wavelet Transform. This is accomplished using a novel method of \"pseudo-residual\" to calculate detail coefficients.\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v01\n\nThis Kernel demonstrates a \"zero feature extraction\" method to solve the problem. Keeps trainable parameter count very low (less than 32,000 trainable parameters solve the whole problem)\n\nThis Kernel demonstrates a \"Leave Two Out\" k-means cross-validation scheme to insure accurate results.\n\nEnjoy, and good luck!",
    "540092": "good job, thank you for sharing",
    "540131": "Good work. But next time please don't spam your comments in all other topics. One is enough.",
    "540297": "What is k-means cross validation?\n\nEdit: after reading, it is just k folds cv, a typo repeated consistently.  \n\nThanks for sharing.",
    "540337": "K-fold, yes you are correct. I guess I was calling it the wrong thing.",
    "540338": "OK, will do...",
    "540339": "You’re welcome. I’m still in the learning curve so all advice is appreciated.",
    "541162": "Interesting idea - still digesting it. \n\nIt looks like you get a bunch of different wavelet choices at each approximation step (in the channels of the cnn layer). How does this multitude of wavelets square with the orthonormality property of the discrete wavelet transform?And the similarity across scales?",
    "542328": "That's a great question. You actually can get not half-bad results by changing \"scale\" to = 1. This is equivalent to changing the \"filters\" value in the Conv1D Keras function to 1. So that means, you're only allowing the CNN to learn one best wavelet for each decomposition layer (similar to techniques used in hand-crated feature extraction discussed in literature). \n\nThe reason why yours is a great question is that I can't seem to figure out how to observe these wavelet shapes during training. This might be an artifact of spreading the epochs/work over several CPU/GPUs thanks to TensorFlow.\n\nI do hope to make a new Kernel (probably for a different competition) that allows visualization of the trained wavelet shapes - but that means giving each layer a unique name, and then loading only some of the trained layers into a smaller network whose purpose it is for us to look inside the Robot's mind (the subject of my research  https://www.linkedin.com/pulse/reading-robot-mind-paul-nussbaum-ph-d-/ ).",
    "542344": "I tweaked the Kernel a bit.\n\n - Increased the number of trained parameters to approx. 90,000.\n - Corrected \"k-fold\" terminology as pointed out by @cpmpml .\n - Explained the \"Van Halen\" activation function (I guess I'm revealing how old I am).\n - Got a marginally better LB score.\n - Only posting it here (as per @khahuras )\n\nhttps://www.kaggle.com/pnussbaum/dwt-earthquake-w-lto-v0301",
    "542442": "So, if you would set number of filters to 1, then each conv output would be a scaled version of a “Mother” wavelet? Although I think that this is a very nice way to look at the problem, I am skeptical of the proposition as there seems to be no force in the optimization moving toward orthonormal. Perhaps something in the loss function could do this? Also, maybe we would get some kind of continuous wavelet transform?\n\np.s. It should be possible to look at the kernel, even in keras. It’s possible in pytorch.",
    "543237": "Let me rephrase your question to see if I understand it. At each layer n of the DWT, you are asking how we make sure that g(n) is orthonormal with h(n). \n\nIn the proposed solution, we do not actually find h(n) (I will look into determining the kernel during training) but instead create a pseudo-residual (pseudo-detail coefficients) that approximates the result of convolving h(n) with the signal. These convolutions (the signal with g(n) and the signal with h(n)) are by design of the \"Van Halen\" activation function, maximum when the other is zero, and vice versa. Convolved with each other, the result should tend towards 0 (orthogonal), although individually, they are not scaled to a magnitude of 1 (normalized). \n\nSo g(n) and h(n) are \"pseudo-orthogonal\" but not \"normalized\" in the current Kernel.\n\nDid I understand your question correctly?"
  },
  "source": "meta"
}