{
  "id": 93451,
  "title": "Issues with my CNN ",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/93451",
  "author_name": "",
  "post_date": "2019-05-27T07:38:44.349296500Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello guys</p>\n\n<p>I am really really sorry for asking silly questions. I am just a beginner and trying to improve my skill in data science. So please, bear with me.</p>\n\n<p>Yesterday when I just created a very small toy CNN with only one 1 conv layer, the network was learning and the predicted values was following the ground truth. But, then I was studying another kernel and tried to expanded my CNN model. It improved the result.</p>\n\n<p>But then I do not who what I did, at some point the network stopped learning. Now, it predicts a single value for all training set. I tried to go back to my previous small network, but it is always giving a single predicted value. What ever I do, change the learning rate, change the activation, change the layers, the network won't learn anything.</p>\n\n<p>I would really appreciate any suggestion regarding the possible causes for these behaviors. Thank you so much.</p>",
  "messages": [
    {
      "id": "537533",
      "postDate": "05/27/2019 07:38:44",
      "content": "<p>Hello guys</p>\n\n<p>I am really really sorry for asking silly questions. I am just a beginner and trying to improve my skill in data science. So please, bear with me.</p>\n\n<p>Yesterday when I just created a very small toy CNN with only one 1 conv layer, the network was learning and the predicted values was following the ground truth. But, then I was studying another kernel and tried to expanded my CNN model. It improved the result.</p>\n\n<p>But then I do not who what I did, at some point the network stopped learning. Now, it predicts a single value for all training set. I tried to go back to my previous small network, but it is always giving a single predicted value. What ever I do, change the learning rate, change the activation, change the layers, the network won't learn anything.</p>\n\n<p>I would really appreciate any suggestion regarding the possible causes for these behaviors. Thank you so much.</p>",
      "rawMarkdown": "Hello guys\n\nI am really really sorry for asking silly questions. I am just a beginner and trying to improve my skill in data science. So please, bear with me.\n\nYesterday when I just created a very small toy CNN with only one 1 conv layer, the network was learning and the predicted values was following the ground truth. But, then I was studying another kernel and tried to expanded my CNN model. It improved the result.\n\nBut then I do not who what I did, at some point the network stopped learning. Now, it predicts a single value for all training set. I tried to go back to my previous small network, but it is always giving a single predicted value. What ever I do, change the learning rate, change the activation, change the layers, the network won't learn anything.\n\nI would really appreciate any suggestion regarding the possible causes for these behaviors. Thank you so much.",
      "votes": null
    },
    {
      "id": "537630",
      "postDate": "05/27/2019 11:17:01",
      "content": "<p>First, you are showing only 4 epochs. The network might be stuck due to vanishing gradient. Cannot tell without more epochs. You can try decreasing learning rate, if that does not help, remove layers by layer and see which layer is causing issues. I recommend removing layer by layer because you might have added a layer of two that is causing your network to do nothing.</p>",
      "rawMarkdown": "First, you are showing only 4 epochs. The network might be stuck due to vanishing gradient. Cannot tell without more epochs. You can try decreasing learning rate, if that does not help, remove layers by layer and see which layer is causing issues. I recommend removing layer by layer because you might have added a layer of two that is causing your network to do nothing.",
      "votes": null
    },
    {
      "id": "537851",
      "postDate": "05/27/2019 18:09:28",
      "content": "<p>I ran for 40 epochs just now, the loss is fixed at 5.7324 or something like that..</p>",
      "rawMarkdown": "I ran for 40 epochs just now, the loss is fixed at 5.7324 or something like that..",
      "votes": null
    },
    {
      "id": "537855",
      "postDate": "05/27/2019 18:29:04",
      "content": "<p>Drop everything but the first CNN block, then try again. You need to learn to debug... you added too much layers. Go layer by layer to figure out where it is going wrong.</p>",
      "rawMarkdown": "Drop everything but the first CNN block, then try again. You need to learn to debug... you added too much layers. Go layer by layer to figure out where it is going wrong.",
      "votes": null
    },
    {
      "id": "537914",
      "postDate": "05/27/2019 20:45:05",
      "content": "<p>Take a look at <a href=\"https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2\">https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2</a></p>\n\n<p>By the way, (off topic question) why do you do permutation followed by flattening? Does it makes sense?</p>",
      "rawMarkdown": "Take a look at https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2\n\nBy the way, (off topic question) why do you do permutation followed by flattening? Does it makes sense?",
      "votes": null
    },
    {
      "id": "537944",
      "postDate": "05/27/2019 22:24:53",
      "content": "<p>I ran into similar difficulties trying to train a CNN.  In addition to the suggestions made by other posters, you might try varying the random number seed before creating and training your model.  I found that with some seed values training got stuck with high loss (perhaps the initial weight settings made it susceptible to the vanishing gradient problem), but with other seeds it was more successful.</p>",
      "rawMarkdown": "I ran into similar difficulties trying to train a CNN.  In addition to the suggestions made by other posters, you might try varying the random number seed before creating and training your model.  I found that with some seed values training got stuck with high loss (perhaps the initial weight settings made it susceptible to the vanishing gradient problem), but with other seeds it was more successful.",
      "votes": null
    },
    {
      "id": "538246",
      "postDate": "05/28/2019 10:24:54",
      "content": "<p>Thank you so much Tim! Apparently softmax activation was the problem. But that is weird. The kernel I was following had softmax and the author says it had ~1.5 in the LB.</p>",
      "rawMarkdown": "Thank you so much Tim! Apparently softmax activation was the problem. But that is weird. The kernel I was following had softmax and the author says it had ~1.5 in the LB.",
      "votes": null
    },
    {
      "id": "538251",
      "postDate": "05/28/2019 10:28:04",
      "content": "<p>Hi Alex! Thank  you for this excellent link. I think I got the bug(?!). Not sure why that is a bug though.\nAnd I think I applied flattening following by permute layer.</p>",
      "rawMarkdown": "Hi Alex! Thank  you for this excellent link. I think I got the bug(?!). Not sure why that is a bug though.\nAnd I think I applied flattening following by permute layer.",
      "votes": null
    },
    {
      "id": "538254",
      "postDate": "05/28/2019 10:31:57",
      "content": "<p>I actually thought in a similar way. But for some reason softmax/sigmoid in the 2nd last FC layer was the issue. Using Relu does the work nicely.</p>",
      "rawMarkdown": "I actually thought in a similar way. But for some reason softmax/sigmoid in the 2nd last FC layer was the issue. Using Relu does the work nicely.",
      "votes": null
    },
    {
      "id": "538580",
      "postDate": "05/28/2019 19:26:47",
      "content": "<p>Read it somewhere,  people don't normally apply Dropout at conv layers.  </p>",
      "rawMarkdown": "Read it somewhere,  people don't normally apply Dropout at conv layers.",
      "votes": null
    },
    {
      "id": "538690",
      "postDate": "05/29/2019 01:44:55",
      "content": "<p>I get the same problem in pytorch, it seems. One thing I noticed in the kernel you refer to is that the author specifically initialized the last Dense (I think it was done to slightly bias toward higher ttf, which isn't easily predicted). Did you have those initializers included?</p>",
      "rawMarkdown": "I get the same problem in pytorch, it seems. One thing I noticed in the kernel you refer to is that the author specifically initialized the last Dense (I think it was done to slightly bias toward higher ttf, which isn't easily predicted). Did you have those initializers included?",
      "votes": null
    },
    {
      "id": "538776",
      "postDate": "05/29/2019 05:23:59",
      "content": "<p>I was just looking into that initialization. It makes a lot of sense now, how important that last custom initialization is. \nAt first I just thought, custom initialization, big deal. I was and still am so naive!</p>",
      "rawMarkdown": "I was just looking into that initialization. It makes a lot of sense now, how important that last custom initialization is. \nAt first I just thought, custom initialization, big deal. I was and still am so naive!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 537630,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "05/27/2019 11:17:01",
      "content": "<p>First, you are showing only 4 epochs. The network might be stuck due to vanishing gradient. Cannot tell without more epochs. You can try decreasing learning rate, if that does not help, remove layers by layer and see which layer is causing issues. I recommend removing layer by layer because you might have added a layer of two that is causing your network to do nothing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 537851,
          "author_name": "rumi2752",
          "author_url": "",
          "post_date": "05/27/2019 18:09:28",
          "content": "<p>I ran for 40 epochs just now, the loss is fixed at 5.7324 or something like that..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 537855,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "05/27/2019 18:29:04",
          "content": "<p>Drop everything but the first CNN block, then try again. You need to learn to debug... you added too much layers. Go layer by layer to figure out where it is going wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538246,
          "author_name": "rumi2752",
          "author_url": "",
          "post_date": "05/28/2019 10:24:54",
          "content": "<p>Thank you so much Tim! Apparently softmax activation was the problem. But that is weird. The kernel I was following had softmax and the author says it had ~1.5 in the LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 537914,
      "author_name": "alexpcli",
      "author_url": "",
      "post_date": "05/27/2019 20:45:05",
      "content": "<p>Take a look at <a href=\"https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2\">https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2</a></p>\n\n<p>By the way, (off topic question) why do you do permutation followed by flattening? Does it makes sense?</p>",
      "votes": null,
      "replies": [
        {
          "id": 538251,
          "author_name": "rumi2752",
          "author_url": "",
          "post_date": "05/28/2019 10:28:04",
          "content": "<p>Hi Alex! Thank  you for this excellent link. I think I got the bug(?!). Not sure why that is a bug though.\nAnd I think I applied flattening following by permute layer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 537944,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "05/27/2019 22:24:53",
      "content": "<p>I ran into similar difficulties trying to train a CNN.  In addition to the suggestions made by other posters, you might try varying the random number seed before creating and training your model.  I found that with some seed values training got stuck with high loss (perhaps the initial weight settings made it susceptible to the vanishing gradient problem), but with other seeds it was more successful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 538254,
          "author_name": "rumi2752",
          "author_url": "",
          "post_date": "05/28/2019 10:31:57",
          "content": "<p>I actually thought in a similar way. But for some reason softmax/sigmoid in the 2nd last FC layer was the issue. Using Relu does the work nicely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538690,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "05/29/2019 01:44:55",
          "content": "<p>I get the same problem in pytorch, it seems. One thing I noticed in the kernel you refer to is that the author specifically initialized the last Dense (I think it was done to slightly bias toward higher ttf, which isn't easily predicted). Did you have those initializers included?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538776,
          "author_name": "rumi2752",
          "author_url": "",
          "post_date": "05/29/2019 05:23:59",
          "content": "<p>I was just looking into that initialization. It makes a lot of sense now, how important that last custom initialization is. \nAt first I just thought, custom initialization, big deal. I was and still am so naive!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 538580,
      "author_name": "joejeo1",
      "author_url": "",
      "post_date": "05/28/2019 19:26:47",
      "content": "<p>Read it somewhere,  people don't normally apply Dropout at conv layers.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "537533": "Hello guys\n\nI am really really sorry for asking silly questions. I am just a beginner and trying to improve my skill in data science. So please, bear with me.\n\nYesterday when I just created a very small toy CNN with only one 1 conv layer, the network was learning and the predicted values was following the ground truth. But, then I was studying another kernel and tried to expanded my CNN model. It improved the result.\n\nBut then I do not who what I did, at some point the network stopped learning. Now, it predicts a single value for all training set. I tried to go back to my previous small network, but it is always giving a single predicted value. What ever I do, change the learning rate, change the activation, change the layers, the network won't learn anything.\n\nI would really appreciate any suggestion regarding the possible causes for these behaviors. Thank you so much.",
    "537630": "First, you are showing only 4 epochs. The network might be stuck due to vanishing gradient. Cannot tell without more epochs. You can try decreasing learning rate, if that does not help, remove layers by layer and see which layer is causing issues. I recommend removing layer by layer because you might have added a layer of two that is causing your network to do nothing.",
    "537851": "I ran for 40 epochs just now, the loss is fixed at 5.7324 or something like that..",
    "537855": "Drop everything but the first CNN block, then try again. You need to learn to debug... you added too much layers. Go layer by layer to figure out where it is going wrong.",
    "537914": "Take a look at https://medium.com/machine-learning-world/how-to-debug-neural-networks-manual-dc2a200f10f2\n\nBy the way, (off topic question) why do you do permutation followed by flattening? Does it makes sense?",
    "537944": "I ran into similar difficulties trying to train a CNN.  In addition to the suggestions made by other posters, you might try varying the random number seed before creating and training your model.  I found that with some seed values training got stuck with high loss (perhaps the initial weight settings made it susceptible to the vanishing gradient problem), but with other seeds it was more successful.",
    "538246": "Thank you so much Tim! Apparently softmax activation was the problem. But that is weird. The kernel I was following had softmax and the author says it had ~1.5 in the LB.",
    "538251": "Hi Alex! Thank  you for this excellent link. I think I got the bug(?!). Not sure why that is a bug though.\nAnd I think I applied flattening following by permute layer.",
    "538254": "I actually thought in a similar way. But for some reason softmax/sigmoid in the 2nd last FC layer was the issue. Using Relu does the work nicely.",
    "538580": "Read it somewhere,  people don't normally apply Dropout at conv layers.",
    "538690": "I get the same problem in pytorch, it seems. One thing I noticed in the kernel you refer to is that the author specifically initialized the last Dense (I think it was done to slightly bias toward higher ttf, which isn't easily predicted). Did you have those initializers included?",
    "538776": "I was just looking into that initialization. It makes a lot of sense now, how important that last custom initialization is. \nAt first I just thought, custom initialization, big deal. I was and still am so naive!"
  },
  "source": "meta"
}