{
  "id": 86453,
  "title": "How to enhance generalization of RNNs ?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/86453",
  "author_name": "",
  "post_date": "2019-03-24T05:31:19.958968500Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Does anyone have any ideas of how to prevent RNN models from overfitting to the small number of training samples ? Here are some ideas :\n- Dropout\n- Reduce model complexity\n- EarlyStopping or ModelCheckpoint callbacks</p>\n\n<p>Please feel free to share any ideas you have to reduce overfitting of RNNs.</p>",
  "messages": [
    {
      "id": "498957",
      "postDate": "03/24/2019 05:31:19",
      "content": "<p>Does anyone have any ideas of how to prevent RNN models from overfitting to the small number of training samples ? Here are some ideas :\n- Dropout\n- Reduce model complexity\n- EarlyStopping or ModelCheckpoint callbacks</p>\n\n<p>Please feel free to share any ideas you have to reduce overfitting of RNNs.</p>",
      "rawMarkdown": "Does anyone have any ideas of how to prevent RNN models from overfitting to the small number of training samples ? Here are some ideas :\n- Dropout\n- Reduce model complexity\n- EarlyStopping or ModelCheckpoint callbacks\n\nPlease feel free to share any ideas you have to reduce overfitting of RNNs.",
      "votes": null
    },
    {
      "id": "504278",
      "postDate": "03/31/2019 10:07:32",
      "content": "<p>Greetings Tarun\nYou're placed much higher in this competition than I am, so perhaps I should be asking you for tips :)</p>\n\n<p>I've been using this <a href=\"https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks#diagnostics\">cheat sheet on machine learning</a> to work with overfitting and underfitting.</p>\n\n<p>For overfitting, the tips there are to get more data, or to <a href=\"https://machinelearningmastery.com/introduction-to-regularization-to-reduce-overfitting-and-improve-generalization-error/\">perform regularisation</a>. You've mentioned dropout, reducing model complexity and callbacks. Of course callbacks and dropout are examples of regularization, but there are also others to take into consideration. I've had some successful improvements of my RNN by utilising dropouts and callbacks as well. </p>",
      "rawMarkdown": "Greetings Tarun\nYou're placed much higher in this competition than I am, so perhaps I should be asking you for tips :)\n\nI've been using this [cheat sheet on machine learning](https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks#diagnostics) to work with overfitting and underfitting.\n\nFor overfitting, the tips there are to get more data, or to [perform regularisation](https://machinelearningmastery.com/introduction-to-regularization-to-reduce-overfitting-and-improve-generalization-error/). You've mentioned dropout, reducing model complexity and callbacks. Of course callbacks and dropout are examples of regularization, but there are also others to take into consideration. I've had some successful improvements of my RNN by utilising dropouts and callbacks as well.",
      "votes": null
    },
    {
      "id": "504392",
      "postDate": "03/31/2019 14:32:07",
      "content": "<p>DevilEars,</p>\n\n<p>Thanks for the links to the \"cheat sheets\", they are very helpful in understanding various aspects of these topics.</p>\n\n<p>Kickback</p>",
      "rawMarkdown": "DevilEars,\n\nThanks for the links to the \"cheat sheets\", they are very helpful in understanding various aspects of these topics.\n\nKickback",
      "votes": null
    },
    {
      "id": "504619",
      "postDate": "03/31/2019 21:53:21",
      "content": "<p>I use two additional strategies to fight overfitting:</p>\n\n<ul>\n<li>Adding some noise to the samples</li>\n<li>Having segments larger than 150,000 (for example 250,000) and then randomly select a subsegment of 150,000. This is a bit like cropping for images but then applied to one dimensional data.</li>\n</ul>\n\n<p>Below is the code I use for this (based on MXNet):</p>\n\n<p>```\nclass Dataset( gluon.data.Dataset ):</p>\n\n<pre><code>def __init__(self, data, noise=True):\n    super().__init__() \n    self.data = data.astype(np.float32)\n    self.noise = noise\n\ndef __len__(self):\n    return len(self.data)\n\ndef __getitem__(self, idx):\n    data = self.data[idx]\n    start = np.random.randint(0, len(data)-150_000)\n    end = start + 150_000\n    x = data[start:end, 0]\n    if self.noise:\n        noise = x*np.random.normal(scale=0.05, size=x.shape)\n        x = x + noise\n    y = data[end, 1]\n    return (x, y)\n</code></pre>\n\n<p>```</p>\n\n<p>It does help against overfitting, however now my MAE-loss won't go below 2.0 ;) </p>",
      "rawMarkdown": "I use two additional strategies to fight overfitting:\n\n* Adding some noise to the samples\n* Having segments larger than 150,000 (for example 250,000) and then randomly select a subsegment of 150,000. This is a bit like cropping for images but then applied to one dimensional data.\n\nBelow is the code I use for this (based on MXNet):\n\n```\nclass Dataset( gluon.data.Dataset ):\n\n    def __init__(self, data, noise=True):\n        super().__init__() \n        self.data = data.astype(np.float32)\n        self.noise = noise\n \n    def __len__(self):\n        return len(self.data)\n    \n    def __getitem__(self, idx):\n        data = self.data[idx]\n        start = np.random.randint(0, len(data)-150_000)\n        end = start + 150_000\n        x = data[start:end, 0]\n        if self.noise:\n            noise = x*np.random.normal(scale=0.05, size=x.shape)\n            x = x + noise\n        y = data[end, 1]\n        return (x, y)\n```\n\nIt does help against overfitting, however now my MAE-loss won't go below 2.0 ;)",
      "votes": null
    },
    {
      "id": "512321",
      "postDate": "04/10/2019 17:26:22",
      "content": "<p>Thanks Peter !</p>",
      "rawMarkdown": "Thanks Peter !",
      "votes": null
    },
    {
      "id": "514359",
      "postDate": "04/11/2019 15:35:47",
      "content": "<p>Thanks Peter - am i right to understand that point 2 is equivalent to creating training samples with overlap? (ie training instance 1 might cover row 200k to 350k, and training instance 2 might cover row 150k to 300k) So you get more training examples this way</p>",
      "rawMarkdown": "Thanks Peter - am i right to understand that point 2 is equivalent to creating training samples with overlap? (ie training instance 1 might cover row 200k to 350k, and training instance 2 might cover row 150k to 300k) So you get more training examples this way",
      "votes": null
    },
    {
      "id": "514682",
      "postDate": "04/11/2019 20:55:31",
      "content": "<p>In each epoch you draw one random consecutive segment of 150_000 out of the larger 250_000 segment. So you'll have many more variations than with a simple overlap strategy. In this case you would have 100_000 different samples. For example the following are examples of possible samples: \n<code>\na[10:150_010]\na[11:150_011]\na[1001:151_001]\n</code>\n(of course you could indeed see it as overlap with an overlap of 149_999 :)</p>",
      "rawMarkdown": "In each epoch you draw one random consecutive segment of 150_000 out of the larger 250_000 segment. So you'll have many more variations than with a simple overlap strategy. In this case you would have 100_000 different samples. For example the following are examples of possible samples: \n```\na[10:150_010]\na[11:150_011]\na[1001:151_001]\n```\n(of course you could indeed see it as overlap with an overlap of 149_999 :)",
      "votes": null
    },
    {
      "id": "514864",
      "postDate": "04/12/2019 01:38:17",
      "content": "<p>Got it thanks Peter! Is there a particular reason for choosing 250k?</p>",
      "rawMarkdown": "Got it thanks Peter! Is there a particular reason for choosing 250k?",
      "votes": null
    },
    {
      "id": "515156",
      "postDate": "04/12/2019 08:53:58",
      "content": "<p>Just one of the many hyper parameters that I tuned based on limited to no real insights  :)  The only  things I considered are:</p>\n\n<ul>\n<li><p>if you make it too big, you have only limited number of sections to split between train and validate and as a result a less well representative validation set. So setting it to 1000k would not be desirable.</p></li>\n<li><p>if setting it to a very small value (like 160k), the randomly drawn samples are too similar and still overfitting could happen.</p></li>\n</ul>\n\n<p>So I tried some values between 200k and 500k and 250k seems to avoid overfitting (in combination with dropouts) and still a reasonable representative validation set.</p>",
      "rawMarkdown": "Just one of the many hyper parameters that I tuned based on limited to no real insights  :)  The only  things I considered are:\n\n-  if you make it too big, you have only limited number of sections to split between train and validate and as a result a less well representative validation set. So setting it to 1000k would not be desirable.\n\n-  if setting it to a very small value (like 160k), the randomly drawn samples are too similar and still overfitting could happen.\n\nSo I tried some values between 200k and 500k and 250k seems to avoid overfitting (in combination with dropouts) and still a reasonable representative validation set.",
      "votes": null
    },
    {
      "id": "518318",
      "postDate": "04/17/2019 04:50:46",
      "content": "<p>I don't see exactly how RNN's are used. Training samples are ~130K in length and span basically no time relative to the earthquake period (10+ seconds). So the pieces inserted into the RNN have no significant causal evolution. </p>",
      "rawMarkdown": "I don't see exactly how RNN's are used. Training samples are ~130K in length and span basically no time relative to the earthquake period (10+ seconds). So the pieces inserted into the RNN have no significant causal evolution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 504278,
      "author_name": "devilears",
      "author_url": "",
      "post_date": "03/31/2019 10:07:32",
      "content": "<p>Greetings Tarun\nYou're placed much higher in this competition than I am, so perhaps I should be asking you for tips :)</p>\n\n<p>I've been using this <a href=\"https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks#diagnostics\">cheat sheet on machine learning</a> to work with overfitting and underfitting.</p>\n\n<p>For overfitting, the tips there are to get more data, or to <a href=\"https://machinelearningmastery.com/introduction-to-regularization-to-reduce-overfitting-and-improve-generalization-error/\">perform regularisation</a>. You've mentioned dropout, reducing model complexity and callbacks. Of course callbacks and dropout are examples of regularization, but there are also others to take into consideration. I've had some successful improvements of my RNN by utilising dropouts and callbacks as well. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 504392,
      "author_name": "bkosar1640",
      "author_url": "",
      "post_date": "03/31/2019 14:32:07",
      "content": "<p>DevilEars,</p>\n\n<p>Thanks for the links to the \"cheat sheets\", they are very helpful in understanding various aspects of these topics.</p>\n\n<p>Kickback</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 504619,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "03/31/2019 21:53:21",
      "content": "<p>I use two additional strategies to fight overfitting:</p>\n\n<ul>\n<li>Adding some noise to the samples</li>\n<li>Having segments larger than 150,000 (for example 250,000) and then randomly select a subsegment of 150,000. This is a bit like cropping for images but then applied to one dimensional data.</li>\n</ul>\n\n<p>Below is the code I use for this (based on MXNet):</p>\n\n<p>```\nclass Dataset( gluon.data.Dataset ):</p>\n\n<pre><code>def __init__(self, data, noise=True):\n    super().__init__() \n    self.data = data.astype(np.float32)\n    self.noise = noise\n\ndef __len__(self):\n    return len(self.data)\n\ndef __getitem__(self, idx):\n    data = self.data[idx]\n    start = np.random.randint(0, len(data)-150_000)\n    end = start + 150_000\n    x = data[start:end, 0]\n    if self.noise:\n        noise = x*np.random.normal(scale=0.05, size=x.shape)\n        x = x + noise\n    y = data[end, 1]\n    return (x, y)\n</code></pre>\n\n<p>```</p>\n\n<p>It does help against overfitting, however now my MAE-loss won't go below 2.0 ;) </p>",
      "votes": null,
      "replies": [
        {
          "id": 512321,
          "author_name": "tarunpaparaju",
          "author_url": "",
          "post_date": "04/10/2019 17:26:22",
          "content": "<p>Thanks Peter !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 514359,
          "author_name": "heisenger",
          "author_url": "",
          "post_date": "04/11/2019 15:35:47",
          "content": "<p>Thanks Peter - am i right to understand that point 2 is equivalent to creating training samples with overlap? (ie training instance 1 might cover row 200k to 350k, and training instance 2 might cover row 150k to 300k) So you get more training examples this way</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 514682,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "04/11/2019 20:55:31",
          "content": "<p>In each epoch you draw one random consecutive segment of 150_000 out of the larger 250_000 segment. So you'll have many more variations than with a simple overlap strategy. In this case you would have 100_000 different samples. For example the following are examples of possible samples: \n<code>\na[10:150_010]\na[11:150_011]\na[1001:151_001]\n</code>\n(of course you could indeed see it as overlap with an overlap of 149_999 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 514864,
          "author_name": "heisenger",
          "author_url": "",
          "post_date": "04/12/2019 01:38:17",
          "content": "<p>Got it thanks Peter! Is there a particular reason for choosing 250k?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 515156,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "04/12/2019 08:53:58",
          "content": "<p>Just one of the many hyper parameters that I tuned based on limited to no real insights  :)  The only  things I considered are:</p>\n\n<ul>\n<li><p>if you make it too big, you have only limited number of sections to split between train and validate and as a result a less well representative validation set. So setting it to 1000k would not be desirable.</p></li>\n<li><p>if setting it to a very small value (like 160k), the randomly drawn samples are too similar and still overfitting could happen.</p></li>\n</ul>\n\n<p>So I tried some values between 200k and 500k and 250k seems to avoid overfitting (in combination with dropouts) and still a reasonable representative validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 518318,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "04/17/2019 04:50:46",
      "content": "<p>I don't see exactly how RNN's are used. Training samples are ~130K in length and span basically no time relative to the earthquake period (10+ seconds). So the pieces inserted into the RNN have no significant causal evolution. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "498957": "Does anyone have any ideas of how to prevent RNN models from overfitting to the small number of training samples ? Here are some ideas :\n- Dropout\n- Reduce model complexity\n- EarlyStopping or ModelCheckpoint callbacks\n\nPlease feel free to share any ideas you have to reduce overfitting of RNNs.",
    "504278": "Greetings Tarun\nYou're placed much higher in this competition than I am, so perhaps I should be asking you for tips :)\n\nI've been using this [cheat sheet on machine learning](https://stanford.edu/~shervine/teaching/cs-229/cheatsheet-machine-learning-tips-and-tricks#diagnostics) to work with overfitting and underfitting.\n\nFor overfitting, the tips there are to get more data, or to [perform regularisation](https://machinelearningmastery.com/introduction-to-regularization-to-reduce-overfitting-and-improve-generalization-error/). You've mentioned dropout, reducing model complexity and callbacks. Of course callbacks and dropout are examples of regularization, but there are also others to take into consideration. I've had some successful improvements of my RNN by utilising dropouts and callbacks as well.",
    "504392": "DevilEars,\n\nThanks for the links to the \"cheat sheets\", they are very helpful in understanding various aspects of these topics.\n\nKickback",
    "504619": "I use two additional strategies to fight overfitting:\n\n* Adding some noise to the samples\n* Having segments larger than 150,000 (for example 250,000) and then randomly select a subsegment of 150,000. This is a bit like cropping for images but then applied to one dimensional data.\n\nBelow is the code I use for this (based on MXNet):\n\n```\nclass Dataset( gluon.data.Dataset ):\n\n    def __init__(self, data, noise=True):\n        super().__init__() \n        self.data = data.astype(np.float32)\n        self.noise = noise\n \n    def __len__(self):\n        return len(self.data)\n    \n    def __getitem__(self, idx):\n        data = self.data[idx]\n        start = np.random.randint(0, len(data)-150_000)\n        end = start + 150_000\n        x = data[start:end, 0]\n        if self.noise:\n            noise = x*np.random.normal(scale=0.05, size=x.shape)\n            x = x + noise\n        y = data[end, 1]\n        return (x, y)\n```\n\nIt does help against overfitting, however now my MAE-loss won't go below 2.0 ;)",
    "512321": "Thanks Peter !",
    "514359": "Thanks Peter - am i right to understand that point 2 is equivalent to creating training samples with overlap? (ie training instance 1 might cover row 200k to 350k, and training instance 2 might cover row 150k to 300k) So you get more training examples this way",
    "514682": "In each epoch you draw one random consecutive segment of 150_000 out of the larger 250_000 segment. So you'll have many more variations than with a simple overlap strategy. In this case you would have 100_000 different samples. For example the following are examples of possible samples: \n```\na[10:150_010]\na[11:150_011]\na[1001:151_001]\n```\n(of course you could indeed see it as overlap with an overlap of 149_999 :)",
    "514864": "Got it thanks Peter! Is there a particular reason for choosing 250k?",
    "515156": "Just one of the many hyper parameters that I tuned based on limited to no real insights  :)  The only  things I considered are:\n\n-  if you make it too big, you have only limited number of sections to split between train and validate and as a result a less well representative validation set. So setting it to 1000k would not be desirable.\n\n-  if setting it to a very small value (like 160k), the randomly drawn samples are too similar and still overfitting could happen.\n\nSo I tried some values between 200k and 500k and 250k seems to avoid overfitting (in combination with dropouts) and still a reasonable representative validation set.",
    "518318": "I don't see exactly how RNN's are used. Training samples are ~130K in length and span basically no time relative to the earthquake period (10+ seconds). So the pieces inserted into the RNN have no significant causal evolution."
  },
  "source": "meta"
}