{
  "id": 401676,
  "title": "Any idea for this hypothesis?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/401676",
  "author_name": "",
  "post_date": "2023-04-14T11:37:41.620140600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I trained a model on all data, which results in poor performance than 200 batches. </p>\n<p>I think that this result is created by similarity of each data, which can make overfitting even though the number of iterations is low than 200 batches. (the ratio of iteration between all and 200 is lower than 3, which means I have tried to avoid overfitting while changing the number of batches.)</p>\n<p>This means that we don't need to train with all data in terms of correlation of the data.</p>\n<p>Any idea of this situation or is there something missing that I am? </p>\n<p>[Edit 1] : I found something, Manoj et al. [1] \"The GRU model has fewer gates compared to LSTM and has been found to outperform LSTM when dealing with smaller datasets.\"</p>\n<p>Therefore, I can assume that I should change model structure if hyperparameters aren't wrong.</p>\n<p>[1] Deep BLSTM-GRU Model for Monthly Rainfall Prediction: A Case Study of Simtokha, Bhutan</p>",
  "messages": [
    {
      "id": "2221568",
      "postDate": "04/14/2023 11:37:41",
      "content": "<p>I trained a model on all data, which results in poor performance than 200 batches. </p>\n<p>I think that this result is created by similarity of each data, which can make overfitting even though the number of iterations is low than 200 batches. (the ratio of iteration between all and 200 is lower than 3, which means I have tried to avoid overfitting while changing the number of batches.)</p>\n<p>This means that we don't need to train with all data in terms of correlation of the data.</p>\n<p>Any idea of this situation or is there something missing that I am? </p>\n<p>[Edit 1] : I found something, Manoj et al. [1] \"The GRU model has fewer gates compared to LSTM and has been found to outperform LSTM when dealing with smaller datasets.\"</p>\n<p>Therefore, I can assume that I should change model structure if hyperparameters aren't wrong.</p>\n<p>[1] Deep BLSTM-GRU Model for Monthly Rainfall Prediction: A Case Study of Simtokha, Bhutan</p>",
      "rawMarkdown": "I trained a model on all data, which results in poor performance than 200 batches. \n\nI think that this result is created by similarity of each data, which can make overfitting even though the number of iterations is low than 200 batches. (the ratio of iteration between all and 200 is lower than 3, which means I have tried to avoid overfitting while changing the number of batches.)\n\nThis means that we don't need to train with all data in terms of correlation of the data.\n\nAny idea of this situation or is there something missing that I am? \n\n[Edit 1] : I found something, Manoj et al. [1] \"The GRU model has fewer gates compared to LSTM and has been found to outperform LSTM when dealing with smaller datasets.\"\n\nTherefore, I can assume that I should change model structure if hyperparameters aren't wrong.\n\n[1] Deep BLSTM-GRU Model for Monthly Rainfall Prediction: A Case Study of Simtokha, Bhutan",
      "votes": null
    },
    {
      "id": "2222020",
      "postDate": "04/14/2023 19:40:33",
      "content": "<p>I trained for 650 batches and can confirm that it is not worse than 200, perhaps your model is too small or has some other issue?</p>",
      "rawMarkdown": "I trained for 650 batches and can confirm that it is not worse than 200, perhaps your model is too small or has some other issue?",
      "votes": null
    },
    {
      "id": "2222149",
      "postDate": "04/15/2023 00:46:11",
      "content": "<p>I think you are right. I experimented with a new model that has lower parameters than before, which results in lower performance even on 200 (tested numbers: 2.5m -&gt; 1.2m). I don't need to modify the hyperparameters of models anymore but should focus on building a new model.</p>",
      "rawMarkdown": "I think you are right. I experimented with a new model that has lower parameters than before, which results in lower performance even on 200 (tested numbers: 2.5m -> 1.2m). I don't need to modify the hyperparameters of models anymore but should focus on building a new model.",
      "votes": null
    },
    {
      "id": "2228434",
      "postDate": "04/20/2023 14:48:29",
      "content": "<p>Hello there,</p>\n<p>I suggest you try to:</p>\n<ul>\n<li>Analyzing the model's training/validation loss and accuracy to see if overfitting is indeed the issue</li>\n<li>Experimenting with different model architectures, hyperparameters and optimizer to see if the performance improves (did you manage to increase it's val score? How did it translate to the leaderboard?)</li>\n</ul>\n<p>The Devastator.</p>",
      "rawMarkdown": "Hello there,\n\nI suggest you try to:\n\n- Analyzing the model's training/validation loss and accuracy to see if overfitting is indeed the issue\n- Experimenting with different model architectures, hyperparameters and optimizer to see if the performance improves (did you manage to increase it's val score? How did it translate to the leaderboard?)\n\n\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2228455",
      "postDate": "04/20/2023 15:04:10",
      "content": "<p>Thank you for your suggestion. Several days ago I was only trying to evaluate my model on LB, I realized that it shouldn't do that.</p>\n<p>About other architecture, after this competition, I found out many kagglers use a variant of a transformer model. I'm going to analyze and apply this. </p>\n<p>Thanks.</p>",
      "rawMarkdown": "Thank you for your suggestion. Several days ago I was only trying to evaluate my model on LB, I realized that it shouldn't do that.\n\nAbout other architecture, after this competition, I found out many kagglers use a variant of a transformer model. I'm going to analyze and apply this. \n\nThanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2222020,
      "author_name": "dipamc77",
      "author_url": "",
      "post_date": "04/14/2023 19:40:33",
      "content": "<p>I trained for 650 batches and can confirm that it is not worse than 200, perhaps your model is too small or has some other issue?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2222149,
          "author_name": "hyunsoolee1010",
          "author_url": "",
          "post_date": "04/15/2023 00:46:11",
          "content": "<p>I think you are right. I experimented with a new model that has lower parameters than before, which results in lower performance even on 200 (tested numbers: 2.5m -&gt; 1.2m). I don't need to modify the hyperparameters of models anymore but should focus on building a new model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2228434,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "04/20/2023 14:48:29",
      "content": "<p>Hello there,</p>\n<p>I suggest you try to:</p>\n<ul>\n<li>Analyzing the model's training/validation loss and accuracy to see if overfitting is indeed the issue</li>\n<li>Experimenting with different model architectures, hyperparameters and optimizer to see if the performance improves (did you manage to increase it's val score? How did it translate to the leaderboard?)</li>\n</ul>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2228455,
          "author_name": "hyunsoolee1010",
          "author_url": "",
          "post_date": "04/20/2023 15:04:10",
          "content": "<p>Thank you for your suggestion. Several days ago I was only trying to evaluate my model on LB, I realized that it shouldn't do that.</p>\n<p>About other architecture, after this competition, I found out many kagglers use a variant of a transformer model. I'm going to analyze and apply this. </p>\n<p>Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2221568": "I trained a model on all data, which results in poor performance than 200 batches. \n\nI think that this result is created by similarity of each data, which can make overfitting even though the number of iterations is low than 200 batches. (the ratio of iteration between all and 200 is lower than 3, which means I have tried to avoid overfitting while changing the number of batches.)\n\nThis means that we don't need to train with all data in terms of correlation of the data.\n\nAny idea of this situation or is there something missing that I am? \n\n[Edit 1] : I found something, Manoj et al. [1] \"The GRU model has fewer gates compared to LSTM and has been found to outperform LSTM when dealing with smaller datasets.\"\n\nTherefore, I can assume that I should change model structure if hyperparameters aren't wrong.\n\n[1] Deep BLSTM-GRU Model for Monthly Rainfall Prediction: A Case Study of Simtokha, Bhutan",
    "2222020": "I trained for 650 batches and can confirm that it is not worse than 200, perhaps your model is too small or has some other issue?",
    "2222149": "I think you are right. I experimented with a new model that has lower parameters than before, which results in lower performance even on 200 (tested numbers: 2.5m -> 1.2m). I don't need to modify the hyperparameters of models anymore but should focus on building a new model.",
    "2228434": "Hello there,\n\nI suggest you try to:\n\n- Analyzing the model's training/validation loss and accuracy to see if overfitting is indeed the issue\n- Experimenting with different model architectures, hyperparameters and optimizer to see if the performance improves (did you manage to increase it's val score? How did it translate to the leaderboard?)\n\n\n\nThe Devastator.",
    "2228455": "Thank you for your suggestion. Several days ago I was only trying to evaluate my model on LB, I realized that it shouldn't do that.\n\nAbout other architecture, after this competition, I found out many kagglers use a variant of a transformer model. I'm going to analyze and apply this. \n\nThanks."
  },
  "source": "meta"
}