{
  "id": 396851,
  "title": "Tedious but mandatory checks (pytorch)",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/396851",
  "author_name": "",
  "post_date": "2023-03-23T05:10:13.960102100Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am using pytorch: It seems in all competitions like this one, we all go through the same boring/tedious but unavoidable checks. <br>\nFirst, we need to overfit: make sure a model is powerful enough. Then, check dropouts, weight decay in AdamW, etc etc, mixup, and on and on, all the standard machinery.</p>\n<p>Maybe we could share to try to cut down on the boring and have more time for fun?</p>\n<p>Here is my finding so far. \"Any\" trivial LSTM-like pytorch model overfits around 500 epochs. 1d-convs+LSTM takes longer: around 600-700 epochs. <br>\nTried weight-decay=0. Have not tried dropouts yet. Trying mixup now.</p>\n<p>Anyone?</p>",
  "messages": [
    {
      "id": "2193114",
      "postDate": "03/23/2023 05:10:13",
      "content": "<p>I am using pytorch: It seems in all competitions like this one, we all go through the same boring/tedious but unavoidable checks. <br>\nFirst, we need to overfit: make sure a model is powerful enough. Then, check dropouts, weight decay in AdamW, etc etc, mixup, and on and on, all the standard machinery.</p>\n<p>Maybe we could share to try to cut down on the boring and have more time for fun?</p>\n<p>Here is my finding so far. \"Any\" trivial LSTM-like pytorch model overfits around 500 epochs. 1d-convs+LSTM takes longer: around 600-700 epochs. <br>\nTried weight-decay=0. Have not tried dropouts yet. Trying mixup now.</p>\n<p>Anyone?</p>",
      "rawMarkdown": "I am using pytorch: It seems in all competitions like this one, we all go through the same boring/tedious but unavoidable checks. \nFirst, we need to overfit: make sure a model is powerful enough. Then, check dropouts, weight decay in AdamW, etc etc, mixup, and on and on, all the standard machinery.\n\nMaybe we could share to try to cut down on the boring and have more time for fun?\n\nHere is my finding so far. \"Any\" trivial LSTM-like pytorch model overfits around 500 epochs. 1d-convs+LSTM takes longer: around 600-700 epochs. \nTried weight-decay=0. Have not tried dropouts yet. Trying mixup now.\n\nAnyone?",
      "votes": null
    },
    {
      "id": "2193494",
      "postDate": "03/23/2023 10:41:12",
      "content": "<p>textbook mixup (mixing labels, not losses), beta=alpha=0.5 with 0.5 probability gives a small but noticeable improvement in val metrics.</p>",
      "rawMarkdown": "textbook mixup (mixing labels, not losses), beta=alpha=0.5 with 0.5 probability gives a small but noticeable improvement in val metrics.",
      "votes": null
    },
    {
      "id": "2194398",
      "postDate": "03/23/2023 22:41:02",
      "content": "<p>Adding 1d-BatchNorm(s) really speeds up overfitting but val metrics become much worse.</p>",
      "rawMarkdown": "Adding 1d-BatchNorm(s) really speeds up overfitting but val metrics become much worse.",
      "votes": null
    },
    {
      "id": "2194399",
      "postDate": "03/23/2023 22:41:59",
      "content": "<p>Has anyone tried data dropouts and/or adding noise to reduce overfitting?</p>",
      "rawMarkdown": "Has anyone tried data dropouts and/or adding noise to reduce overfitting?",
      "votes": null
    },
    {
      "id": "2194510",
      "postDate": "03/24/2023 02:00:26",
      "content": "<p>I tried adding dropout multiple times, with some changes to my data preprocessing between attempts. It did more harm than good every time. This may be a peculiarity of the model architecture I'm using though, so your mileage may vary.</p>\n<p>I've had some success with a couple data augmentation approaches that could be viewed as ways to add noise. Overfitting is still a major problem for me though.</p>",
      "rawMarkdown": "I tried adding dropout multiple times, with some changes to my data preprocessing between attempts. It did more harm than good every time. This may be a peculiarity of the model architecture I'm using though, so your mileage may vary.\n\nI've had some success with a couple data augmentation approaches that could be viewed as ways to add noise. Overfitting is still a major problem for me though.",
      "votes": null
    },
    {
      "id": "2194842",
      "postDate": "03/24/2023 07:34:49",
      "content": "<p>I noticed that model-level dropouts (in the head or lstm) do not improve anything, just making the training longer, and not better val scores.</p>",
      "rawMarkdown": "I noticed that model-level dropouts (in the head or lstm) do not improve anything, just making the training longer, and not better val scores.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2193494,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "03/23/2023 10:41:12",
      "content": "<p>textbook mixup (mixing labels, not losses), beta=alpha=0.5 with 0.5 probability gives a small but noticeable improvement in val metrics.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2194398,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "03/23/2023 22:41:02",
      "content": "<p>Adding 1d-BatchNorm(s) really speeds up overfitting but val metrics become much worse.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2194399,
      "author_name": "dmitrykonovalov",
      "author_url": "",
      "post_date": "03/23/2023 22:41:59",
      "content": "<p>Has anyone tried data dropouts and/or adding noise to reduce overfitting?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2194510,
          "author_name": "jsday96",
          "author_url": "",
          "post_date": "03/24/2023 02:00:26",
          "content": "<p>I tried adding dropout multiple times, with some changes to my data preprocessing between attempts. It did more harm than good every time. This may be a peculiarity of the model architecture I'm using though, so your mileage may vary.</p>\n<p>I've had some success with a couple data augmentation approaches that could be viewed as ways to add noise. Overfitting is still a major problem for me though.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2194842,
              "author_name": "dmitrykonovalov",
              "author_url": "",
              "post_date": "03/24/2023 07:34:49",
              "content": "<p>I noticed that model-level dropouts (in the head or lstm) do not improve anything, just making the training longer, and not better val scores.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2193114": "I am using pytorch: It seems in all competitions like this one, we all go through the same boring/tedious but unavoidable checks. \nFirst, we need to overfit: make sure a model is powerful enough. Then, check dropouts, weight decay in AdamW, etc etc, mixup, and on and on, all the standard machinery.\n\nMaybe we could share to try to cut down on the boring and have more time for fun?\n\nHere is my finding so far. \"Any\" trivial LSTM-like pytorch model overfits around 500 epochs. 1d-convs+LSTM takes longer: around 600-700 epochs. \nTried weight-decay=0. Have not tried dropouts yet. Trying mixup now.\n\nAnyone?",
    "2193494": "textbook mixup (mixing labels, not losses), beta=alpha=0.5 with 0.5 probability gives a small but noticeable improvement in val metrics.",
    "2194398": "Adding 1d-BatchNorm(s) really speeds up overfitting but val metrics become much worse.",
    "2194399": "Has anyone tried data dropouts and/or adding noise to reduce overfitting?",
    "2194510": "I tried adding dropout multiple times, with some changes to my data preprocessing between attempts. It did more harm than good every time. This may be a peculiarity of the model architecture I'm using though, so your mileage may vary.\n\nI've had some success with a couple data augmentation approaches that could be viewed as ways to add noise. Overfitting is still a major problem for me though.",
    "2194842": "I noticed that model-level dropouts (in the head or lstm) do not improve anything, just making the training longer, and not better val scores."
  },
  "source": "meta"
}