{
  "id": 203319,
  "title": "Samples that are different from other",
  "url": "/competitions/rfcx-species-audio-detection/discussion/203319",
  "author_name": "",
  "post_date": "2020-12-14T18:53:34.447755400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi to all!<br>\nAs I see we have a group of samples that are different from others (for example, this can be seen in public notebooks when most of these samples are in the validation set and our val_loss \"explodes\" from 0, xxx to 1000, xxx ). Have you tried to find such samples and how do they differ?</p>",
  "messages": [
    {
      "id": "1112628",
      "postDate": "12/14/2020 18:53:34",
      "content": "<p>Hi to all!<br>\nAs I see we have a group of samples that are different from others (for example, this can be seen in public notebooks when most of these samples are in the validation set and our val_loss \"explodes\" from 0, xxx to 1000, xxx ). Have you tried to find such samples and how do they differ?</p>",
      "rawMarkdown": "Hi to all!\nAs I see we have a group of samples that are different from others (for example, this can be seen in public notebooks when most of these samples are in the validation set and our val_loss \"explodes\" from 0, xxx to 1000, xxx ). Have you tried to find such samples and how do they differ?",
      "votes": null
    },
    {
      "id": "1113253",
      "postDate": "12/15/2020 10:23:08",
      "content": "<p>can you name few of these samples so its easy to check this behavior when used in validation </p>",
      "rawMarkdown": "can you name few of these samples so its easy to check this behavior when used in validation",
      "votes": null
    },
    {
      "id": "1113815",
      "postDate": "12/15/2020 17:54:53",
      "content": "<p>No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a>, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!</p>",
      "rawMarkdown": "No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!",
      "votes": null
    },
    {
      "id": "1115739",
      "postDate": "12/16/2020 14:25:50",
      "content": "<blockquote>\n  <p>No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a>, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!</p>\n</blockquote>\n<p>I guess this depends also on the optimization you are using, and not only on the samples that are falling into the different folds. From what I see you are using a fast training method (super-convergence), with a very big network (ResNet-101), so it is reasonable to see jumps on the validation (even on training) loss/metrics, especially at the beginning of the training.</p>\n<p>I think you can fix this issue by just simply looking at input samples when the validation loss has a huge value.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
      "rawMarkdown": "> No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!\n\nI guess this depends also on the optimization you are using, and not only on the samples that are falling into the different folds. From what I see you are using a fast training method (super-convergence), with a very big network (ResNet-101), so it is reasonable to see jumps on the validation (even on training) loss/metrics, especially at the beginning of the training.\n\nI think you can fix this issue by just simply looking at input samples when the validation loss has a huge value.\n\nBest,\n\nGuglielmo",
      "votes": null
    },
    {
      "id": "1115806",
      "postDate": "12/16/2020 15:35:40",
      "content": "<p>Thanks Guglielmo! I will try to</p>",
      "rawMarkdown": "Thanks Guglielmo! I will try to",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1113253,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "12/15/2020 10:23:08",
      "content": "<p>can you name few of these samples so its easy to check this behavior when used in validation </p>",
      "votes": null,
      "replies": [
        {
          "id": 1113815,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "12/15/2020 17:54:53",
          "content": "<p>No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a>, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115739,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "12/16/2020 14:25:50",
          "content": "<blockquote>\n  <p>No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook <a href=\"https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu\" target=\"_blank\">https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu</a>, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!</p>\n</blockquote>\n<p>I guess this depends also on the optimization you are using, and not only on the samples that are falling into the different folds. From what I see you are using a fast training method (super-convergence), with a very big network (ResNet-101), so it is reasonable to see jumps on the validation (even on training) loss/metrics, especially at the beginning of the training.</p>\n<p>I think you can fix this issue by just simply looking at input samples when the validation loss has a huge value.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1115806,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "12/16/2020 15:35:40",
          "content": "<p>Thanks Guglielmo! I will try to</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1112628": "Hi to all!\nAs I see we have a group of samples that are different from others (for example, this can be seen in public notebooks when most of these samples are in the validation set and our val_loss \"explodes\" from 0, xxx to 1000, xxx ). Have you tried to find such samples and how do they differ?",
    "1113253": "can you name few of these samples so its easy to check this behavior when used in validation",
    "1113815": "No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!",
    "1115739": "> No, I did not identify them. I thought that such a group of samples should be for the following reasons: When using Stratified 5-Fold we get different results during training. After some modifications to this notebook https://www.kaggle.com/yosshi999/rfcx-train-resnet50-with-tpu, I noticed that SOMETIMES val_loss grows sharply after several epochs (This is best seen when using resnet101). At first I thought that there was some error in the implementation, but then I realized that there is a clear relationship between the samples that fall into KFold!\n\nI guess this depends also on the optimization you are using, and not only on the samples that are falling into the different folds. From what I see you are using a fast training method (super-convergence), with a very big network (ResNet-101), so it is reasonable to see jumps on the validation (even on training) loss/metrics, especially at the beginning of the training.\n\nI think you can fix this issue by just simply looking at input samples when the validation loss has a huge value.\n\nBest,\n\nGuglielmo",
    "1115806": "Thanks Guglielmo! I will try to"
  },
  "source": "meta"
}