{
  "id": 174528,
  "title": "Are you all using the same transforms in train, validation, OOF and test?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174528",
  "author_name": "",
  "post_date": "2020-08-14T00:20:52.911189900Z",
  "votes": 1,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I have seen a few notebooks not doing transforms during validation, and just using them during train, oof, and test.  What do you find is working best using the same transforms all across the board or do you do something different for some of the phases?</p>\n<p>My personal workflow is I augment train, oof, and test but not validation. I do put validation through Normalization though, and I am wondering if people use things like <code>color_constancy</code> on validation as well, since that is sort of normalization as opposed to an augmentation.</p>",
  "messages": [
    {
      "id": "969790",
      "postDate": "08/14/2020 00:20:52",
      "content": "<p>I have seen a few notebooks not doing transforms during validation, and just using them during train, oof, and test.  What do you find is working best using the same transforms all across the board or do you do something different for some of the phases?</p>\n<p>My personal workflow is I augment train, oof, and test but not validation. I do put validation through Normalization though, and I am wondering if people use things like <code>color_constancy</code> on validation as well, since that is sort of normalization as opposed to an augmentation.</p>",
      "rawMarkdown": "I have seen a few notebooks not doing transforms during validation, and just using them during train, oof, and test.  What do you find is working best using the same transforms all across the board or do you do something different for some of the phases?\n\nMy personal workflow is I augment train, oof, and test but not validation. I do put validation through Normalization though, and I am wondering if people use things like `color_constancy` on validation as well, since that is sort of normalization as opposed to an augmentation.",
      "votes": null
    },
    {
      "id": "969930",
      "postDate": "08/14/2020 04:28:28",
      "content": "<p>No. when making OOF or predicting TEST, I don't apply transforms<br>\ntransforms are applied only for training </p>",
      "rawMarkdown": "No. when making OOF or predicting TEST, I don't apply transforms\ntransforms are applied only for training",
      "votes": null
    },
    {
      "id": "969935",
      "postDate": "08/14/2020 04:32:46",
      "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> so you don't do TTA.  TTA seems to help at least for me.</p>",
      "rawMarkdown": "deepkim so you don't do TTA.  TTA seems to help at least for me.",
      "votes": null
    },
    {
      "id": "970095",
      "postDate": "08/14/2020 07:37:33",
      "content": "<p>No, TTA really is helping, I am using <code>train_transforms</code> on train, test and oof while <code>validation_transforms</code> on validation.</p>",
      "rawMarkdown": "No, TTA really is helping, I am using `train_transforms` on train, test and oof while `validation_transforms` on validation.",
      "votes": null
    },
    {
      "id": "970103",
      "postDate": "08/14/2020 07:54:50",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> how do your validation transforms differ? just normalization of some sort?</p>",
      "rawMarkdown": "sarques how do your validation transforms differ? just normalization of some sort?",
      "votes": null
    },
    {
      "id": "970108",
      "postDate": "08/14/2020 07:59:22",
      "content": "<p>I added a bit of Flip in validation with normalization, it gave a little bit better results than without it. In train, I am using RandomResize, Flip, Hair Augmentation etc.</p>",
      "rawMarkdown": "I added a bit of Flip in validation with normalization, it gave a little bit better results than without it. In train, I am using RandomResize, Flip, Hair Augmentation etc.",
      "votes": null
    },
    {
      "id": "970611",
      "postDate": "08/14/2020 15:55:04",
      "content": "<p>No. I use less transforms in TTA than in training.</p>",
      "rawMarkdown": "No. I use less transforms in TTA than in training.",
      "votes": null
    },
    {
      "id": "970852",
      "postDate": "08/14/2020 22:04:02",
      "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> So I assume you have tried to use your train transforms on validation?  My own experience is that using the same transforms on validation as on train gives the best result.  Noticeably, as in you can see in the first few epochs how it's working.  I tried to just do normalization on validation, and some other things like color constancy, but it was never as good as using the same transform set.</p>",
      "rawMarkdown": "sarques So I assume you have tried to use your train transforms on validation?  My own experience is that using the same transforms on validation as on train gives the best result.  Noticeably, as in you can see in the first few epochs how it's working.  I tried to just do normalization on validation, and some other things like color constancy, but it was never as good as using the same transform set.",
      "votes": null
    },
    {
      "id": "970866",
      "postDate": "08/14/2020 22:49:12",
      "content": "<p>what's the difference between your oof and your validation set?</p>",
      "rawMarkdown": "what's the difference between your oof and your validation set?",
      "votes": null
    },
    {
      "id": "970870",
      "postDate": "08/14/2020 22:59:02",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> not sure who you are asking. My OOF uses same transforms as train and validation.  Only difference is OOF is using TTA.</p>",
      "rawMarkdown": "optimo not sure who you are asking. My OOF uses same transforms as train and validation.  Only difference is OOF is using TTA.",
      "votes": null
    },
    {
      "id": "970882",
      "postDate": "08/15/2020 00:03:12",
      "content": "<p>I have the same question as Optimo.  When we say CV score we means the score obtained on OOF predictions.  You make a difference between CV and OOF where no one else does AFAIK.  Please enlighten us.</p>",
      "rawMarkdown": "I have the same question as Optimo.  When we say CV score we means the score obtained on OOF predictions.  You make a difference between CV and OOF where no one else does AFAIK.  Please enlighten us.",
      "votes": null
    },
    {
      "id": "970886",
      "postDate": "08/15/2020 00:28:35",
      "content": "<p>I think topic author has 3 different predictions</p>\n<ul>\n<li>OOF with TTA </li>\n<li>Validation set (by verbose option active)</li>\n<li>Test set (public part)</li>\n</ul>\n<p>In this case oof predictions metric differ from validation part (as validation part during training doesn't have augmentations - to make training/validation faster and less memory consuming, and doesn't have \"internal\" TTA blending)</p>\n<hr>\n<blockquote>\n  <p>You make a difference between CV and OOF where no one else does AFAIK</p>\n</blockquote>\n<p>CV should \"mimic\" test set part - and in this case, yes  CV score should be considered with TTA (if you use it for Test part predictions).</p>",
      "rawMarkdown": "I think topic author has 3 different predictions\n- OOF with TTA \n- Validation set (by verbose option active)\n- Test set (public part)\n\nIn this case oof predictions metric differ from validation part (as validation part during training doesn't have augmentations - to make training/validation faster and less memory consuming, and doesn't have \"internal\" TTA blending)\n\n---\n\n> You make a difference between CV and OOF where no one else does AFAIK\n\nCV should \"mimic\" test set part - and in this case, yes  CV score should be considered with TTA (if you use it for Test part predictions).",
      "votes": null
    },
    {
      "id": "970896",
      "postDate": "08/15/2020 00:48:14",
      "content": "<p>I guess I would call this validation my \"early stopping metric\", but I do understand why you make such a distinction</p>",
      "rawMarkdown": "I guess I would call this validation my \"early stopping metric\", but I do understand why you make such a distinction",
      "votes": null
    },
    {
      "id": "970910",
      "postDate": "08/15/2020 02:02:59",
      "content": "<p>I do train, validation, OOF+TTA, test+TTA</p>",
      "rawMarkdown": "I do train, validation, OOF+TTA, test+TTA",
      "votes": null
    },
    {
      "id": "970912",
      "postDate": "08/15/2020 02:03:35",
      "content": "<p>Yes I use it for finding my best model/early stopping</p>",
      "rawMarkdown": "Yes I use it for finding my best model/early stopping",
      "votes": null
    },
    {
      "id": "970913",
      "postDate": "08/15/2020 02:04:30",
      "content": "<p>Same, when I refer to CV I mean OOF</p>",
      "rawMarkdown": "Same, when I refer to CV I mean OOF",
      "votes": null
    },
    {
      "id": "970984",
      "postDate": "08/15/2020 04:23:59",
      "content": "<p>Ohh, yes, sorry for the bad explanation. So you are doing normalization and color constancy on validation and it somehow is not giving better results, right? Since you said using the same transform set for both train and validation is giving better results, I guess we need same transforms for all three, i.e., train, oof and test.</p>",
      "rawMarkdown": "Ohh, yes, sorry for the bad explanation. So you are doing normalization and color constancy on validation and it somehow is not giving better results, right? Since you said using the same transform set for both train and validation is giving better results, I guess we need same transforms for all three, i.e., train, oof and test.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969930,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "08/14/2020 04:28:28",
      "content": "<p>No. when making OOF or predicting TEST, I don't apply transforms<br>\ntransforms are applied only for training </p>",
      "votes": null,
      "replies": [
        {
          "id": 969935,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/14/2020 04:32:46",
          "content": "<p><a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> so you don't do TTA.  TTA seems to help at least for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970095,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/14/2020 07:37:33",
          "content": "<p>No, TTA really is helping, I am using <code>train_transforms</code> on train, test and oof while <code>validation_transforms</code> on validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970103,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/14/2020 07:54:50",
          "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> how do your validation transforms differ? just normalization of some sort?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970108,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/14/2020 07:59:22",
          "content": "<p>I added a bit of Flip in validation with normalization, it gave a little bit better results than without it. In train, I am using RandomResize, Flip, Hair Augmentation etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970852,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/14/2020 22:04:02",
          "content": "<p><a href=\"https://www.kaggle.com/sarques\" target=\"_blank\">@sarques</a> So I assume you have tried to use your train transforms on validation?  My own experience is that using the same transforms on validation as on train gives the best result.  Noticeably, as in you can see in the first few epochs how it's working.  I tried to just do normalization on validation, and some other things like color constancy, but it was never as good as using the same transform set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970984,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/15/2020 04:23:59",
          "content": "<p>Ohh, yes, sorry for the bad explanation. So you are doing normalization and color constancy on validation and it somehow is not giving better results, right? Since you said using the same transform set for both train and validation is giving better results, I guess we need same transforms for all three, i.e., train, oof and test.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970611,
      "author_name": "bfishh",
      "author_url": "",
      "post_date": "08/14/2020 15:55:04",
      "content": "<p>No. I use less transforms in TTA than in training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 970866,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/14/2020 22:49:12",
      "content": "<p>what's the difference between your oof and your validation set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 970870,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/14/2020 22:59:02",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> not sure who you are asking. My OOF uses same transforms as train and validation.  Only difference is OOF is using TTA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970882,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2020 00:03:12",
          "content": "<p>I have the same question as Optimo.  When we say CV score we means the score obtained on OOF predictions.  You make a difference between CV and OOF where no one else does AFAIK.  Please enlighten us.</p>",
          "votes": null,
          "replies": [
            {
              "id": 970913,
              "author_name": "brianfeeny",
              "author_url": "",
              "post_date": "08/15/2020 02:04:30",
              "content": "<p>Same, when I refer to CV I mean OOF</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 970886,
          "author_name": "kyakovlev",
          "author_url": "",
          "post_date": "08/15/2020 00:28:35",
          "content": "<p>I think topic author has 3 different predictions</p>\n<ul>\n<li>OOF with TTA </li>\n<li>Validation set (by verbose option active)</li>\n<li>Test set (public part)</li>\n</ul>\n<p>In this case oof predictions metric differ from validation part (as validation part during training doesn't have augmentations - to make training/validation faster and less memory consuming, and doesn't have \"internal\" TTA blending)</p>\n<hr>\n<blockquote>\n  <p>You make a difference between CV and OOF where no one else does AFAIK</p>\n</blockquote>\n<p>CV should \"mimic\" test set part - and in this case, yes  CV score should be considered with TTA (if you use it for Test part predictions).</p>",
          "votes": null,
          "replies": [
            {
              "id": 970910,
              "author_name": "brianfeeny",
              "author_url": "",
              "post_date": "08/15/2020 02:02:59",
              "content": "<p>I do train, validation, OOF+TTA, test+TTA</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 970896,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/15/2020 00:48:14",
          "content": "<p>I guess I would call this validation my \"early stopping metric\", but I do understand why you make such a distinction</p>",
          "votes": null,
          "replies": [
            {
              "id": 970912,
              "author_name": "brianfeeny",
              "author_url": "",
              "post_date": "08/15/2020 02:03:35",
              "content": "<p>Yes I use it for finding my best model/early stopping</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "969790": "I have seen a few notebooks not doing transforms during validation, and just using them during train, oof, and test.  What do you find is working best using the same transforms all across the board or do you do something different for some of the phases?\n\nMy personal workflow is I augment train, oof, and test but not validation. I do put validation through Normalization though, and I am wondering if people use things like `color_constancy` on validation as well, since that is sort of normalization as opposed to an augmentation.",
    "969930": "No. when making OOF or predicting TEST, I don't apply transforms\ntransforms are applied only for training",
    "969935": "deepkim so you don't do TTA.  TTA seems to help at least for me.",
    "970095": "No, TTA really is helping, I am using `train_transforms` on train, test and oof while `validation_transforms` on validation.",
    "970103": "sarques how do your validation transforms differ? just normalization of some sort?",
    "970108": "I added a bit of Flip in validation with normalization, it gave a little bit better results than without it. In train, I am using RandomResize, Flip, Hair Augmentation etc.",
    "970611": "No. I use less transforms in TTA than in training.",
    "970852": "sarques So I assume you have tried to use your train transforms on validation?  My own experience is that using the same transforms on validation as on train gives the best result.  Noticeably, as in you can see in the first few epochs how it's working.  I tried to just do normalization on validation, and some other things like color constancy, but it was never as good as using the same transform set.",
    "970866": "what's the difference between your oof and your validation set?",
    "970870": "optimo not sure who you are asking. My OOF uses same transforms as train and validation.  Only difference is OOF is using TTA.",
    "970882": "I have the same question as Optimo.  When we say CV score we means the score obtained on OOF predictions.  You make a difference between CV and OOF where no one else does AFAIK.  Please enlighten us.",
    "970886": "I think topic author has 3 different predictions\n- OOF with TTA \n- Validation set (by verbose option active)\n- Test set (public part)\n\nIn this case oof predictions metric differ from validation part (as validation part during training doesn't have augmentations - to make training/validation faster and less memory consuming, and doesn't have \"internal\" TTA blending)\n\n---\n\n> You make a difference between CV and OOF where no one else does AFAIK\n\nCV should \"mimic\" test set part - and in this case, yes  CV score should be considered with TTA (if you use it for Test part predictions).",
    "970896": "I guess I would call this validation my \"early stopping metric\", but I do understand why you make such a distinction",
    "970910": "I do train, validation, OOF+TTA, test+TTA",
    "970912": "Yes I use it for finding my best model/early stopping",
    "970913": "Same, when I refer to CV I mean OOF",
    "970984": "Ohh, yes, sorry for the bad explanation. So you are doing normalization and color constancy on validation and it somehow is not giving better results, right? Since you said using the same transform set for both train and validation is giving better results, I guess we need same transforms for all three, i.e., train, oof and test."
  },
  "source": "meta"
}