{
  "id": 77266,
  "title": "One idea about difference between train set and test set.",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77266",
  "author_name": "sheep",
  "post_date": "2019-01-11T02:34:12.408000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It seem that most of the CVs in the competition are not consist with public LB. After the competition's ending, I find that my Private LB seems consist with public LB.</p>\n\n<p>One idea is that the train set and the test set comes from different batch(different Lab or somehow).</p>\n\n<p>The average size of nucleus(number of pixel in 512) in train set, test set and hpa dataset seem totally different.\n<img src=\"https://www.kaggleusercontent.com/kf/8176479/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..4l5QSJIb8PUd9fO5mHqE2w.aUViIr3EnLqanjoP6ZJq9SWJ1YEWMvXh3Ip5N2YyAE2JvdMwqeIq3evc92acyvcpRG04VLpxQAvZRn_Xf4lqgR7YPeOvXHPBgTp5CdLHI_Xp2Zr8tR46PmWe8W0iCsH-CF7q_EoMMyB7YqCVW0D_Z9jlbnsO5f7QtUjVMPyeXWo.Ktj4dnxuWWhMEnJm-77Gsg/__results___files/__results___16_2.png\" alt=\"enter image description here\"></p>\n\n<p>As for train set and test set, there are two peak exist, a combination of two different distribution.\nI guess these difference are introduced by different setting of microscope, what we call 'Batch effect' in the field of bioinformatic. </p>\n\n<p>Using a 16 time crop and resize TTA based on full TIFF image of test set, give me about 0.05 boost in PB(maybe I use a week model).</p>\n\n<p>Please post here if your CV are consist with LB or you know whey they are different.</p>",
  "messages": [
    {
      "id": 453995,
      "postDate": "2019-01-11T02:34:12.410Z",
      "content": "<p>It seem that most of the CVs in the competition are not consist with public LB. After the competition's ending, I find that my Private LB seems consist with public LB.</p>\n\n<p>One idea is that the train set and the test set comes from different batch(different Lab or somehow).</p>\n\n<p>The average size of nucleus(number of pixel in 512) in train set, test set and hpa dataset seem totally different.\n<img src=\"https://www.kaggleusercontent.com/kf/8176479/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..4l5QSJIb8PUd9fO5mHqE2w.aUViIr3EnLqanjoP6ZJq9SWJ1YEWMvXh3Ip5N2YyAE2JvdMwqeIq3evc92acyvcpRG04VLpxQAvZRn_Xf4lqgR7YPeOvXHPBgTp5CdLHI_Xp2Zr8tR46PmWe8W0iCsH-CF7q_EoMMyB7YqCVW0D_Z9jlbnsO5f7QtUjVMPyeXWo.Ktj4dnxuWWhMEnJm-77Gsg/__results___files/__results___16_2.png\" alt=\"enter image description here\"></p>\n\n<p>As for train set and test set, there are two peak exist, a combination of two different distribution.\nI guess these difference are introduced by different setting of microscope, what we call 'Batch effect' in the field of bioinformatic. </p>\n\n<p>Using a 16 time crop and resize TTA based on full TIFF image of test set, give me about 0.05 boost in PB(maybe I use a week model).</p>\n\n<p>Please post here if your CV are consist with LB or you know whey they are different.</p>",
      "rawMarkdown": "It seem that most of the CVs in the competition are not consist with public LB. After the competition's ending, I find that my Private LB seems consist with public LB.\n\nOne idea is that the train set and the test set comes from different batch(different Lab or somehow).\n\nThe average size of nucleus(number of pixel in 512) in train set, test set and hpa dataset seem totally different.\n![enter image description here][1]\n\n\n  [1]: https://www.kaggleusercontent.com/kf/8176479/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..4l5QSJIb8PUd9fO5mHqE2w.aUViIr3EnLqanjoP6ZJq9SWJ1YEWMvXh3Ip5N2YyAE2JvdMwqeIq3evc92acyvcpRG04VLpxQAvZRn_Xf4lqgR7YPeOvXHPBgTp5CdLHI_Xp2Zr8tR46PmWe8W0iCsH-CF7q_EoMMyB7YqCVW0D_Z9jlbnsO5f7QtUjVMPyeXWo.Ktj4dnxuWWhMEnJm-77Gsg/__results___files/__results___16_2.png\n\nAs for train set and test set, there are two peak exist, a combination of two different distribution.\nI guess these difference are introduced by different setting of microscope, what we call 'Batch effect' in the field of bioinformatic. \n\nUsing a 16 time crop and resize TTA based on full TIFF image of test set, give me about 0.05 boost in PB(maybe I use a week model).\n\nPlease post here if your CV are consist with LB or you know whey they are different.\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "453995": "It seem that most of the CVs in the competition are not consist with public LB. After the competition's ending, I find that my Private LB seems consist with public LB.\n\nOne idea is that the train set and the test set comes from different batch(different Lab or somehow).\n\nThe average size of nucleus(number of pixel in 512) in train set, test set and hpa dataset seem totally different.\n![enter image description here][1]\n\n\n  [1]: https://www.kaggleusercontent.com/kf/8176479/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..4l5QSJIb8PUd9fO5mHqE2w.aUViIr3EnLqanjoP6ZJq9SWJ1YEWMvXh3Ip5N2YyAE2JvdMwqeIq3evc92acyvcpRG04VLpxQAvZRn_Xf4lqgR7YPeOvXHPBgTp5CdLHI_Xp2Zr8tR46PmWe8W0iCsH-CF7q_EoMMyB7YqCVW0D_Z9jlbnsO5f7QtUjVMPyeXWo.Ktj4dnxuWWhMEnJm-77Gsg/__results___files/__results___16_2.png\n\nAs for train set and test set, there are two peak exist, a combination of two different distribution.\nI guess these difference are introduced by different setting of microscope, what we call 'Batch effect' in the field of bioinformatic. \n\nUsing a 16 time crop and resize TTA based on full TIFF image of test set, give me about 0.05 boost in PB(maybe I use a week model).\n\nPlease post here if your CV are consist with LB or you know whey they are different.\n"
  }
}