{
  "id": 71147,
  "title": "What to trust?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/71147",
  "author_name": "",
  "post_date": "2018-11-10T16:11:22.417413300Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>If the train distribution is different from the public LB distribution, and the private LB distribution is likely different from public, what validation strategy can we use to be confident that one model is better than another? Right now I'm just doing a random validation split, I saw some people mention they're trying a stratified split, was wondering if anyone had thoughts on other possible approaches.</p>",
  "messages": [
    {
      "id": "418792",
      "postDate": "11/10/2018 16:11:22",
      "content": "<p>If the train distribution is different from the public LB distribution, and the private LB distribution is likely different from public, what validation strategy can we use to be confident that one model is better than another? Right now I'm just doing a random validation split, I saw some people mention they're trying a stratified split, was wondering if anyone had thoughts on other possible approaches.</p>",
      "rawMarkdown": "If the train distribution is different from the public LB distribution, and the private LB distribution is likely different from public, what validation strategy can we use to be confident that one model is better than another? Right now I'm just doing a random validation split, I saw some people mention they're trying a stratified split, was wondering if anyone had thoughts on other possible approaches.",
      "votes": null
    },
    {
      "id": "418902",
      "postDate": "11/10/2018 20:48:18",
      "content": "<p>This is what I use, it came from one of the starter kernels. It works well enough for me so far. I also plan to try the method that <a href=\"/trentb\">@trentb</a> posted in the thread he created - multilabel stratification python package.</p>\n\n<pre>train_df, valid_df = train_test_split(train_df, \n                 test_size = 0.1,\n                 random_state=42,\n                  # hack to make stratification work                  \n                 stratify = train_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</pre>",
      "rawMarkdown": "This is what I use, it came from one of the starter kernels. It works well enough for me so far. I also plan to try the method that @trentb posted in the thread he created - multilabel stratification python package.\n\n<pre>train_df, valid_df = train_test_split(train_df, \n                 test_size = 0.1,\n                 random_state=42,\n                  # hack to make stratification work                  \n                 stratify = train_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</pre>",
      "votes": null
    },
    {
      "id": "419015",
      "postDate": "11/11/2018 05:04:49",
      "content": "<p>I don't quite understand the last line. Can you explain it? Thanks</p>",
      "rawMarkdown": "I don't quite understand the last line. Can you explain it? Thanks",
      "votes": null
    },
    {
      "id": "419071",
      "postDate": "11/11/2018 07:28:07",
      "content": "<p>I'm not sure exactly how it work myself. I copied it from one of the early kernels that was posted. It looked like it worked ok and I used it for my top score.</p>",
      "rawMarkdown": "I'm not sure exactly how it work myself. I copied it from one of the early kernels that was posted. It looked like it worked ok and I used it for my top score.",
      "votes": null
    },
    {
      "id": "419087",
      "postDate": "11/11/2018 08:47:28",
      "content": "<p>You are stratifying only taking into account the label 27, rods and rings.</p>",
      "rawMarkdown": "You are stratifying only taking into account the label 27, rods and rings.",
      "votes": null
    },
    {
      "id": "419460",
      "postDate": "11/12/2018 02:38:15",
      "content": "<p>If I understood it right, sounds like it's based on string value of target, not actual targets....</p>\n\n<p>So <code>x[:3]</code> are the first 3 characters, which is quite strange. It may take two targets if there are two starting targets with length 1. It may take one target if the first target is length 2.</p>\n\n<p>If the stratifier considers strings, then it's also considering some sort of combination of targets, but not entirely. It's weird. And 27 was considered as the same as an image with only 0. </p>",
      "rawMarkdown": "If I understood it right, sounds like it's based on string value of target, not actual targets....\n\nSo `x[:3]` are the first 3 characters, which is quite strange. It may take two targets if there are two starting targets with length 1. It may take one target if the first target is length 2.\n\nIf the stratifier considers strings, then it's also considering some sort of combination of targets, but not entirely. It's weird. And 27 was considered as the same as an image with only 0.",
      "votes": null
    },
    {
      "id": "419480",
      "postDate": "11/12/2018 03:43:44",
      "content": "<p>I've switched to the implemention that <a href=\"/trentb\">@trentb</a> posted in another thread. Using his library, doing a 20% split the train/valid historgrams visually match. In this example I've dropped many 0's and 25's to even the data out a bit.</p>\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\nimport gc\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>\n\n<p><img src=\"http://brians.network/images/ttsplit.png\" alt=\"\"></p>",
      "rawMarkdown": "I've switched to the implemention that @trentb posted in another thread. Using his library, doing a 20% split the train/valid historgrams visually match. In this example I've dropped many 0's and 25's to even the data out a bit.\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\nimport gc\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>\n\n![](http://brians.network/images/ttsplit.png)",
      "votes": null
    },
    {
      "id": "441349",
      "postDate": "12/18/2018 15:15:32",
      "content": "<p>Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation? I'm already using stratified split and weighted sampling to try to account for the class imbalance.</p>",
      "rawMarkdown": "Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation? I'm already using stratified split and weighted sampling to try to account for the class imbalance.",
      "votes": null
    },
    {
      "id": "441380",
      "postDate": "12/18/2018 15:48:52",
      "content": "<p>I generally put more stock into my local validation once I know that it's working properly. Then try to find some trends between it and the public, then decide which scores deserve the most of my attention. For this one I think that I'll mostly focus on my local validation. </p>\n\n<p>At the end I usually submit one with my best local validation and one that has both good local and public scores.</p>",
      "rawMarkdown": "I generally put more stock into my local validation once I know that it's working properly. Then try to find some trends between it and the public, then decide which scores deserve the most of my attention. For this one I think that I'll mostly focus on my local validation. \n\nAt the end I usually submit one with my best local validation and one that has both good local and public scores.",
      "votes": null
    },
    {
      "id": "441455",
      "postDate": "12/18/2018 17:12:50",
      "content": "<blockquote>\n  <p>Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation?</p>\n</blockquote>\n\n<p>While I don't know all the details, what you describe sounds like a textbook definition of over-fitting. Unless you can control that in some way (dropout, regularization, different weights [less spread?], different network depth/width), I would suggest you move on to the next idea.</p>",
      "rawMarkdown": "&gt; Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation?\n\nWhile I don't know all the details, what you describe sounds like a textbook definition of over-fitting. Unless you can control that in some way (dropout, regularization, different weights [less spread?], different network depth/width), I would suggest you move on to the next idea.",
      "votes": null
    },
    {
      "id": "441808",
      "postDate": "12/19/2018 05:21:46",
      "content": "<p>What I did to see if my model is stable or not: <br>\nI select fixed thresholds for both <code>validation</code> and <code>submission</code> then compare the GAP between them. My best GAP is around 0.05. <br>\nI don't trust my local CV. Some models give a good local CV but a big GAP.  Some models I expected to be good but they dont. Some methods work well with a model but not with another. Deeper does not mean it is better. <br>\nWhat an interesting competition is !</p>",
      "rawMarkdown": "What I did to see if my model is stable or not:  \nI select fixed thresholds for both `validation` and `submission` then compare the GAP between them. My best GAP is around 0.05.  \nI don't trust my local CV. Some models give a good local CV but a big GAP.  Some models I expected to be good but they dont. Some methods work well with a model but not with another. Deeper does not mean it is better.   \nWhat an interesting competition is !",
      "votes": null
    },
    {
      "id": "441858",
      "postDate": "12/19/2018 07:08:32",
      "content": "<p>I agree with @Nguyen Xuan Bac : the CV-LB gap looks to me an important metric to discriminate between models. \nThe choice of the final submissions will be, at least for me, one of the most difficult among the competitions I have done so far</p>",
      "rawMarkdown": "I agree with @Nguyen Xuan Bac : the CV-LB gap looks to me an important metric to discriminate between models. \nThe choice of the final submissions will be, at least for me, one of the most difficult among the competitions I have done so far",
      "votes": null
    },
    {
      "id": "442522",
      "postDate": "12/20/2018 04:05:13",
      "content": "<p>@Tilii Indeed I was overfitting. I increased my regularization (more aggressive data augmentation) and found myself with a much better result!</p>",
      "rawMarkdown": "Tilii Indeed I was overfitting. I increased my regularization (more aggressive data augmentation) and found myself with a much better result!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 418902,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/10/2018 20:48:18",
      "content": "<p>This is what I use, it came from one of the starter kernels. It works well enough for me so far. I also plan to try the method that <a href=\"/trentb\">@trentb</a> posted in the thread he created - multilabel stratification python package.</p>\n\n<pre>train_df, valid_df = train_test_split(train_df, \n                 test_size = 0.1,\n                 random_state=42,\n                  # hack to make stratification work                  \n                 stratify = train_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</pre>",
      "votes": null,
      "replies": [
        {
          "id": 419015,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/11/2018 05:04:49",
          "content": "<p>I don't quite understand the last line. Can you explain it? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419071,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/11/2018 07:28:07",
          "content": "<p>I'm not sure exactly how it work myself. I copied it from one of the early kernels that was posted. It looked like it worked ok and I used it for my top score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419087,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "11/11/2018 08:47:28",
          "content": "<p>You are stratifying only taking into account the label 27, rods and rings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419460,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/12/2018 02:38:15",
          "content": "<p>If I understood it right, sounds like it's based on string value of target, not actual targets....</p>\n\n<p>So <code>x[:3]</code> are the first 3 characters, which is quite strange. It may take two targets if there are two starting targets with length 1. It may take one target if the first target is length 2.</p>\n\n<p>If the stratifier considers strings, then it's also considering some sort of combination of targets, but not entirely. It's weird. And 27 was considered as the same as an image with only 0. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419480,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/12/2018 03:43:44",
          "content": "<p>I've switched to the implemention that <a href=\"/trentb\">@trentb</a> posted in another thread. Using his library, doing a 20% split the train/valid historgrams visually match. In this example I've dropped many 0's and 25's to even the data out a bit.</p>\n\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\nimport gc\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>\n\n<p><img src=\"http://brians.network/images/ttsplit.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441349,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "12/18/2018 15:15:32",
      "content": "<p>Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation? I'm already using stratified split and weighted sampling to try to account for the class imbalance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441380,
          "author_name": "florianm",
          "author_url": "",
          "post_date": "12/18/2018 15:48:52",
          "content": "<p>I generally put more stock into my local validation once I know that it's working properly. Then try to find some trends between it and the public, then decide which scores deserve the most of my attention. For this one I think that I'll mostly focus on my local validation. </p>\n\n<p>At the end I usually submit one with my best local validation and one that has both good local and public scores.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441455,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "12/18/2018 17:12:50",
          "content": "<blockquote>\n  <p>Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation?</p>\n</blockquote>\n\n<p>While I don't know all the details, what you describe sounds like a textbook definition of over-fitting. Unless you can control that in some way (dropout, regularization, different weights [less spread?], different network depth/width), I would suggest you move on to the next idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441808,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "12/19/2018 05:21:46",
          "content": "<p>What I did to see if my model is stable or not: <br>\nI select fixed thresholds for both <code>validation</code> and <code>submission</code> then compare the GAP between them. My best GAP is around 0.05. <br>\nI don't trust my local CV. Some models give a good local CV but a big GAP.  Some models I expected to be good but they dont. Some methods work well with a model but not with another. Deeper does not mean it is better. <br>\nWhat an interesting competition is !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441858,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "12/19/2018 07:08:32",
          "content": "<p>I agree with @Nguyen Xuan Bac : the CV-LB gap looks to me an important metric to discriminate between models. \nThe choice of the final submissions will be, at least for me, one of the most difficult among the competitions I have done so far</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442522,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "12/20/2018 04:05:13",
          "content": "<p>@Tilii Indeed I was overfitting. I increased my regularization (more aggressive data augmentation) and found myself with a much better result!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "418792": "If the train distribution is different from the public LB distribution, and the private LB distribution is likely different from public, what validation strategy can we use to be confident that one model is better than another? Right now I'm just doing a random validation split, I saw some people mention they're trying a stratified split, was wondering if anyone had thoughts on other possible approaches.",
    "418902": "This is what I use, it came from one of the starter kernels. It works well enough for me so far. I also plan to try the method that @trentb posted in the thread he created - multilabel stratification python package.\n\n<pre>train_df, valid_df = train_test_split(train_df, \n                 test_size = 0.1,\n                 random_state=42,\n                  # hack to make stratification work                  \n                 stratify = train_df['Target'].map(lambda x: x[:3] if '27' not in x else '0'))\n</pre>",
    "419015": "I don't quite understand the last line. Can you explain it? Thanks",
    "419071": "I'm not sure exactly how it work myself. I copied it from one of the early kernels that was posted. It looked like it worked ok and I used it for my top score.",
    "419087": "You are stratifying only taking into account the label 27, rods and rings.",
    "419460": "If I understood it right, sounds like it's based on string value of target, not actual targets....\n\nSo `x[:3]` are the first 3 characters, which is quite strange. It may take two targets if there are two starting targets with length 1. It may take one target if the first target is length 2.\n\nIf the stratifier considers strings, then it's also considering some sort of combination of targets, but not entirely. It's weird. And 27 was considered as the same as an image with only 0.",
    "419480": "I've switched to the implemention that @trentb posted in another thread. Using his library, doing a 20% split the train/valid historgrams visually match. In this example I've dropped many 0's and 25's to even the data out a bit.\n<pre>from iterstrat.ml_stratifiers import MultilabelStratifiedShuffleSplit\nimport numpy as np\nimport gc\n\nmsss = MultilabelStratifiedShuffleSplit(n_splits=1, test_size=0.2, random_state=42)\ntrain_df_orig = train_df.copy()\nX = train_df_orig['Id'].tolist()\ny = train_df_orig['target_vec_float'].tolist()\n\nfor train_index, test_index in msss.split(X,y): #it should only do one iteration\n    print(\"TRAIN:\", train_index, \"TEST:\", test_index)\n    train_df = train_df_orig.loc[train_df_orig.index.intersection(train_index)].copy()\n    valid_df = train_df_orig.loc[train_df_orig.index.intersection(test_index)].copy()\ngc.collect()\n</pre>\n\n![](http://brians.network/images/ttsplit.png)",
    "441349": "Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation? I'm already using stratified split and weighted sampling to try to account for the class imbalance.",
    "441380": "I generally put more stock into my local validation once I know that it's working properly. Then try to find some trends between it and the public, then decide which scores deserve the most of my attention. For this one I think that I'll mostly focus on my local validation. \n\nAt the end I usually submit one with my best local validation and one that has both good local and public scores.",
    "441455": "&gt; Coming back to this--I just trained a model that was 0.15 better than my previous best on local validation, but did 0.30 worse on the public LB. What are people doing in this situation?\n\nWhile I don't know all the details, what you describe sounds like a textbook definition of over-fitting. Unless you can control that in some way (dropout, regularization, different weights [less spread?], different network depth/width), I would suggest you move on to the next idea.",
    "441808": "What I did to see if my model is stable or not:  \nI select fixed thresholds for both `validation` and `submission` then compare the GAP between them. My best GAP is around 0.05.  \nI don't trust my local CV. Some models give a good local CV but a big GAP.  Some models I expected to be good but they dont. Some methods work well with a model but not with another. Deeper does not mean it is better.   \nWhat an interesting competition is !",
    "441858": "I agree with @Nguyen Xuan Bac : the CV-LB gap looks to me an important metric to discriminate between models. \nThe choice of the final submissions will be, at least for me, one of the most difficult among the competitions I have done so far",
    "442522": "Tilii Indeed I was overfitting. I increased my regularization (more aggressive data augmentation) and found myself with a much better result!"
  },
  "source": "meta"
}