{
  "id": 111731,
  "title": "Beware of Pandas value_counts method for validation split",
  "url": "/competitions/understanding_cloud_organization/discussion/111731",
  "author_name": "",
  "post_date": "2019-10-08T13:21:39.452422300Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In several notebooks, (<a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">here</a>, <a href=\"https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch\">here</a>, <a href=\"https://www.kaggle.com/hung96ad/fine-turning-model-with-adabound\">here</a>), the validation dataset is constructed using the following dataframe:\n<code>id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\\nreset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'})</code></p>\n\n<p>I used a very similar code snippet to set up k-fold cross validation and had major issues with it yesterday.\nIt turns out that <code>value_counts</code> doesn't handle ties in a deterministic way (see e.g. <a href=\"https://github.com/pandas-dev/pandas/issues/15833\">here</a>), meaning that the validation set differs from run to run.</p>",
  "messages": [
    {
      "id": "644191",
      "postDate": "10/08/2019 13:21:39",
      "content": "<p>In several notebooks, (<a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">here</a>, <a href=\"https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch\">here</a>, <a href=\"https://www.kaggle.com/hung96ad/fine-turning-model-with-adabound\">here</a>), the validation dataset is constructed using the following dataframe:\n<code>id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\\nreset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'})</code></p>\n\n<p>I used a very similar code snippet to set up k-fold cross validation and had major issues with it yesterday.\nIt turns out that <code>value_counts</code> doesn't handle ties in a deterministic way (see e.g. <a href=\"https://github.com/pandas-dev/pandas/issues/15833\">here</a>), meaning that the validation set differs from run to run.</p>",
      "rawMarkdown": "In several notebooks, ([here](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools), [here](https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch), [here](https://www.kaggle.com/hung96ad/fine-turning-model-with-adabound)), the validation dataset is constructed using the following dataframe:\n`id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\\nreset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'})`\n\nI used a very similar code snippet to set up k-fold cross validation and had major issues with it yesterday.\nIt turns out that `value_counts` doesn't handle ties in a deterministic way (see e.g. [here](https://github.com/pandas-dev/pandas/issues/15833)), meaning that the validation set differs from run to run.",
      "votes": null
    },
    {
      "id": "644705",
      "postDate": "10/09/2019 07:56:42",
      "content": "<p>Then you could just add a <code>sort_value</code> to sort by <code>count</code> and <code>ìmg_id</code> in the second place to fix this:</p>\n\n<p><code>id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\ reset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'}).sort_values(['count', 'img_id'])</code></p>",
      "rawMarkdown": "Then you could just add a `sort_value` to sort by `count` and `ìmg_id` in the second place to fix this:\n\n`id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\ reset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'}).sort_values(['count', 'img_id'])`",
      "votes": null
    },
    {
      "id": "645248",
      "postDate": "10/09/2019 23:43:52",
      "content": "<p>Thanks so much for this notice Welinder. This helps me.</p>",
      "rawMarkdown": "Thanks so much for this notice Welinder. This helps me.",
      "votes": null
    },
    {
      "id": "645351",
      "postDate": "10/10/2019 03:07:49",
      "content": "<p>Someone find a good validation strategy? I can't make my val score stay close to LB. Now I'm getting 0.581/0.626(LB)</p>",
      "rawMarkdown": "Someone find a good validation strategy? I can't make my val score stay close to LB. Now I'm getting 0.581/0.626(LB)",
      "votes": null
    },
    {
      "id": "645372",
      "postDate": "10/10/2019 03:54:08",
      "content": "<p>Your validation score may be calculated on smoothed dice (2*(A intersect B)+ eps)/(|A| + |B| + eps) when training, not competition dice. You need to add the original dice in your validation strategy, because for both empty ground truth and empty prediction the dice score is undefined. Then LB score will stay close to validation score.</p>",
      "rawMarkdown": "Your validation score may be calculated on smoothed dice (2*(A intersect B)+ eps)/(|A| + |B| + eps) when training, not competition dice. You need to add the original dice in your validation strategy, because for both empty ground truth and empty prediction the dice score is undefined. Then LB score will stay close to validation score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 644705,
      "author_name": "jottbe",
      "author_url": "",
      "post_date": "10/09/2019 07:56:42",
      "content": "<p>Then you could just add a <code>sort_value</code> to sort by <code>count</code> and <code>ìmg_id</code> in the second place to fix this:</p>\n\n<p><code>id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\ reset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'}).sort_values(['count', 'img_id'])</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 645248,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "10/09/2019 23:43:52",
      "content": "<p>Thanks so much for this notice Welinder. This helps me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 645351,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "10/10/2019 03:07:49",
      "content": "<p>Someone find a good validation strategy? I can't make my val score stay close to LB. Now I'm getting 0.581/0.626(LB)</p>",
      "votes": null,
      "replies": [
        {
          "id": 645372,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/10/2019 03:54:08",
          "content": "<p>Your validation score may be calculated on smoothed dice (2*(A intersect B)+ eps)/(|A| + |B| + eps) when training, not competition dice. You need to add the original dice in your validation strategy, because for both empty ground truth and empty prediction the dice score is undefined. Then LB score will stay close to validation score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "644191": "In several notebooks, ([here](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools), [here](https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch), [here](https://www.kaggle.com/hung96ad/fine-turning-model-with-adabound)), the validation dataset is constructed using the following dataframe:\n`id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\\nreset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'})`\n\nI used a very similar code snippet to set up k-fold cross validation and had major issues with it yesterday.\nIt turns out that `value_counts` doesn't handle ties in a deterministic way (see e.g. [here](https://github.com/pandas-dev/pandas/issues/15833)), meaning that the validation set differs from run to run.",
    "644705": "Then you could just add a `sort_value` to sort by `count` and `ìmg_id` in the second place to fix this:\n\n`id_mask_count = train.loc[train['EncodedPixels'].isnull() == False, 'Image_Label'].apply(lambda x: x.split('_')[0]).value_counts().\\ reset_index().rename(columns={'index': 'img_id', 'Image_Label': 'count'}).sort_values(['count', 'img_id'])`",
    "645248": "Thanks so much for this notice Welinder. This helps me.",
    "645351": "Someone find a good validation strategy? I can't make my val score stay close to LB. Now I'm getting 0.581/0.626(LB)",
    "645372": "Your validation score may be calculated on smoothed dice (2*(A intersect B)+ eps)/(|A| + |B| + eps) when training, not competition dice. You need to add the original dice in your validation strategy, because for both empty ground truth and empty prediction the dice score is undefined. Then LB score will stay close to validation score."
  },
  "source": "meta"
}