{
  "id": 285546,
  "title": "Best way to split folds for cross-validation",
  "url": "/competitions/sartorius-cell-instance-segmentation/discussion/285546",
  "author_name": "",
  "post_date": "2021-11-05T06:29:56.002920500Z",
  "votes": 55,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I noticed none of the public notebooks are using cross-validation so I decided to share my approach.</p>\n<p>Data should be split into folds on image level, not annotation level. I took a single row for every image from the original training dataframe.</p>\n<p><code>df_images = df.groupby('id').first().reset_index()</code></p>\n<p>then I do stratified splits on cell type…</p>\n<pre><code>skf = StratifiedKFold(n_splits=n_splits, shuffle=shuffle, random_state=random_state)\nfor fold, (_, val_idx) in enumerate(skf.split(X=df_images, y=df_images['cell_type']), 1):\n    df_images.loc[val_idx, 'fold'] = fold\ndf_images['fold'] = df_images['fold'].astype(np.uint8)\n</code></pre>\n<p>then I save folds and id columns to data directory.</p>\n<p><code>df_images[['id', 'fold']].to_csv(f'{settings.DATA_PATH}/train_folds.csv', index=False)</code></p>\n<p>Saved file consumes less disk space with this way. When training starts, I merge train_folds.csv with train.csv like this</p>\n<pre><code>df_train = pd.read_csv(f'{settings.DATA_PATH}/train.csv')\ndf_train_folds = pd.read_csv(f'{settings.DATA_PATH}/train_folds.csv')\ndf_train = df_train.merge(df_train_folds, how='left', on='id')\n</code></pre>\n<p>This is the distribution of cell types in folds.</p>\n<pre><code>Training set images are split into 5 stratified folds\nFold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n</code></pre>",
  "messages": [
    {
      "id": "1571770",
      "postDate": "11/05/2021 06:29:56",
      "content": "<p>I noticed none of the public notebooks are using cross-validation so I decided to share my approach.</p>\n<p>Data should be split into folds on image level, not annotation level. I took a single row for every image from the original training dataframe.</p>\n<p><code>df_images = df.groupby('id').first().reset_index()</code></p>\n<p>then I do stratified splits on cell type…</p>\n<pre><code>skf = StratifiedKFold(n_splits=n_splits, shuffle=shuffle, random_state=random_state)\nfor fold, (_, val_idx) in enumerate(skf.split(X=df_images, y=df_images['cell_type']), 1):\n    df_images.loc[val_idx, 'fold'] = fold\ndf_images['fold'] = df_images['fold'].astype(np.uint8)\n</code></pre>\n<p>then I save folds and id columns to data directory.</p>\n<p><code>df_images[['id', 'fold']].to_csv(f'{settings.DATA_PATH}/train_folds.csv', index=False)</code></p>\n<p>Saved file consumes less disk space with this way. When training starts, I merge train_folds.csv with train.csv like this</p>\n<pre><code>df_train = pd.read_csv(f'{settings.DATA_PATH}/train.csv')\ndf_train_folds = pd.read_csv(f'{settings.DATA_PATH}/train_folds.csv')\ndf_train = df_train.merge(df_train_folds, how='left', on='id')\n</code></pre>\n<p>This is the distribution of cell types in folds.</p>\n<pre><code>Training set images are split into 5 stratified folds\nFold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n</code></pre>",
      "rawMarkdown": "I noticed none of the public notebooks are using cross-validation so I decided to share my approach.\n\nData should be split into folds on image level, not annotation level. I took a single row for every image from the original training dataframe.\n\n`df_images = df.groupby('id').first().reset_index()`\n\nthen I do stratified splits on cell type...\n\n```\nskf = StratifiedKFold(n_splits=n_splits, shuffle=shuffle, random_state=random_state)\nfor fold, (_, val_idx) in enumerate(skf.split(X=df_images, y=df_images['cell_type']), 1):\n    df_images.loc[val_idx, 'fold'] = fold\ndf_images['fold'] = df_images['fold'].astype(np.uint8)\n```\n\nthen I save folds and id columns to data directory.\n\n`df_images[['id', 'fold']].to_csv(f'{settings.DATA_PATH}/train_folds.csv', index=False)`\n\nSaved file consumes less disk space with this way. When training starts, I merge train_folds.csv with train.csv like this\n\n```\ndf_train = pd.read_csv(f'{settings.DATA_PATH}/train.csv')\ndf_train_folds = pd.read_csv(f'{settings.DATA_PATH}/train_folds.csv')\ndf_train = df_train.merge(df_train_folds, how='left', on='id')\n```\n\nThis is the distribution of cell types in folds.\n\n```\nTraining set images are split into 5 stratified folds\nFold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n```",
      "votes": null
    },
    {
      "id": "1571928",
      "postDate": "11/05/2021 09:29:43",
      "content": "<p>Nice mate, i have the approach as well. But as i see, i grouped at cell level and img lvl. Do you think this is correct ? Here is the distribution per fold as well.</p>\n<pre><code>n_folds = 5\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\ncell_types = list(df_base['cell_type'].unique())\ndf_images = df_base.groupby([\"id\", \"cell_type\"]).agg({'annotation': 'count'}).sort_values(\"annotation\", ascending=False).reset_index()\n\nfor fold_num, (trn_index, val_index) in enumerate(skf.split(df_images, df_images['cell_type'])):\n    trn_df, val_df = df_images.iloc[trn_index], df_images.iloc[val_index]\n\n    df_train = df_base[df_base['id'].isin(trn_df['id'])]\n    df_val = df_base[df_base['id'].isin(val_df['id'])]\n\n    per_fold_train, per_fold_val = {}, {}\n\n    for cell in cell_types:\n        per_fold_train[cell] = sum(trn_df['cell_type'] == cell)\n        per_fold_val[cell] = sum(val_df['cell_type'] == cell)\n\n    total_tr, total_val = sum(per_fold_train.values()), sum(per_fold_val.values())\n    total = total_tr + total_val\n\n    ds_train = CellDataset(config.TRAIN_PATH, df_train, resize=False, transforms=get_transform(train=True))\n    ds_val = CellDataset(config.TRAIN_PATH, df_val, resize=False, transforms=get_transform(train=False))\n\n    print(f\"\\n---------Fold {fold_num+1}---------\")\n    print(f\"Images in train set:           {len(trn_df)}\")\n    print(f\"Annotations in train set:      {len(df_train)}\")\n    print(f\"Images in validation set:      {len(val_df)}\")\n    print(f\"Annotations in validation set: {len(df_val)}\")\n\n---------Fold 1---------\nImages in train set:           484\nAnnotations in train set:      58130\nImages in validation set:      122\nAnnotations in validation set: 15455\n</code></pre>",
      "rawMarkdown": "Nice mate, i have the approach as well. But as i see, i grouped at cell level and img lvl. Do you think this is correct ? Here is the distribution per fold as well.\n\n```\nn_folds = 5\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\ncell_types = list(df_base['cell_type'].unique())\ndf_images = df_base.groupby([\"id\", \"cell_type\"]).agg({'annotation': 'count'}).sort_values(\"annotation\", ascending=False).reset_index()\n\nfor fold_num, (trn_index, val_index) in enumerate(skf.split(df_images, df_images['cell_type'])):\n    trn_df, val_df = df_images.iloc[trn_index], df_images.iloc[val_index]\n    \n    df_train = df_base[df_base['id'].isin(trn_df['id'])]\n    df_val = df_base[df_base['id'].isin(val_df['id'])]\n    \n    per_fold_train, per_fold_val = {}, {}\n    \n    for cell in cell_types:\n        per_fold_train[cell] = sum(trn_df['cell_type'] == cell)\n        per_fold_val[cell] = sum(val_df['cell_type'] == cell)\n        \n    total_tr, total_val = sum(per_fold_train.values()), sum(per_fold_val.values())\n    total = total_tr + total_val\n    \n    ds_train = CellDataset(config.TRAIN_PATH, df_train, resize=False, transforms=get_transform(train=True))\n    ds_val = CellDataset(config.TRAIN_PATH, df_val, resize=False, transforms=get_transform(train=False))\n    \n    print(f\"\\n---------Fold {fold_num+1}---------\")\n    print(f\"Images in train set:           {len(trn_df)}\")\n    print(f\"Annotations in train set:      {len(df_train)}\")\n    print(f\"Images in validation set:      {len(val_df)}\")\n    print(f\"Annotations in validation set: {len(df_val)}\")\n\n---------Fold 1---------\nImages in train set:           484\nAnnotations in train set:      58130\nImages in validation set:      122\nAnnotations in validation set: 15455\n\n```",
      "votes": null
    },
    {
      "id": "1573258",
      "postDate": "11/06/2021 12:48:12",
      "content": "<p>I think it's also good to compare the number of masked pixels for each cell type at each fold. If the deviation between the folds is big, repeated cross validation usually gives a big boost.</p>",
      "rawMarkdown": "I think it's also good to compare the number of masked pixels for each cell type at each fold. If the deviation between the folds is big, repeated cross validation usually gives a big boost.",
      "votes": null
    },
    {
      "id": "1575189",
      "postDate": "11/08/2021 08:17:48",
      "content": "<p>You are right but I think splitting them by their cell type should make the number of annotated objects and pixels consisted between folds because those two things are dependent to cell type.</p>",
      "rawMarkdown": "You are right but I think splitting them by their cell type should make the number of annotated objects and pixels consisted between folds because those two things are dependent to cell type.",
      "votes": null
    },
    {
      "id": "1575560",
      "postDate": "11/08/2021 13:23:16",
      "content": "<p>I guess grouping by sample_id is another approach.</p>",
      "rawMarkdown": "I guess grouping by sample_id is another approach.",
      "votes": null
    },
    {
      "id": "1578412",
      "postDate": "11/11/2021 02:05:28",
      "content": "<p>I found that sample id's align with id, so it won't change the results anyhow. </p>",
      "rawMarkdown": "I found that sample id's align with id, so it won't change the results anyhow.",
      "votes": null
    },
    {
      "id": "1586254",
      "postDate": "11/17/2021 23:52:12",
      "content": "<p>When split by the cell type, the number of instances per fold is as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Fold / cell_type</th>\n<th>0</th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>astro</td>\n<td>2122</td>\n<td>2023</td>\n<td>2064</td>\n<td>2316</td>\n<td>1997</td>\n</tr>\n<tr>\n<td>cort</td>\n<td>2197</td>\n<td>2262</td>\n<td>2272</td>\n<td>2053</td>\n<td>2011</td>\n</tr>\n<tr>\n<td>shsy5y</td>\n<td>9822</td>\n<td>10340</td>\n<td>10868</td>\n<td>10147</td>\n<td>11109</td>\n</tr>\n</tbody>\n</table>\n<p>The per fold deviation is 6, 5.5, 5 percent, respectively. I'm not suggesting anything, just reporting the numbers.</p>",
      "rawMarkdown": "When split by the cell type, the number of instances per fold is as follows:\n\n| Fold / cell_type  | 0 | 1 | 2 | 3 | 4 | \n| --- | --- | --- |\n|  astro |  2122 | 2023 | 2064 | 2316 | 1997 |\n| cort | 2197 | 2262 | 2272 | 2053 | 2011 |\n| shsy5y | 9822 | 10340 | 10868 | 10147 | 11109\n\nThe per fold deviation is 6, 5.5, 5 percent, respectively. I'm not suggesting anything, just reporting the numbers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1571928,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "11/05/2021 09:29:43",
      "content": "<p>Nice mate, i have the approach as well. But as i see, i grouped at cell level and img lvl. Do you think this is correct ? Here is the distribution per fold as well.</p>\n<pre><code>n_folds = 5\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\ncell_types = list(df_base['cell_type'].unique())\ndf_images = df_base.groupby([\"id\", \"cell_type\"]).agg({'annotation': 'count'}).sort_values(\"annotation\", ascending=False).reset_index()\n\nfor fold_num, (trn_index, val_index) in enumerate(skf.split(df_images, df_images['cell_type'])):\n    trn_df, val_df = df_images.iloc[trn_index], df_images.iloc[val_index]\n\n    df_train = df_base[df_base['id'].isin(trn_df['id'])]\n    df_val = df_base[df_base['id'].isin(val_df['id'])]\n\n    per_fold_train, per_fold_val = {}, {}\n\n    for cell in cell_types:\n        per_fold_train[cell] = sum(trn_df['cell_type'] == cell)\n        per_fold_val[cell] = sum(val_df['cell_type'] == cell)\n\n    total_tr, total_val = sum(per_fold_train.values()), sum(per_fold_val.values())\n    total = total_tr + total_val\n\n    ds_train = CellDataset(config.TRAIN_PATH, df_train, resize=False, transforms=get_transform(train=True))\n    ds_val = CellDataset(config.TRAIN_PATH, df_val, resize=False, transforms=get_transform(train=False))\n\n    print(f\"\\n---------Fold {fold_num+1}---------\")\n    print(f\"Images in train set:           {len(trn_df)}\")\n    print(f\"Annotations in train set:      {len(df_train)}\")\n    print(f\"Images in validation set:      {len(val_df)}\")\n    print(f\"Annotations in validation set: {len(df_val)}\")\n\n---------Fold 1---------\nImages in train set:           484\nAnnotations in train set:      58130\nImages in validation set:      122\nAnnotations in validation set: 15455\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1573258,
      "author_name": "tolgadincer",
      "author_url": "",
      "post_date": "11/06/2021 12:48:12",
      "content": "<p>I think it's also good to compare the number of masked pixels for each cell type at each fold. If the deviation between the folds is big, repeated cross validation usually gives a big boost.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1575189,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/08/2021 08:17:48",
          "content": "<p>You are right but I think splitting them by their cell type should make the number of annotated objects and pixels consisted between folds because those two things are dependent to cell type.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1586254,
          "author_name": "tolgadincer",
          "author_url": "",
          "post_date": "11/17/2021 23:52:12",
          "content": "<p>When split by the cell type, the number of instances per fold is as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Fold / cell_type</th>\n<th>0</th>\n<th>1</th>\n<th>2</th>\n<th>3</th>\n<th>4</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>astro</td>\n<td>2122</td>\n<td>2023</td>\n<td>2064</td>\n<td>2316</td>\n<td>1997</td>\n</tr>\n<tr>\n<td>cort</td>\n<td>2197</td>\n<td>2262</td>\n<td>2272</td>\n<td>2053</td>\n<td>2011</td>\n</tr>\n<tr>\n<td>shsy5y</td>\n<td>9822</td>\n<td>10340</td>\n<td>10868</td>\n<td>10147</td>\n<td>11109</td>\n</tr>\n</tbody>\n</table>\n<p>The per fold deviation is 6, 5.5, 5 percent, respectively. I'm not suggesting anything, just reporting the numbers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1575560,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "11/08/2021 13:23:16",
      "content": "<p>I guess grouping by sample_id is another approach.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1578412,
          "author_name": "phantnguyen",
          "author_url": "",
          "post_date": "11/11/2021 02:05:28",
          "content": "<p>I found that sample id's align with id, so it won't change the results anyhow. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1571770": "I noticed none of the public notebooks are using cross-validation so I decided to share my approach.\n\nData should be split into folds on image level, not annotation level. I took a single row for every image from the original training dataframe.\n\n`df_images = df.groupby('id').first().reset_index()`\n\nthen I do stratified splits on cell type...\n\n```\nskf = StratifiedKFold(n_splits=n_splits, shuffle=shuffle, random_state=random_state)\nfor fold, (_, val_idx) in enumerate(skf.split(X=df_images, y=df_images['cell_type']), 1):\n    df_images.loc[val_idx, 'fold'] = fold\ndf_images['fold'] = df_images['fold'].astype(np.uint8)\n```\n\nthen I save folds and id columns to data directory.\n\n`df_images[['id', 'fold']].to_csv(f'{settings.DATA_PATH}/train_folds.csv', index=False)`\n\nSaved file consumes less disk space with this way. When training starts, I merge train_folds.csv with train.csv like this\n\n```\ndf_train = pd.read_csv(f'{settings.DATA_PATH}/train.csv')\ndf_train_folds = pd.read_csv(f'{settings.DATA_PATH}/train_folds.csv')\ndf_train = df_train.merge(df_train_folds, how='left', on='id')\n```\n\nThis is the distribution of cell types in folds.\n\n```\nTraining set images are split into 5 stratified folds\nFold 1 (122, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 27}\nFold 2 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 3 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 4 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\nFold 5 (121, 10) - {'cort': 64, 'shsy5y': 31, 'astro': 26}\n```",
    "1571928": "Nice mate, i have the approach as well. But as i see, i grouped at cell level and img lvl. Do you think this is correct ? Here is the distribution per fold as well.\n\n```\nn_folds = 5\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\ncell_types = list(df_base['cell_type'].unique())\ndf_images = df_base.groupby([\"id\", \"cell_type\"]).agg({'annotation': 'count'}).sort_values(\"annotation\", ascending=False).reset_index()\n\nfor fold_num, (trn_index, val_index) in enumerate(skf.split(df_images, df_images['cell_type'])):\n    trn_df, val_df = df_images.iloc[trn_index], df_images.iloc[val_index]\n    \n    df_train = df_base[df_base['id'].isin(trn_df['id'])]\n    df_val = df_base[df_base['id'].isin(val_df['id'])]\n    \n    per_fold_train, per_fold_val = {}, {}\n    \n    for cell in cell_types:\n        per_fold_train[cell] = sum(trn_df['cell_type'] == cell)\n        per_fold_val[cell] = sum(val_df['cell_type'] == cell)\n        \n    total_tr, total_val = sum(per_fold_train.values()), sum(per_fold_val.values())\n    total = total_tr + total_val\n    \n    ds_train = CellDataset(config.TRAIN_PATH, df_train, resize=False, transforms=get_transform(train=True))\n    ds_val = CellDataset(config.TRAIN_PATH, df_val, resize=False, transforms=get_transform(train=False))\n    \n    print(f\"\\n---------Fold {fold_num+1}---------\")\n    print(f\"Images in train set:           {len(trn_df)}\")\n    print(f\"Annotations in train set:      {len(df_train)}\")\n    print(f\"Images in validation set:      {len(val_df)}\")\n    print(f\"Annotations in validation set: {len(df_val)}\")\n\n---------Fold 1---------\nImages in train set:           484\nAnnotations in train set:      58130\nImages in validation set:      122\nAnnotations in validation set: 15455\n\n```",
    "1573258": "I think it's also good to compare the number of masked pixels for each cell type at each fold. If the deviation between the folds is big, repeated cross validation usually gives a big boost.",
    "1575189": "You are right but I think splitting them by their cell type should make the number of annotated objects and pixels consisted between folds because those two things are dependent to cell type.",
    "1575560": "I guess grouping by sample_id is another approach.",
    "1578412": "I found that sample id's align with id, so it won't change the results anyhow.",
    "1586254": "When split by the cell type, the number of instances per fold is as follows:\n\n| Fold / cell_type  | 0 | 1 | 2 | 3 | 4 | \n| --- | --- | --- |\n|  astro |  2122 | 2023 | 2064 | 2316 | 1997 |\n| cort | 2197 | 2262 | 2272 | 2053 | 2011 |\n| shsy5y | 9822 | 10340 | 10868 | 10147 | 11109\n\nThe per fold deviation is 6, 5.5, 5 percent, respectively. I'm not suggesting anything, just reporting the numbers."
  },
  "source": "meta"
}