{
  "id": 100090,
  "title": "Understanding Test & (Collecting) Validation Strategies",
  "url": "/competitions/recursion-cellular-image-classification/discussion/100090",
  "author_name": "",
  "post_date": "2019-07-16T14:57:14.662994700Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<ol>\n<li><p>I wonder why the length of <code>test</code> folder is twice the length of <code>test.csv</code>?\nAnswer by Phuc Le: Each well has 2 images from 2 sites (2 different patches)</p></li>\n<li><p>Also, what are some of the validation strategies people are trying?</p></li>\n</ol>\n\n<p>Thank you.</p>\n\n<p>Edit: Follow up on Q2: How to make use of <code>train_controls.csv</code> and <code>test_controls.csv</code>? </p>",
  "messages": [
    {
      "id": "577268",
      "postDate": "07/16/2019 14:57:14",
      "content": "<p>Hi guys,</p>\n\n<ol>\n<li><p>I wonder why the length of <code>test</code> folder is twice the length of <code>test.csv</code>?\nAnswer by Phuc Le: Each well has 2 images from 2 sites (2 different patches)</p></li>\n<li><p>Also, what are some of the validation strategies people are trying?</p></li>\n</ol>\n\n<p>Thank you.</p>\n\n<p>Edit: Follow up on Q2: How to make use of <code>train_controls.csv</code> and <code>test_controls.csv</code>? </p>",
      "rawMarkdown": "Hi guys,\n\n1. I wonder why the length of `test` folder is twice the length of `test.csv`?\nAnswer by Phuc Le: Each well has 2 images from 2 sites (2 different patches)\n\n2. Also, what are some of the validation strategies people are trying?\n\nThank you.\n\nEdit: Follow up on Q2: How to make use of `train_controls.csv` and `test_controls.csv`?",
      "votes": null
    },
    {
      "id": "577280",
      "postDate": "07/16/2019 15:08:12",
      "content": "<ol>\n<li>Each well has 2 images from 2 sites (2 different patches)</li>\n<li>I stratified by cell types and group by experiments.</li>\n</ol>",
      "rawMarkdown": "1. Each well has 2 images from 2 sites (2 different patches)\n2. I stratified by cell types and group by experiments.",
      "votes": null
    },
    {
      "id": "577295",
      "postDate": "07/16/2019 15:19:58",
      "content": "<p>thanks, Phuc! Sorry this might be a noob question. But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this? </p>\n\n<p>also, what does it mean by stratified by <code>cell_types</code> and group by <code>experiments</code>?</p>\n\n<p>say an instance is <code>HEPG2-01</code>, we have <code>cell_type</code> of <code>HEPG2</code> and <code>experiment</code> of <code>01</code>. Since we have 51 instances, do we validate by according to the <code>24 in HUVEC, 11 in RPE, 11 in HepG2, and 5 in U2OS</code> introduced? Say having 20 HUVEC in <code>train</code>, 4 HUVEC in <code>test</code>, we then split 20 HUVEC into 80:20 ratio into training and validation set?</p>\n\n<p>Thanks, once again! </p>",
      "rawMarkdown": "thanks, Phuc! Sorry this might be a noob question. But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this? \n\nalso, what does it mean by stratified by `cell_types` and group by `experiments`?\n\nsay an instance is `HEPG2-01`, we have `cell_type` of `HEPG2` and `experiment` of `01`. Since we have 51 instances, do we validate by according to the `24 in HUVEC, 11 in RPE, 11 in HepG2, and 5 in U2OS` introduced? Say having 20 HUVEC in `train`, 4 HUVEC in `test`, we then split 20 HUVEC into 80:20 ratio into training and validation set?\n\nThanks, once again!",
      "votes": null
    },
    {
      "id": "577391",
      "postDate": "07/16/2019 16:43:51",
      "content": "<p><code>\ndef stratified_groups_kfold(df, n_splits=5, random_state=0):\n    df['cell'] = df['experiment']\n    all_exps = pd.Series(df['experiment'].unique())\n    all_cells = all_exps.apply(lambda x: x.split('-')[0])\n    folds = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n    for idx_tr, idx_val in folds.split(all_cells, all_cells):\n        exps_tr = all_exps[idx_tr]\n        exps_val = all_exps[idx_val]\n        idx_tr_new = df[df['experiment'].isin(exps_tr)].index.to_numpy()\n        idx_val_new = df[df['experiment'].isin(exps_val)].index.to_numpy()\n        yield idx_tr_new, idx_val_new\n</code>\nHere how I stratified by all cells and group by experiments</p>",
      "rawMarkdown": "```\ndef stratified_groups_kfold(df, n_splits=5, random_state=0):\n    df['cell'] = df['experiment']\n    all_exps = pd.Series(df['experiment'].unique())\n    all_cells = all_exps.apply(lambda x: x.split('-')[0])\n    folds = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n    for idx_tr, idx_val in folds.split(all_cells, all_cells):\n        exps_tr = all_exps[idx_tr]\n        exps_val = all_exps[idx_val]\n        idx_tr_new = df[df['experiment'].isin(exps_tr)].index.to_numpy()\n        idx_val_new = df[df['experiment'].isin(exps_val)].index.to_numpy()\n        yield idx_tr_new, idx_val_new\n```\nHere how I stratified by all cells and group by experiments",
      "votes": null
    },
    {
      "id": "577404",
      "postDate": "07/16/2019 16:50:01",
      "content": "<p>&gt; But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this?</p>\n\n<p>Very likely. The better consistency between 2 sites, the better the models.\nSimplest way to deal with this is just average the prediction from 2 sites. -&gt; My PB gains 3% from doing this, likely more if I have stronger model.</p>\n\n<p>I think how you deal with the consistency between 2 sites in training and testing will be an important factor in this competition. Will need more time to test it though.</p>",
      "rawMarkdown": "&gt; But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this?\n\nVery likely. The better consistency between 2 sites, the better the models.\nSimplest way to deal with this is just average the prediction from 2 sites. -&gt; My PB gains 3% from doing this, likely more if I have stronger model.\n\nI think how you deal with the consistency between 2 sites in training and testing will be an important factor in this competition. Will need more time to test it though.",
      "votes": null
    },
    {
      "id": "577429",
      "postDate": "07/16/2019 17:12:02",
      "content": "<p>thank you very much, Phuc! Here is how i am creating the validation. Sorry for the long script. I am basically just hard coding. :( </p>\n\n<p>Basic idea is for each experiment id in the test set, divide half to get the number of batches for validation.</p>\n\n<p>Anyone: please feel free to help me \"optimize\" my pandas code to fewer lines.</p>\n\n<p>```\ntr_experiments = train_df['experiment'].unique()\nte_experiments = test_df['experiment'].unique()</p>\n\n<p>experiments = np.concatenate([tr_experiments, te_experiments])</p>\n\n<p>labels = [1] * train_df['experiment'].nunique()\nlabels += [0] * test_df['experiment'].nunique()</p>\n\n<p>exp_df = pd.DataFrame({'id': experiments,\n                       'labels': labels})</p>\n\n<p>exp_df = exp_df.sort_values('id').reset_index(drop=True)</p>\n\n<p>exp_df['batch'] = exp_df['id'].apply(lambda o: o.split('-')[0])</p>\n\n<p>g = exp_df.groupby(['batch'])['labels'].value_counts()</p>\n\n<p>g = g.unstack().reset_index()</p>\n\n<p>g[2] = g[0] // 2</p>\n\n<p>exp_df = pd.merge(exp_df, g, on='batch')</p>\n\n<p>exp_df = exp_df.rename(columns={2: 'splits'})</p>\n\n<p>exp_df.drop([0, 1], axis=1, inplace=True)</p>\n\n<p>tmp = exp_df[exp_df['labels'] == 1]</p>\n\n<p>exp_g = tmp.groupby(['batch']).groups</p>\n\n<p>val_cnt = dict(zip(tmp.groupby('batch')['splits'].mean().keys(),\n              tmp.groupby('batch')['splits'].mean().values))</p>\n\n<p>val_indices = []</p>\n\n<p>for i, batch in enumerate(exp_g):\n    val_indices += exp_g[batch][-val_cnt[batch]:],</p>\n\n<p>val_idx = [v for vs in val_indices for v in vs]</p>\n\n<p>exp_df['val']  = 0</p>\n\n<p>exp_df.loc[exp_df.index.isin(val_idx), 'val'] = 1</p>\n\n<p>exp_df = exp_df[['id', 'val']].copy()</p>\n\n<p>exp_df\n```</p>",
      "rawMarkdown": "thank you very much, Phuc! Here is how i am creating the validation. Sorry for the long script. I am basically just hard coding. :( \n\nBasic idea is for each experiment id in the test set, divide half to get the number of batches for validation.\n\nAnyone: please feel free to help me \"optimize\" my pandas code to fewer lines.\n\n```\ntr_experiments = train_df['experiment'].unique()\nte_experiments = test_df['experiment'].unique()\n\nexperiments = np.concatenate([tr_experiments, te_experiments])\n\nlabels = [1] * train_df['experiment'].nunique()\nlabels += [0] * test_df['experiment'].nunique()\n\nexp_df = pd.DataFrame({'id': experiments,\n                       'labels': labels})\n\nexp_df = exp_df.sort_values('id').reset_index(drop=True)\n\nexp_df['batch'] = exp_df['id'].apply(lambda o: o.split('-')[0])\n\ng = exp_df.groupby(['batch'])['labels'].value_counts()\n\ng = g.unstack().reset_index()\n\ng[2] = g[0] // 2\n\nexp_df = pd.merge(exp_df, g, on='batch')\n\nexp_df = exp_df.rename(columns={2: 'splits'})\n\nexp_df.drop([0, 1], axis=1, inplace=True)\n\ntmp = exp_df[exp_df['labels'] == 1]\n\nexp_g = tmp.groupby(['batch']).groups\n\nval_cnt = dict(zip(tmp.groupby('batch')['splits'].mean().keys(),\n              tmp.groupby('batch')['splits'].mean().values))\n\nval_indices = []\n\nfor i, batch in enumerate(exp_g):\n    val_indices += exp_g[batch][-val_cnt[batch]:],\n    \nval_idx = [v for vs in val_indices for v in vs]\n\nexp_df['val']  = 0\n\nexp_df.loc[exp_df.index.isin(val_idx), 'val'] = 1\n\nexp_df = exp_df[['id', 'val']].copy()\n\nexp_df\n```",
      "votes": null
    },
    {
      "id": "577464",
      "postDate": "07/16/2019 17:53:41",
      "content": "<p>Amit Kumar Jaiswal's validation strategy is currently face loss.</p>",
      "rawMarkdown": "Amit Kumar Jaiswal's validation strategy is currently face loss.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 577280,
      "author_name": "lkhphuc",
      "author_url": "",
      "post_date": "07/16/2019 15:08:12",
      "content": "<ol>\n<li>Each well has 2 images from 2 sites (2 different patches)</li>\n<li>I stratified by cell types and group by experiments.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 577295,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/16/2019 15:19:58",
          "content": "<p>thanks, Phuc! Sorry this might be a noob question. But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this? </p>\n\n<p>also, what does it mean by stratified by <code>cell_types</code> and group by <code>experiments</code>?</p>\n\n<p>say an instance is <code>HEPG2-01</code>, we have <code>cell_type</code> of <code>HEPG2</code> and <code>experiment</code> of <code>01</code>. Since we have 51 instances, do we validate by according to the <code>24 in HUVEC, 11 in RPE, 11 in HepG2, and 5 in U2OS</code> introduced? Say having 20 HUVEC in <code>train</code>, 4 HUVEC in <code>test</code>, we then split 20 HUVEC into 80:20 ratio into training and validation set?</p>\n\n<p>Thanks, once again! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577391,
          "author_name": "lkhphuc",
          "author_url": "",
          "post_date": "07/16/2019 16:43:51",
          "content": "<p><code>\ndef stratified_groups_kfold(df, n_splits=5, random_state=0):\n    df['cell'] = df['experiment']\n    all_exps = pd.Series(df['experiment'].unique())\n    all_cells = all_exps.apply(lambda x: x.split('-')[0])\n    folds = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n    for idx_tr, idx_val in folds.split(all_cells, all_cells):\n        exps_tr = all_exps[idx_tr]\n        exps_val = all_exps[idx_val]\n        idx_tr_new = df[df['experiment'].isin(exps_tr)].index.to_numpy()\n        idx_val_new = df[df['experiment'].isin(exps_val)].index.to_numpy()\n        yield idx_tr_new, idx_val_new\n</code>\nHere how I stratified by all cells and group by experiments</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577404,
          "author_name": "lkhphuc",
          "author_url": "",
          "post_date": "07/16/2019 16:50:01",
          "content": "<p>&gt; But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this?</p>\n\n<p>Very likely. The better consistency between 2 sites, the better the models.\nSimplest way to deal with this is just average the prediction from 2 sites. -&gt; My PB gains 3% from doing this, likely more if I have stronger model.</p>\n\n<p>I think how you deal with the consistency between 2 sites in training and testing will be an important factor in this competition. Will need more time to test it though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 577429,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/16/2019 17:12:02",
          "content": "<p>thank you very much, Phuc! Here is how i am creating the validation. Sorry for the long script. I am basically just hard coding. :( </p>\n\n<p>Basic idea is for each experiment id in the test set, divide half to get the number of batches for validation.</p>\n\n<p>Anyone: please feel free to help me \"optimize\" my pandas code to fewer lines.</p>\n\n<p>```\ntr_experiments = train_df['experiment'].unique()\nte_experiments = test_df['experiment'].unique()</p>\n\n<p>experiments = np.concatenate([tr_experiments, te_experiments])</p>\n\n<p>labels = [1] * train_df['experiment'].nunique()\nlabels += [0] * test_df['experiment'].nunique()</p>\n\n<p>exp_df = pd.DataFrame({'id': experiments,\n                       'labels': labels})</p>\n\n<p>exp_df = exp_df.sort_values('id').reset_index(drop=True)</p>\n\n<p>exp_df['batch'] = exp_df['id'].apply(lambda o: o.split('-')[0])</p>\n\n<p>g = exp_df.groupby(['batch'])['labels'].value_counts()</p>\n\n<p>g = g.unstack().reset_index()</p>\n\n<p>g[2] = g[0] // 2</p>\n\n<p>exp_df = pd.merge(exp_df, g, on='batch')</p>\n\n<p>exp_df = exp_df.rename(columns={2: 'splits'})</p>\n\n<p>exp_df.drop([0, 1], axis=1, inplace=True)</p>\n\n<p>tmp = exp_df[exp_df['labels'] == 1]</p>\n\n<p>exp_g = tmp.groupby(['batch']).groups</p>\n\n<p>val_cnt = dict(zip(tmp.groupby('batch')['splits'].mean().keys(),\n              tmp.groupby('batch')['splits'].mean().values))</p>\n\n<p>val_indices = []</p>\n\n<p>for i, batch in enumerate(exp_g):\n    val_indices += exp_g[batch][-val_cnt[batch]:],</p>\n\n<p>val_idx = [v for vs in val_indices for v in vs]</p>\n\n<p>exp_df['val']  = 0</p>\n\n<p>exp_df.loc[exp_df.index.isin(val_idx), 'val'] = 1</p>\n\n<p>exp_df = exp_df[['id', 'val']].copy()</p>\n\n<p>exp_df\n```</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 577464,
      "author_name": "wjshenggggg",
      "author_url": "",
      "post_date": "07/16/2019 17:53:41",
      "content": "<p>Amit Kumar Jaiswal's validation strategy is currently face loss.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "577268": "Hi guys,\n\n1. I wonder why the length of `test` folder is twice the length of `test.csv`?\nAnswer by Phuc Le: Each well has 2 images from 2 sites (2 different patches)\n\n2. Also, what are some of the validation strategies people are trying?\n\nThank you.\n\nEdit: Follow up on Q2: How to make use of `train_controls.csv` and `test_controls.csv`?",
    "577280": "1. Each well has 2 images from 2 sites (2 different patches)\n2. I stratified by cell types and group by experiments.",
    "577295": "thanks, Phuc! Sorry this might be a noob question. But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this? \n\nalso, what does it mean by stratified by `cell_types` and group by `experiments`?\n\nsay an instance is `HEPG2-01`, we have `cell_type` of `HEPG2` and `experiment` of `01`. Since we have 51 instances, do we validate by according to the `24 in HUVEC, 11 in RPE, 11 in HepG2, and 5 in U2OS` introduced? Say having 20 HUVEC in `train`, 4 HUVEC in `test`, we then split 20 HUVEC into 80:20 ratio into training and validation set?\n\nThanks, once again!",
    "577391": "```\ndef stratified_groups_kfold(df, n_splits=5, random_state=0):\n    df['cell'] = df['experiment']\n    all_exps = pd.Series(df['experiment'].unique())\n    all_cells = all_exps.apply(lambda x: x.split('-')[0])\n    folds = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=random_state)\n    for idx_tr, idx_val in folds.split(all_cells, all_cells):\n        exps_tr = all_exps[idx_tr]\n        exps_val = all_exps[idx_val]\n        idx_tr_new = df[df['experiment'].isin(exps_tr)].index.to_numpy()\n        idx_val_new = df[df['experiment'].isin(exps_val)].index.to_numpy()\n        yield idx_tr_new, idx_val_new\n```\nHere how I stratified by all cells and group by experiments",
    "577404": "&gt; But is it possible that the model predicts two labels for the 2 sites? If so, how would we deal with this?\n\nVery likely. The better consistency between 2 sites, the better the models.\nSimplest way to deal with this is just average the prediction from 2 sites. -&gt; My PB gains 3% from doing this, likely more if I have stronger model.\n\nI think how you deal with the consistency between 2 sites in training and testing will be an important factor in this competition. Will need more time to test it though.",
    "577429": "thank you very much, Phuc! Here is how i am creating the validation. Sorry for the long script. I am basically just hard coding. :( \n\nBasic idea is for each experiment id in the test set, divide half to get the number of batches for validation.\n\nAnyone: please feel free to help me \"optimize\" my pandas code to fewer lines.\n\n```\ntr_experiments = train_df['experiment'].unique()\nte_experiments = test_df['experiment'].unique()\n\nexperiments = np.concatenate([tr_experiments, te_experiments])\n\nlabels = [1] * train_df['experiment'].nunique()\nlabels += [0] * test_df['experiment'].nunique()\n\nexp_df = pd.DataFrame({'id': experiments,\n                       'labels': labels})\n\nexp_df = exp_df.sort_values('id').reset_index(drop=True)\n\nexp_df['batch'] = exp_df['id'].apply(lambda o: o.split('-')[0])\n\ng = exp_df.groupby(['batch'])['labels'].value_counts()\n\ng = g.unstack().reset_index()\n\ng[2] = g[0] // 2\n\nexp_df = pd.merge(exp_df, g, on='batch')\n\nexp_df = exp_df.rename(columns={2: 'splits'})\n\nexp_df.drop([0, 1], axis=1, inplace=True)\n\ntmp = exp_df[exp_df['labels'] == 1]\n\nexp_g = tmp.groupby(['batch']).groups\n\nval_cnt = dict(zip(tmp.groupby('batch')['splits'].mean().keys(),\n              tmp.groupby('batch')['splits'].mean().values))\n\nval_indices = []\n\nfor i, batch in enumerate(exp_g):\n    val_indices += exp_g[batch][-val_cnt[batch]:],\n    \nval_idx = [v for vs in val_indices for v in vs]\n\nexp_df['val']  = 0\n\nexp_df.loc[exp_df.index.isin(val_idx), 'val'] = 1\n\nexp_df = exp_df[['id', 'val']].copy()\n\nexp_df\n```",
    "577464": "Amit Kumar Jaiswal's validation strategy is currently face loss."
  },
  "source": "meta"
}