{
  "id": 235199,
  "title": "different order different dice-score in valid set",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/235199",
  "author_name": "",
  "post_date": "2021-04-28T08:56:03.612436400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I found that there are big differences between different order of valid set. As comparing to randomly sorted set with the set positive samples first then nagetive, the latter order will return a very low score in validation.  How should I define the principles of the valid set order in order to be more reliable.</p>",
  "messages": [
    {
      "id": "1286652",
      "postDate": "04/28/2021 08:56:03",
      "content": "<p>I found that there are big differences between different order of valid set. As comparing to randomly sorted set with the set positive samples first then nagetive, the latter order will return a very low score in validation.  How should I define the principles of the valid set order in order to be more reliable.</p>",
      "rawMarkdown": "I found that there are big differences between different order of valid set. As comparing to randomly sorted set with the set positive samples first then nagetive, the latter order will return a very low score in validation.  How should I define the principles of the valid set order in order to be more reliable.",
      "votes": null
    },
    {
      "id": "1286661",
      "postDate": "04/28/2021 09:07:04",
      "content": "<p>Interesting. I wonder if that's because batches where the union is 0 of y_pred and y_true are counted as null.</p>\n<p>I then wonder if you then calculate the average Dice over the batch (rather than each tile) you will get this phenomenon, because<br>\n1) If all your negatives are at the end the entire batch is null and does not contribute to the metric, and so only the positives (which are harder) count.<br>\n2) If the negatives are interspersed, their intersection and union do contribute towards the batch-wise Dice.</p>\n<p>I suggest you just randomly shuffle your validation set using a seed and then stick with that for all experiments.</p>",
      "rawMarkdown": "Interesting. I wonder if that's because batches where the union is 0 of y_pred and y_true are counted as null.\n\nI then wonder if you then calculate the average Dice over the batch (rather than each tile) you will get this phenomenon, because\n1) If all your negatives are at the end the entire batch is null and does not contribute to the metric, and so only the positives (which are harder) count.\n2) If the negatives are interspersed, their intersection and union do contribute towards the batch-wise Dice.\n\nI suggest you just randomly shuffle your validation set using a seed and then stick with that for all experiments.",
      "votes": null
    },
    {
      "id": "1286689",
      "postDate": "04/28/2021 10:02:19",
      "content": "<p>Thanks. I accept your suggest and I just found something very interesting: </p>\n<ol>\n<li>if the first is all positives. validation dice score  ~= 0.514</li>\n<li>if the first is all nagatives. validation dice score  ~= 0.309</li>\n<li>if using principle like: one pos one nag cross（pos,nag,pos,nag … nag,pos,nag）. the score ~= 0.542.</li>\n</ol>\n<p>also, seeds are important too. I use one seed(2021) with final score = 0.903. but the other(2020) with socre = 0.928. both have nearly the same LB score.</p>\n<p>Through your advice, I will choose a seed to make each batch has a proper proportion of nag and pos or the score will mislead me.</p>\n<p>ps: valid set with 1433 nag and 687 pos. </p>",
      "rawMarkdown": "Thanks. I accept your suggest and I just found something very interesting: \n1. if the first is all positives. validation dice score  ~= 0.514\n2. if the first is all nagatives. validation dice score  ~= 0.309\n3. if using principle like: one pos one nag cross（pos,nag,pos,nag ... nag,pos,nag）. the score ~= 0.542.\n\nalso, seeds are important too. I use one seed(2021) with final score = 0.903. but the other(2020) with socre = 0.928. both have nearly the same LB score.\n\nThrough your advice, I will choose a seed to make each batch has a proper proportion of nag and pos or the score will mislead me.\n\n\n ps: valid set with 1433 nag and 687 pos.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1286661,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/28/2021 09:07:04",
      "content": "<p>Interesting. I wonder if that's because batches where the union is 0 of y_pred and y_true are counted as null.</p>\n<p>I then wonder if you then calculate the average Dice over the batch (rather than each tile) you will get this phenomenon, because<br>\n1) If all your negatives are at the end the entire batch is null and does not contribute to the metric, and so only the positives (which are harder) count.<br>\n2) If the negatives are interspersed, their intersection and union do contribute towards the batch-wise Dice.</p>\n<p>I suggest you just randomly shuffle your validation set using a seed and then stick with that for all experiments.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1286689,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "04/28/2021 10:02:19",
          "content": "<p>Thanks. I accept your suggest and I just found something very interesting: </p>\n<ol>\n<li>if the first is all positives. validation dice score  ~= 0.514</li>\n<li>if the first is all nagatives. validation dice score  ~= 0.309</li>\n<li>if using principle like: one pos one nag cross（pos,nag,pos,nag … nag,pos,nag）. the score ~= 0.542.</li>\n</ol>\n<p>also, seeds are important too. I use one seed(2021) with final score = 0.903. but the other(2020) with socre = 0.928. both have nearly the same LB score.</p>\n<p>Through your advice, I will choose a seed to make each batch has a proper proportion of nag and pos or the score will mislead me.</p>\n<p>ps: valid set with 1433 nag and 687 pos. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1286652": "I found that there are big differences between different order of valid set. As comparing to randomly sorted set with the set positive samples first then nagetive, the latter order will return a very low score in validation.  How should I define the principles of the valid set order in order to be more reliable.",
    "1286661": "Interesting. I wonder if that's because batches where the union is 0 of y_pred and y_true are counted as null.\n\nI then wonder if you then calculate the average Dice over the batch (rather than each tile) you will get this phenomenon, because\n1) If all your negatives are at the end the entire batch is null and does not contribute to the metric, and so only the positives (which are harder) count.\n2) If the negatives are interspersed, their intersection and union do contribute towards the batch-wise Dice.\n\nI suggest you just randomly shuffle your validation set using a seed and then stick with that for all experiments.",
    "1286689": "Thanks. I accept your suggest and I just found something very interesting: \n1. if the first is all positives. validation dice score  ~= 0.514\n2. if the first is all nagatives. validation dice score  ~= 0.309\n3. if using principle like: one pos one nag cross（pos,nag,pos,nag ... nag,pos,nag）. the score ~= 0.542.\n\nalso, seeds are important too. I use one seed(2021) with final score = 0.903. but the other(2020) with socre = 0.928. both have nearly the same LB score.\n\nThrough your advice, I will choose a seed to make each batch has a proper proportion of nag and pos or the score will mislead me.\n\n\n ps: valid set with 1433 nag and 687 pos."
  },
  "source": "meta"
}