{
  "id": 65884,
  "title": "Trust your CV",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/65884",
  "author_name": "",
  "post_date": "2018-09-15T18:57:20.725122900Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The common wisdom is that you should trust your CV and take only slight notice of public LB scores. In general, in the past, I've been a bit skeptical of that view, and <em>in general</em> I remain a bit skeptical. (If you have a trustworthy CV, you should definitely trust it, but developing a trustworthy validation framework is a difficult task, and you cannot always be confident you have succeeded.) In this particular competition, however, <em>we have no choice but to trust our CV</em>. With a mere 1000 images in the test set, and a ground truth noisy with the difference between rectangles and reality, the public LB contains only a small amount of useful information.  Even with 25,000 images, my CV is finding it hard to make a reliable choice between models.  A score based on 1000 images is for entertainment purposes only.</p>",
  "messages": [
    {
      "id": "387844",
      "postDate": "09/15/2018 18:57:20",
      "content": "<p>The common wisdom is that you should trust your CV and take only slight notice of public LB scores. In general, in the past, I've been a bit skeptical of that view, and <em>in general</em> I remain a bit skeptical. (If you have a trustworthy CV, you should definitely trust it, but developing a trustworthy validation framework is a difficult task, and you cannot always be confident you have succeeded.) In this particular competition, however, <em>we have no choice but to trust our CV</em>. With a mere 1000 images in the test set, and a ground truth noisy with the difference between rectangles and reality, the public LB contains only a small amount of useful information.  Even with 25,000 images, my CV is finding it hard to make a reliable choice between models.  A score based on 1000 images is for entertainment purposes only.</p>",
      "rawMarkdown": "The common wisdom is that you should trust your CV and take only slight notice of public LB scores. In general, in the past, I've been a bit skeptical of that view, and *in general* I remain a bit skeptical. (If you have a trustworthy CV, you should definitely trust it, but developing a trustworthy validation framework is a difficult task, and you cannot always be confident you have succeeded.) In this particular competition, however, *we have no choice but to trust our CV*. With a mere 1000 images in the test set, and a ground truth noisy with the difference between rectangles and reality, the public LB contains only a small amount of useful information.  Even with 25,000 images, my CV is finding it hard to make a reliable choice between models.  A score based on 1000 images is for entertainment purposes only.",
      "votes": null
    },
    {
      "id": "388002",
      "postDate": "09/16/2018 04:50:00",
      "content": "<p>I agree with your wisdom --especially the point about the ground truth noise. This has been my primary focus. For example, annotators are likely to draw a rectangle from one lung edge to the other given the pneumonia is large enough. This is useful to know. I'm still debating whether or not to use some post-processing to adjust rectangles and/or to attempt to correct this noise when training. Possibly I could adjust with the bounding boxes of another trained model. I'm still not sure.</p>\n\n<p>What are your thoughts on the matter?</p>",
      "rawMarkdown": "I agree with your wisdom --especially the point about the ground truth noise. This has been my primary focus. For example, annotators are likely to draw a rectangle from one lung edge to the other given the pneumonia is large enough. This is useful to know. I'm still debating whether or not to use some post-processing to adjust rectangles and/or to attempt to correct this noise when training. Possibly I could adjust with the bounding boxes of another trained model. I'm still not sure.\n\nWhat are your thoughts on the matter?",
      "votes": null
    },
    {
      "id": "388894",
      "postDate": "09/17/2018 19:27:20",
      "content": "<p>How do your local score and LB score compare? For example, I get 0.15 local (it's not CV, just validation on 10%) and only 0.11 LB. Though maybe it's ok if variance is high.</p>\n\n<p>In the case of your CV what's the variance on our folds?</p>",
      "rawMarkdown": "How do your local score and LB score compare? For example, I get 0.15 local (it's not CV, just validation on 10%) and only 0.11 LB. Though maybe it's ok if variance is high.\n\nIn the case of your CV what's the variance on our folds?",
      "votes": null
    },
    {
      "id": "388907",
      "postDate": "09/17/2018 20:02:56",
      "content": "<p>Usually my CV scores are a little better than LB scores, but there is a lot of variation. My best LB score so far was 0.139, and the CV was 0.137 (but CV on a different set of folds was 0.131).  Individual scores among the 5 folds ranged from 0.123 to 0.151.</p>",
      "rawMarkdown": "Usually my CV scores are a little better than LB scores, but there is a lot of variation. My best LB score so far was 0.139, and the CV was 0.137 (but CV on a different set of folds was 0.131).  Individual scores among the 5 folds ranged from 0.123 to 0.151.",
      "votes": null
    },
    {
      "id": "389648",
      "postDate": "09/19/2018 02:58:33",
      "content": "<p>If we take my fold range (from the previous comment) to imply a margin of error range of 0.028, this would mean that LB positions 42 (at 0.165) through 104 (at 0.137) (including all of the bronze range and most of the sliver range) are within a margin of error of one another. Or similarly positions 140 (at 0.126) through 464 (at 0.098), which represent more than a third of the LB.</p>\n\n<p>[EDIT: Presumably the actual margin of error should be much larger, since each of my validation folds has about 5000 images, whereas the test set has only 1000.]</p>",
      "rawMarkdown": "If we take my fold range (from the previous comment) to imply a margin of error range of 0.028, this would mean that LB positions 42 (at 0.165) through 104 (at 0.137) (including all of the bronze range and most of the sliver range) are within a margin of error of one another. Or similarly positions 140 (at 0.126) through 464 (at 0.098), which represent more than a third of the LB.\n\n[EDIT: Presumably the actual margin of error should be much larger, since each of my validation folds has about 5000 images, whereas the test set has only 1000.]",
      "votes": null
    },
    {
      "id": "389968",
      "postDate": "09/19/2018 14:26:14",
      "content": "<p>I agree</p>\n\n<p>Trust your CV and build a good validation strategy.</p>\n\n<p>I tried random 10 folds splitting..and when I submitted the first fold it scored 0.116 while the averaged 10 folds scored 0.142 , suggesting the variation between folds for public LB is really huge. </p>\n\n<p>Building good validation strategy is key here ( along with k folds splitting for CV and eventually oof stacking)  and random train_test_splitting  seems definitely not the way to go. </p>",
      "rawMarkdown": "I agree\n\nTrust your CV and build a good validation strategy.\n\nI tried random 10 folds splitting..and when I submitted the first fold it scored 0.116 while the averaged 10 folds scored 0.142 , suggesting the variation between folds for public LB is really huge. \n\n\nBuilding good validation strategy is key here ( along with k folds splitting for CV and eventually oof stacking)  and random train_test_splitting  seems definitely not the way to go.",
      "votes": null
    },
    {
      "id": "390029",
      "postDate": "09/19/2018 15:42:54",
      "content": "<p>Also I think the fold scores are leptokurtic. After a fair number of 5-fold runs, I see some where one of the 5 is far out of the range of the other 4. I wonder if, instead of scoring the OOF predictions for the whole training set at once, we should take a median of fold scores, or (probably better) a trimmed mean (e.g., with the best and worst dropped), to filter out the influence of outliers on the model selection process.</p>",
      "rawMarkdown": "Also I think the fold scores are leptokurtic. After a fair number of 5-fold runs, I see some where one of the 5 is far out of the range of the other 4. I wonder if, instead of scoring the OOF predictions for the whole training set at once, we should take a median of fold scores, or (probably better) a trimmed mean (e.g., with the best and worst dropped), to filter out the influence of outliers on the model selection process.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 388002,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "09/16/2018 04:50:00",
      "content": "<p>I agree with your wisdom --especially the point about the ground truth noise. This has been my primary focus. For example, annotators are likely to draw a rectangle from one lung edge to the other given the pneumonia is large enough. This is useful to know. I'm still debating whether or not to use some post-processing to adjust rectangles and/or to attempt to correct this noise when training. Possibly I could adjust with the bounding boxes of another trained model. I'm still not sure.</p>\n\n<p>What are your thoughts on the matter?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 388894,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "09/17/2018 19:27:20",
      "content": "<p>How do your local score and LB score compare? For example, I get 0.15 local (it's not CV, just validation on 10%) and only 0.11 LB. Though maybe it's ok if variance is high.</p>\n\n<p>In the case of your CV what's the variance on our folds?</p>",
      "votes": null,
      "replies": [
        {
          "id": 388907,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/17/2018 20:02:56",
          "content": "<p>Usually my CV scores are a little better than LB scores, but there is a lot of variation. My best LB score so far was 0.139, and the CV was 0.137 (but CV on a different set of folds was 0.131).  Individual scores among the 5 folds ranged from 0.123 to 0.151.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 389648,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/19/2018 02:58:33",
          "content": "<p>If we take my fold range (from the previous comment) to imply a margin of error range of 0.028, this would mean that LB positions 42 (at 0.165) through 104 (at 0.137) (including all of the bronze range and most of the sliver range) are within a margin of error of one another. Or similarly positions 140 (at 0.126) through 464 (at 0.098), which represent more than a third of the LB.</p>\n\n<p>[EDIT: Presumably the actual margin of error should be much larger, since each of my validation folds has about 5000 images, whereas the test set has only 1000.]</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 389968,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "09/19/2018 14:26:14",
      "content": "<p>I agree</p>\n\n<p>Trust your CV and build a good validation strategy.</p>\n\n<p>I tried random 10 folds splitting..and when I submitted the first fold it scored 0.116 while the averaged 10 folds scored 0.142 , suggesting the variation between folds for public LB is really huge. </p>\n\n<p>Building good validation strategy is key here ( along with k folds splitting for CV and eventually oof stacking)  and random train_test_splitting  seems definitely not the way to go. </p>",
      "votes": null,
      "replies": [
        {
          "id": 390029,
          "author_name": "aharless",
          "author_url": "",
          "post_date": "09/19/2018 15:42:54",
          "content": "<p>Also I think the fold scores are leptokurtic. After a fair number of 5-fold runs, I see some where one of the 5 is far out of the range of the other 4. I wonder if, instead of scoring the OOF predictions for the whole training set at once, we should take a median of fold scores, or (probably better) a trimmed mean (e.g., with the best and worst dropped), to filter out the influence of outliers on the model selection process.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "387844": "The common wisdom is that you should trust your CV and take only slight notice of public LB scores. In general, in the past, I've been a bit skeptical of that view, and *in general* I remain a bit skeptical. (If you have a trustworthy CV, you should definitely trust it, but developing a trustworthy validation framework is a difficult task, and you cannot always be confident you have succeeded.) In this particular competition, however, *we have no choice but to trust our CV*. With a mere 1000 images in the test set, and a ground truth noisy with the difference between rectangles and reality, the public LB contains only a small amount of useful information.  Even with 25,000 images, my CV is finding it hard to make a reliable choice between models.  A score based on 1000 images is for entertainment purposes only.",
    "388002": "I agree with your wisdom --especially the point about the ground truth noise. This has been my primary focus. For example, annotators are likely to draw a rectangle from one lung edge to the other given the pneumonia is large enough. This is useful to know. I'm still debating whether or not to use some post-processing to adjust rectangles and/or to attempt to correct this noise when training. Possibly I could adjust with the bounding boxes of another trained model. I'm still not sure.\n\nWhat are your thoughts on the matter?",
    "388894": "How do your local score and LB score compare? For example, I get 0.15 local (it's not CV, just validation on 10%) and only 0.11 LB. Though maybe it's ok if variance is high.\n\nIn the case of your CV what's the variance on our folds?",
    "388907": "Usually my CV scores are a little better than LB scores, but there is a lot of variation. My best LB score so far was 0.139, and the CV was 0.137 (but CV on a different set of folds was 0.131).  Individual scores among the 5 folds ranged from 0.123 to 0.151.",
    "389648": "If we take my fold range (from the previous comment) to imply a margin of error range of 0.028, this would mean that LB positions 42 (at 0.165) through 104 (at 0.137) (including all of the bronze range and most of the sliver range) are within a margin of error of one another. Or similarly positions 140 (at 0.126) through 464 (at 0.098), which represent more than a third of the LB.\n\n[EDIT: Presumably the actual margin of error should be much larger, since each of my validation folds has about 5000 images, whereas the test set has only 1000.]",
    "389968": "I agree\n\nTrust your CV and build a good validation strategy.\n\nI tried random 10 folds splitting..and when I submitted the first fold it scored 0.116 while the averaged 10 folds scored 0.142 , suggesting the variation between folds for public LB is really huge. \n\n\nBuilding good validation strategy is key here ( along with k folds splitting for CV and eventually oof stacking)  and random train_test_splitting  seems definitely not the way to go.",
    "390029": "Also I think the fold scores are leptokurtic. After a fair number of 5-fold runs, I see some where one of the 5 is far out of the range of the other 4. I wonder if, instead of scoring the OOF predictions for the whole training set at once, we should take a median of fold scores, or (probably better) a trimmed mean (e.g., with the best and worst dropped), to filter out the influence of outliers on the model selection process."
  },
  "source": "meta"
}