{
  "id": 39973,
  "title": "Train Dice Coefficient vs Validation DC vs LB",
  "url": "/competitions/carvana-image-masking-challenge/discussion/39973",
  "author_name": "",
  "post_date": "2017-09-25T10:02:29.011474700Z",
  "votes": 1,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Friends,\nJust curious to know how do above values look like for you.\nFor me its Train: 0.9966, Validation:0.9966, LB: 0.9963\nThanks</p>",
  "messages": [
    {
      "id": "224142",
      "postDate": "09/25/2017 10:02:29",
      "content": "<p>Hi Friends,\nJust curious to know how do above values look like for you.\nFor me its Train: 0.9966, Validation:0.9966, LB: 0.9963\nThanks</p>",
      "rawMarkdown": "Hi Friends,\nJust curious to know how do above values look like for you.\nFor me its Train: 0.9966, Validation:0.9966, LB: 0.9963\nThanks",
      "votes": null
    },
    {
      "id": "224163",
      "postDate": "09/25/2017 12:03:03",
      "content": "<p>For me</p>\n\n<p>Train:0.9966</p>\n\n<p>Validation:0.9966</p>\n\n<p>LB:0.9968</p>",
      "rawMarkdown": "For me\n\nTrain:0.9966\n\nValidation:0.9966\n\nLB:0.9968",
      "votes": null
    },
    {
      "id": "224191",
      "postDate": "09/25/2017 13:46:19",
      "content": "<p>For me 0.9975 on CV and  low border of 0.9970 on LB. \nHoping that it's because bad cars in public LB) </p>",
      "rawMarkdown": "For me 0.9975 on CV and  low border of 0.9970 on LB. \nHoping that it's because bad cars in public LB)",
      "votes": null
    },
    {
      "id": "224201",
      "postDate": "09/25/2017 14:09:15",
      "content": "<p>you split train/valid by car?</p>\n\n<p>e.g. all 16 views images of car id =xxx only appear in train or validation, but not both?</p>",
      "rawMarkdown": "you split train/valid by car?\n\ne.g. all 16 views images of car id =xxx only appear in train or validation, but not both?",
      "votes": null
    },
    {
      "id": "224210",
      "postDate": "09/25/2017 14:35:11",
      "content": "<p>For my UNet:</p>\n\n<ul>\n<li>CV = 0.9971 </li>\n<li>LB = 0.9972</li>\n</ul>",
      "rawMarkdown": "For my UNet:\n\n - CV = 0.9971 \n - LB = 0.9972",
      "votes": null
    },
    {
      "id": "224223",
      "postDate": "09/25/2017 15:04:45",
      "content": "<p>Train: 0.9956, Validation: 0.9959, LB: 0.9968</p>",
      "rawMarkdown": "Train: 0.9956, Validation: 0.9959, LB: 0.9968",
      "votes": null
    },
    {
      "id": "224233",
      "postDate": "09/25/2017 15:50:05",
      "content": "<p>no, I splited all images, so in valid and train may be the same car. </p>\n\n<p>At first I was spliting by car, but then decided to split by images. Have no time to try the split by car. Maybe I had to split by car, so there would be more variance in splits and better generalization...</p>\n\n<p>Actually that's an idea to train two pipelines on 1) split by car and 2) split by all images and then ensemble it? </p>",
      "rawMarkdown": "no, I splited all images, so in valid and train may be the same car. \n\nAt first I was spliting by car, but then decided to split by images. Have no time to try the split by car. Maybe I had to split by car, so there would be more variance in splits and better generalization...\n\nActually that's an idea to train two pipelines on 1) split by car and 2) split by all images and then ensemble it?",
      "votes": null
    },
    {
      "id": "224296",
      "postDate": "09/25/2017 20:43:14",
      "content": "<p>If you do not split by cars, your cross-validation is not representative :/</p>",
      "rawMarkdown": "If you do not split by cars, your cross-validation is not representative :/",
      "votes": null
    },
    {
      "id": "224306",
      "postDate": "09/25/2017 21:32:38",
      "content": "<p>For my UNet: Train 0.9961, Validation: 0.9961, LB: 0.9966 :( </p>\n\n<p>One of my major problem that I could not improve my result is because of running my code in a cluster as I do not have my own powerful machine to run the code. However, I have not permission to bring changes into Keres Optimiser file to use accumulator gradient to give my network with bigger batch size. </p>",
      "rawMarkdown": "For my UNet: Train 0.9961, Validation: 0.9961, LB: 0.9966 :( \n\nOne of my major problem that I could not improve my result is because of running my code in a cluster as I do not have my own powerful machine to run the code. However, I have not permission to bring changes into Keres Optimiser file to use accumulator gradient to give my network with bigger batch size.",
      "votes": null
    },
    {
      "id": "224317",
      "postDate": "09/25/2017 22:28:03",
      "content": "<p>unet</p>\n\n<p>train: 9969, val: 9971, lb: 9958...</p>",
      "rawMarkdown": "unet\n\ntrain: 9969, val: 9971, lb: 9958...",
      "votes": null
    },
    {
      "id": "224425",
      "postDate": "09/26/2017 09:34:35",
      "content": "<p>can you explain in some details why? </p>",
      "rawMarkdown": "can you explain in some details why?",
      "votes": null
    },
    {
      "id": "224435",
      "postDate": "09/26/2017 10:09:45",
      "content": "<p>By splitting the same car into train and validation you may get biased score because your model may remember some features of the car and \"overfit\" to it even if car has another angle or something. Splitting by car makes sure that you always validate on new cars that weren't in the train set (this is more representative of our situation where we never see the test vehicles)</p>",
      "rawMarkdown": "By splitting the same car into train and validation you may get biased score because your model may remember some features of the car and \"overfit\" to it even if car has another angle or something. Splitting by car makes sure that you always validate on new cars that weren't in the train set (this is more representative of our situation where we never see the test vehicles)",
      "votes": null
    },
    {
      "id": "224443",
      "postDate": "09/26/2017 10:45:49",
      "content": "<p>thank you.   Yeah , now I see, that splitting by angle would be appropriate if we had to train on angles 1-8 and predict angles 8-16 in test.</p>",
      "rawMarkdown": "thank you.   Yeah , now I see, that splitting by angle would be appropriate if we had to train on angles 1-8 and predict angles 8-16 in test.",
      "votes": null
    },
    {
      "id": "224460",
      "postDate": "09/26/2017 12:12:55",
      "content": "<p>on our UNET CV 9971 or 9965, and LB 9969</p>\n\n<p>One thing you might want to note: when calculating the dice_coeff of CV, you're probably calculating the soft dice_coeff of your model with the raw predictions, however in the LB it's the hard dice_coeff after rounding the softmax to 1. For this reason, the LB should be slightly above the dice_coeff</p>\n\n<p>I guess you could use <code>dice_coeff_hard</code> metric if you wanted:</p>\n\n<p><code>\n  def dice_coeff_hard(y_true, y_pred):\n      smooth = 1.\n      y_true_f = K.flatten(y_true)\n      y_pred_f = K.round(K.flatten(y_pred))\n      intersection = K.sum(y_true_f * y_pred_f)\n      score = (2. * intersection + smooth) / (K.sum(y_true_f) + K.sum(y_pred_f) + smooth)\n      return score\n</code></p>",
      "rawMarkdown": "on our UNET CV 9971 or 9965, and LB 9969\n\nOne thing you might want to note: when calculating the dice_coeff of CV, you're probably calculating the soft dice_coeff of your model with the raw predictions, however in the LB it's the hard dice_coeff after rounding the softmax to 1. For this reason, the LB should be slightly above the dice_coeff\n\nI guess you could use `dice_coeff_hard` metric if you wanted:\n\n```\n  def dice_coeff_hard(y_true, y_pred):\n      smooth = 1.\n      y_true_f = K.flatten(y_true)\n      y_pred_f = K.round(K.flatten(y_pred))\n      intersection = K.sum(y_true_f * y_pred_f)\n      score = (2. * intersection + smooth) / (K.sum(y_true_f) + K.sum(y_pred_f) + smooth)\n      return score\n```",
      "votes": null
    },
    {
      "id": "224565",
      "postDate": "09/26/2017 19:38:14",
      "content": "<p>use this in pesudo labelling. some view are more correct than others.</p>\n\n<p>label and train in some view , test in others. and if we reverse and repeat, the results should converge?</p>",
      "rawMarkdown": "use this in pesudo labelling. some view are more correct than others.\n\nlabel and train in some view , test in others. and if we reverse and repeat, the results should converge?",
      "votes": null
    },
    {
      "id": "224629",
      "postDate": "09/27/2017 03:50:55",
      "content": "<p>pre-train an initial 1024unet on resized 1024x1024 images. Then finetune on crops of 1024x1024 from full resolution. this gives VC and LB of about 0.9970 for me.</p>",
      "rawMarkdown": "pre-train an initial 1024unet on resized 1024x1024 images. Then finetune on crops of 1024x1024 from full resolution. this gives VC and LB of about 0.9970 for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224163,
      "author_name": "zhangsongwei",
      "author_url": "",
      "post_date": "09/25/2017 12:03:03",
      "content": "<p>For me</p>\n\n<p>Train:0.9966</p>\n\n<p>Validation:0.9966</p>\n\n<p>LB:0.9968</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224191,
      "author_name": "heyt0ny",
      "author_url": "",
      "post_date": "09/25/2017 13:46:19",
      "content": "<p>For me 0.9975 on CV and  low border of 0.9970 on LB. \nHoping that it's because bad cars in public LB) </p>",
      "votes": null,
      "replies": [
        {
          "id": 224201,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/25/2017 14:09:15",
          "content": "<p>you split train/valid by car?</p>\n\n<p>e.g. all 16 views images of car id =xxx only appear in train or validation, but not both?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224233,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "09/25/2017 15:50:05",
          "content": "<p>no, I splited all images, so in valid and train may be the same car. </p>\n\n<p>At first I was spliting by car, but then decided to split by images. Have no time to try the split by car. Maybe I had to split by car, so there would be more variance in splits and better generalization...</p>\n\n<p>Actually that's an idea to train two pipelines on 1) split by car and 2) split by all images and then ensemble it? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224296,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "09/25/2017 20:43:14",
          "content": "<p>If you do not split by cars, your cross-validation is not representative :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224425,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "09/26/2017 09:34:35",
          "content": "<p>can you explain in some details why? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224435,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "09/26/2017 10:09:45",
          "content": "<p>By splitting the same car into train and validation you may get biased score because your model may remember some features of the car and \"overfit\" to it even if car has another angle or something. Splitting by car makes sure that you always validate on new cars that weren't in the train set (this is more representative of our situation where we never see the test vehicles)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224443,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "09/26/2017 10:45:49",
          "content": "<p>thank you.   Yeah , now I see, that splitting by angle would be appropriate if we had to train on angles 1-8 and predict angles 8-16 in test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 224565,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "09/26/2017 19:38:14",
          "content": "<p>use this in pesudo labelling. some view are more correct than others.</p>\n\n<p>label and train in some view , test in others. and if we reverse and repeat, the results should converge?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224210,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "09/25/2017 14:35:11",
      "content": "<p>For my UNet:</p>\n\n<ul>\n<li>CV = 0.9971 </li>\n<li>LB = 0.9972</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224223,
      "author_name": "truepk",
      "author_url": "",
      "post_date": "09/25/2017 15:04:45",
      "content": "<p>Train: 0.9956, Validation: 0.9959, LB: 0.9968</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224306,
      "author_name": "svesal",
      "author_url": "",
      "post_date": "09/25/2017 21:32:38",
      "content": "<p>For my UNet: Train 0.9961, Validation: 0.9961, LB: 0.9966 :( </p>\n\n<p>One of my major problem that I could not improve my result is because of running my code in a cluster as I do not have my own powerful machine to run the code. However, I have not permission to bring changes into Keres Optimiser file to use accumulator gradient to give my network with bigger batch size. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224317,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "09/25/2017 22:28:03",
      "content": "<p>unet</p>\n\n<p>train: 9969, val: 9971, lb: 9958...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224460,
      "author_name": "burgalon",
      "author_url": "",
      "post_date": "09/26/2017 12:12:55",
      "content": "<p>on our UNET CV 9971 or 9965, and LB 9969</p>\n\n<p>One thing you might want to note: when calculating the dice_coeff of CV, you're probably calculating the soft dice_coeff of your model with the raw predictions, however in the LB it's the hard dice_coeff after rounding the softmax to 1. For this reason, the LB should be slightly above the dice_coeff</p>\n\n<p>I guess you could use <code>dice_coeff_hard</code> metric if you wanted:</p>\n\n<p><code>\n  def dice_coeff_hard(y_true, y_pred):\n      smooth = 1.\n      y_true_f = K.flatten(y_true)\n      y_pred_f = K.round(K.flatten(y_pred))\n      intersection = K.sum(y_true_f * y_pred_f)\n      score = (2. * intersection + smooth) / (K.sum(y_true_f) + K.sum(y_pred_f) + smooth)\n      return score\n</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 224629,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "09/27/2017 03:50:55",
      "content": "<p>pre-train an initial 1024unet on resized 1024x1024 images. Then finetune on crops of 1024x1024 from full resolution. this gives VC and LB of about 0.9970 for me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "224142": "Hi Friends,\nJust curious to know how do above values look like for you.\nFor me its Train: 0.9966, Validation:0.9966, LB: 0.9963\nThanks",
    "224163": "For me\n\nTrain:0.9966\n\nValidation:0.9966\n\nLB:0.9968",
    "224191": "For me 0.9975 on CV and  low border of 0.9970 on LB. \nHoping that it's because bad cars in public LB)",
    "224201": "you split train/valid by car?\n\ne.g. all 16 views images of car id =xxx only appear in train or validation, but not both?",
    "224210": "For my UNet:\n\n - CV = 0.9971 \n - LB = 0.9972",
    "224223": "Train: 0.9956, Validation: 0.9959, LB: 0.9968",
    "224233": "no, I splited all images, so in valid and train may be the same car. \n\nAt first I was spliting by car, but then decided to split by images. Have no time to try the split by car. Maybe I had to split by car, so there would be more variance in splits and better generalization...\n\nActually that's an idea to train two pipelines on 1) split by car and 2) split by all images and then ensemble it?",
    "224296": "If you do not split by cars, your cross-validation is not representative :/",
    "224306": "For my UNet: Train 0.9961, Validation: 0.9961, LB: 0.9966 :( \n\nOne of my major problem that I could not improve my result is because of running my code in a cluster as I do not have my own powerful machine to run the code. However, I have not permission to bring changes into Keres Optimiser file to use accumulator gradient to give my network with bigger batch size.",
    "224317": "unet\n\ntrain: 9969, val: 9971, lb: 9958...",
    "224425": "can you explain in some details why?",
    "224435": "By splitting the same car into train and validation you may get biased score because your model may remember some features of the car and \"overfit\" to it even if car has another angle or something. Splitting by car makes sure that you always validate on new cars that weren't in the train set (this is more representative of our situation where we never see the test vehicles)",
    "224443": "thank you.   Yeah , now I see, that splitting by angle would be appropriate if we had to train on angles 1-8 and predict angles 8-16 in test.",
    "224460": "on our UNET CV 9971 or 9965, and LB 9969\n\nOne thing you might want to note: when calculating the dice_coeff of CV, you're probably calculating the soft dice_coeff of your model with the raw predictions, however in the LB it's the hard dice_coeff after rounding the softmax to 1. For this reason, the LB should be slightly above the dice_coeff\n\nI guess you could use `dice_coeff_hard` metric if you wanted:\n\n```\n  def dice_coeff_hard(y_true, y_pred):\n      smooth = 1.\n      y_true_f = K.flatten(y_true)\n      y_pred_f = K.round(K.flatten(y_pred))\n      intersection = K.sum(y_true_f * y_pred_f)\n      score = (2. * intersection + smooth) / (K.sum(y_true_f) + K.sum(y_pred_f) + smooth)\n      return score\n```",
    "224565": "use this in pesudo labelling. some view are more correct than others.\n\nlabel and train in some view , test in others. and if we reverse and repeat, the results should converge?",
    "224629": "pre-train an initial 1024unet on resized 1024x1024 images. Then finetune on crops of 1024x1024 from full resolution. this gives VC and LB of about 0.9970 for me."
  },
  "source": "meta"
}