{
  "id": 111795,
  "title": "Shakeup possible?",
  "url": "/competitions/understanding_cloud_organization/discussion/111795",
  "author_name": "",
  "post_date": "2019-10-09T03:07:14.411340Z",
  "votes": 6,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I did some experiments on CV. The metric is the exact Dice Coefficient (with dice=1 for both empty prediction and ground truth). The CV-LB relationship seems not strongly concur, when I tried different optimisation strategies for post-processing. \nHas anyone seen the same thing as me?</p>",
  "messages": [
    {
      "id": "644581",
      "postDate": "10/09/2019 03:07:14",
      "content": "<p>I did some experiments on CV. The metric is the exact Dice Coefficient (with dice=1 for both empty prediction and ground truth). The CV-LB relationship seems not strongly concur, when I tried different optimisation strategies for post-processing. \nHas anyone seen the same thing as me?</p>",
      "rawMarkdown": "I did some experiments on CV. The metric is the exact Dice Coefficient (with dice=1 for both empty prediction and ground truth). The CV-LB relationship seems not strongly concur, when I tried different optimisation strategies for post-processing. \nHas anyone seen the same thing as me?",
      "votes": null
    },
    {
      "id": "653025",
      "postDate": "10/19/2019 20:11:00",
      "content": "<p>I agree with you. My validation and lb also do not seem to correlate, so big shake-up is possible.</p>",
      "rawMarkdown": "I agree with you. My validation and lb also do not seem to correlate, so big shake-up is possible.",
      "votes": null
    },
    {
      "id": "653205",
      "postDate": "10/20/2019 03:47:11",
      "content": "<p>That's what I concern. Given public LB has 25% of test data, equivalently to roughly 1000 images. Then, by correcting only 4 rows in submission.csv as empty (from a prediction of non-empty mask), and if it matches ground-truth, we can improve 0.001 public LB. Given a lot of people have same LB scores, shakeup is inevitable. </p>",
      "rawMarkdown": "That's what I concern. Given public LB has 25% of test data, equivalently to roughly 1000 images. Then, by correcting only 4 rows in submission.csv as empty (from a prediction of non-empty mask), and if it matches ground-truth, we can improve 0.001 public LB. Given a lot of people have same LB scores, shakeup is inevitable.",
      "votes": null
    },
    {
      "id": "674460",
      "postDate": "11/16/2019 14:38:55",
      "content": "<p>Going back to this topic. I still cannot improve CV to better than 0.662, corresponding to LB 0.671. Not sure somebody else around me on LB has better/worse CV? I still think shake up is coming. But I may be wrong...</p>",
      "rawMarkdown": "Going back to this topic. I still cannot improve CV to better than 0.662, corresponding to LB 0.671. Not sure somebody else around me on LB has better/worse CV? I still think shake up is coming. But I may be wrong...",
      "votes": null
    },
    {
      "id": "674484",
      "postDate": "11/16/2019 14:58:40",
      "content": "<p>since the Public LB is only 10% of the data, scores can be quite noisy. For me the same CV score can vary on the LB by about .0025. As you pointed out, each image on the pub LB is worth more than 0.001.</p>",
      "rawMarkdown": "since the Public LB is only 10% of the data, scores can be quite noisy. For me the same CV score can vary on the LB by about .0025. As you pointed out, each image on the pub LB is worth more than 0.001.",
      "votes": null
    },
    {
      "id": "674495",
      "postDate": "11/16/2019 15:26:55",
      "content": "<p>It’s 25%, not 10%. </p>",
      "rawMarkdown": "It’s 25%, not 10%.",
      "votes": null
    },
    {
      "id": "674504",
      "postDate": "11/16/2019 15:41:02",
      "content": "<p>Sorry, 10% of the whole dataset. 60% train, 10% Public test, 30% Private test. </p>",
      "rawMarkdown": "Sorry, 10% of the whole dataset. 60% train, 10% Public test, 30% Private test.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 653025,
      "author_name": "anuragtr",
      "author_url": "",
      "post_date": "10/19/2019 20:11:00",
      "content": "<p>I agree with you. My validation and lb also do not seem to correlate, so big shake-up is possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 653205,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "10/20/2019 03:47:11",
          "content": "<p>That's what I concern. Given public LB has 25% of test data, equivalently to roughly 1000 images. Then, by correcting only 4 rows in submission.csv as empty (from a prediction of non-empty mask), and if it matches ground-truth, we can improve 0.001 public LB. Given a lot of people have same LB scores, shakeup is inevitable. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674460,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "11/16/2019 14:38:55",
      "content": "<p>Going back to this topic. I still cannot improve CV to better than 0.662, corresponding to LB 0.671. Not sure somebody else around me on LB has better/worse CV? I still think shake up is coming. But I may be wrong...</p>",
      "votes": null,
      "replies": [
        {
          "id": 674484,
          "author_name": "robga",
          "author_url": "",
          "post_date": "11/16/2019 14:58:40",
          "content": "<p>since the Public LB is only 10% of the data, scores can be quite noisy. For me the same CV score can vary on the LB by about .0025. As you pointed out, each image on the pub LB is worth more than 0.001.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674495,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "11/16/2019 15:26:55",
          "content": "<p>It’s 25%, not 10%. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674504,
          "author_name": "robga",
          "author_url": "",
          "post_date": "11/16/2019 15:41:02",
          "content": "<p>Sorry, 10% of the whole dataset. 60% train, 10% Public test, 30% Private test. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "644581": "I did some experiments on CV. The metric is the exact Dice Coefficient (with dice=1 for both empty prediction and ground truth). The CV-LB relationship seems not strongly concur, when I tried different optimisation strategies for post-processing. \nHas anyone seen the same thing as me?",
    "653025": "I agree with you. My validation and lb also do not seem to correlate, so big shake-up is possible.",
    "653205": "That's what I concern. Given public LB has 25% of test data, equivalently to roughly 1000 images. Then, by correcting only 4 rows in submission.csv as empty (from a prediction of non-empty mask), and if it matches ground-truth, we can improve 0.001 public LB. Given a lot of people have same LB scores, shakeup is inevitable.",
    "674460": "Going back to this topic. I still cannot improve CV to better than 0.662, corresponding to LB 0.671. Not sure somebody else around me on LB has better/worse CV? I still think shake up is coming. But I may be wrong...",
    "674484": "since the Public LB is only 10% of the data, scores can be quite noisy. For me the same CV score can vary on the LB by about .0025. As you pointed out, each image on the pub LB is worth more than 0.001.",
    "674495": "It’s 25%, not 10%.",
    "674504": "Sorry, 10% of the whole dataset. 60% train, 10% Public test, 30% Private test."
  },
  "source": "meta"
}