{
  "id": 68333,
  "title": "Is there any point in Re training?",
  "url": "/competitions/airbus-ship-detection/discussion/68333",
  "author_name": "",
  "post_date": "2018-10-11T15:54:41.884025700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>New set of training images are obtained out of  old Test and Train images. New test images seem to be quite different from Train data,probably that could be the reason why  genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test , if we continue to train old model some how get its accuracy to about 98 ,i doubt if we can ever go above .8 score with new test data . \nSo i request Organizers team to please validate if Train data  and Test data has any degree of similarity so that model trained on trained data is able to generalize well for unseen test data</p>",
  "messages": [
    {
      "id": "402379",
      "postDate": "10/11/2018 15:54:41",
      "content": "<p>New set of training images are obtained out of  old Test and Train images. New test images seem to be quite different from Train data,probably that could be the reason why  genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test , if we continue to train old model some how get its accuracy to about 98 ,i doubt if we can ever go above .8 score with new test data . \nSo i request Organizers team to please validate if Train data  and Test data has any degree of similarity so that model trained on trained data is able to generalize well for unseen test data</p>",
      "rawMarkdown": "New set of training images are obtained out of  old Test and Train images. New test images seem to be quite different from Train data,probably that could be the reason why  genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test , if we continue to train old model some how get its accuracy to about 98 ,i doubt if we can ever go above .8 score with new test data . \nSo i request Organizers team to please validate if Train data  and Test data has any degree of similarity so that model trained on trained data is able to generalize well for unseen test data",
      "votes": null
    },
    {
      "id": "403018",
      "postDate": "10/12/2018 18:12:21",
      "content": "<blockquote>\n  <p>New test images seem to be quite different from Train data,probably that could be the reason why genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test </p>\n</blockquote>\n\n<p>I don't think that's the case. It seems like for the old test set, the images used for public leaderboard were drawn uniformly from the whole data pool, thus the ship/no-ship ratio in the public test set was roughly the same as in the full set (85/15).</p>\n\n<p>For this new set, my best classifier predicted that only about 2.7k/15k test images contain ship. The ratio isn't that much of a different from before, however in the 12% data they used for public testing 48% of them contain ship! So it seems like they intentionally skewed the distribution for public testing, thus making it harder to get high score since you'll now have much less free IoU thanks to the vast amount of no-ship images.</p>",
      "rawMarkdown": "&gt; New test images seem to be quite different from Train data,probably that could be the reason why genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test \n\nI don't think that's the case. It seems like for the old test set, the images used for public leaderboard were drawn uniformly from the whole data pool, thus the ship/no-ship ratio in the public test set was roughly the same as in the full set (85/15).\n\nFor this new set, my best classifier predicted that only about 2.7k/15k test images contain ship. The ratio isn't that much of a different from before, however in the 12% data they used for public testing 48% of them contain ship! So it seems like they intentionally skewed the distribution for public testing, thus making it harder to get high score since you'll now have much less free IoU thanks to the vast amount of no-ship images.",
      "votes": null
    },
    {
      "id": "403061",
      "postDate": "10/12/2018 20:02:51",
      "content": "<p>Also, the former test set was overlapping with the train set. Your model was predicting on boats it had already learned to detect, just at a different place in the pictures</p>",
      "rawMarkdown": "Also, the former test set was overlapping with the train set. Your model was predicting on boats it had already learned to detect, just at a different place in the pictures",
      "votes": null
    },
    {
      "id": "403224",
      "postDate": "10/13/2018 05:27:27",
      "content": "<p>can you please explain how you arrived at the number 48%</p>",
      "rawMarkdown": "can you please explain how you arrived at the number 48%",
      "votes": null
    },
    {
      "id": "403238",
      "postDate": "10/13/2018 06:13:01",
      "content": "<p>There are many early submission with the same score of 0.52, I think they simply uploaded a blank file. Similarly there were a lot of early 0.847 submissions before the leak.</p>\n\n<p>I'm not 100% sure though, since I don't want to waste a submission to confirm that information, but I think I read a comment somewhere confirmed that too.</p>",
      "rawMarkdown": "There are many early submission with the same score of 0.52, I think they simply uploaded a blank file. Similarly there were a lot of early 0.847 submissions before the leak.\n\nI'm not 100% sure though, since I don't want to waste a submission to confirm that information, but I think I read a comment somewhere confirmed that too.",
      "votes": null
    },
    {
      "id": "403344",
      "postDate": "10/13/2018 12:06:18",
      "content": "<p>makes sense. Thank you!</p>",
      "rawMarkdown": "makes sense. Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 403018,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "10/12/2018 18:12:21",
      "content": "<blockquote>\n  <p>New test images seem to be quite different from Train data,probably that could be the reason why genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test </p>\n</blockquote>\n\n<p>I don't think that's the case. It seems like for the old test set, the images used for public leaderboard were drawn uniformly from the whole data pool, thus the ship/no-ship ratio in the public test set was roughly the same as in the full set (85/15).</p>\n\n<p>For this new set, my best classifier predicted that only about 2.7k/15k test images contain ship. The ratio isn't that much of a different from before, however in the 12% data they used for public testing 48% of them contain ship! So it seems like they intentionally skewed the distribution for public testing, thus making it harder to get high score since you'll now have much less free IoU thanks to the vast amount of no-ship images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 403224,
          "author_name": "ravivadapalli",
          "author_url": "",
          "post_date": "10/13/2018 05:27:27",
          "content": "<p>can you please explain how you arrived at the number 48%</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403238,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "10/13/2018 06:13:01",
          "content": "<p>There are many early submission with the same score of 0.52, I think they simply uploaded a blank file. Similarly there were a lot of early 0.847 submissions before the leak.</p>\n\n<p>I'm not 100% sure though, since I don't want to waste a submission to confirm that information, but I think I read a comment somewhere confirmed that too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403344,
          "author_name": "ravivadapalli",
          "author_url": "",
          "post_date": "10/13/2018 12:06:18",
          "content": "<p>makes sense. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 403061,
      "author_name": "mxdbld",
      "author_url": "",
      "post_date": "10/12/2018 20:02:51",
      "content": "<p>Also, the former test set was overlapping with the train set. Your model was predicting on boats it had already learned to detect, just at a different place in the pictures</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "402379": "New set of training images are obtained out of  old Test and Train images. New test images seem to be quite different from Train data,probably that could be the reason why  genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test , if we continue to train old model some how get its accuracy to about 98 ,i doubt if we can ever go above .8 score with new test data . \nSo i request Organizers team to please validate if Train data  and Test data has any degree of similarity so that model trained on trained data is able to generalize well for unseen test data",
    "403018": "&gt; New test images seem to be quite different from Train data,probably that could be the reason why genuinely trained Model which was able to achieve test accuracy above 90 pct ,is just able to achieve 70 pct of accuracy against new test \n\nI don't think that's the case. It seems like for the old test set, the images used for public leaderboard were drawn uniformly from the whole data pool, thus the ship/no-ship ratio in the public test set was roughly the same as in the full set (85/15).\n\nFor this new set, my best classifier predicted that only about 2.7k/15k test images contain ship. The ratio isn't that much of a different from before, however in the 12% data they used for public testing 48% of them contain ship! So it seems like they intentionally skewed the distribution for public testing, thus making it harder to get high score since you'll now have much less free IoU thanks to the vast amount of no-ship images.",
    "403061": "Also, the former test set was overlapping with the train set. Your model was predicting on boats it had already learned to detect, just at a different place in the pictures",
    "403224": "can you please explain how you arrived at the number 48%",
    "403238": "There are many early submission with the same score of 0.52, I think they simply uploaded a blank file. Similarly there were a lot of early 0.847 submissions before the leak.\n\nI'm not 100% sure though, since I don't want to waste a submission to confirm that information, but I think I read a comment somewhere confirmed that too.",
    "403344": "makes sense. Thank you!"
  },
  "source": "meta"
}