{
  "id": 69704,
  "title": "Stage 2: To train or not to train....",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/69704",
  "author_name": "",
  "post_date": "2018-10-26T07:51:54.356850800Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>So with stage 2 data now available it turns out that there is also stage 2 training data. From the original competition rules I concluded that there would be only stage 2 test data....from reading the other discussion topics I'am not the only one confused about that.</p>\n\n<p>Anyway .. it is a given fact that there is stage 2 training data now...so what to do with it?</p>\n\n<p>My trained model for stage 1 performed already better then I personally expected compared to what the model scored when training locally....so I think just using the model from stage 1 to generate the submission for the stage 2 test data. Also because stage 2 LB only uses 1 % of the test data to give an indication of the model score...so it is a bit of a gamble to retrain with stage 2 data.</p>\n\n<p>Curious to know what you will be doing?</p>",
  "messages": [
    {
      "id": "410520",
      "postDate": "10/26/2018 07:51:54",
      "content": "<p>So with stage 2 data now available it turns out that there is also stage 2 training data. From the original competition rules I concluded that there would be only stage 2 test data....from reading the other discussion topics I'am not the only one confused about that.</p>\n\n<p>Anyway .. it is a given fact that there is stage 2 training data now...so what to do with it?</p>\n\n<p>My trained model for stage 1 performed already better then I personally expected compared to what the model scored when training locally....so I think just using the model from stage 1 to generate the submission for the stage 2 test data. Also because stage 2 LB only uses 1 % of the test data to give an indication of the model score...so it is a bit of a gamble to retrain with stage 2 data.</p>\n\n<p>Curious to know what you will be doing?</p>",
      "rawMarkdown": "So with stage 2 data now available it turns out that there is also stage 2 training data. From the original competition rules I concluded that there would be only stage 2 test data....from reading the other discussion topics I'am not the only one confused about that.\n\nAnyway .. it is a given fact that there is stage 2 training data now...so what to do with it?\n\nMy trained model for stage 1 performed already better then I personally expected compared to what the model scored when training locally....so I think just using the model from stage 1 to generate the submission for the stage 2 test data. Also because stage 2 LB only uses 1 % of the test data to give an indication of the model score...so it is a bit of a gamble to retrain with stage 2 data.\n\nCurious to know what you will be doing?",
      "votes": null
    },
    {
      "id": "410532",
      "postDate": "10/26/2018 08:11:14",
      "content": "<p>I thought 1% is just a joke....</p>",
      "rawMarkdown": "I thought 1% is just a joke....",
      "votes": null
    },
    {
      "id": "410543",
      "postDate": "10/26/2018 08:31:34",
      "content": "<p>Nope...sorry about that.\nIt says so on the leaderbord.</p>",
      "rawMarkdown": "Nope...sorry about that.\nIt says so on the leaderbord.",
      "votes": null
    },
    {
      "id": "410561",
      "postDate": "10/26/2018 09:06:17",
      "content": "<p>Use 80% for training, 20% for validation... </p>",
      "rawMarkdown": "Use 80% for training, 20% for validation...",
      "votes": null
    },
    {
      "id": "410571",
      "postDate": "10/26/2018 09:28:11",
      "content": "<p>Maybe it is true,I don't know too~</p>",
      "rawMarkdown": "Maybe it is true,I don't know too~",
      "votes": null
    },
    {
      "id": "410710",
      "postDate": "10/26/2018 14:03:58",
      "content": "<p>Hi Moshel,</p>\n\n<p>Thanks yes I 'am aware of using the 80/20 split for that. My code that was uploaded has to limited automatic tuning of parameters to be usable for stage 2 training data...so I will stick with my model for stage 1.</p>",
      "rawMarkdown": "Hi Moshel,\n\nThanks yes I 'am aware of using the 80/20 split for that. My code that was uploaded has to limited automatic tuning of parameters to be usable for stage 2 training data...so I will stick with my model for stage 1.",
      "votes": null
    },
    {
      "id": "410744",
      "postDate": "10/26/2018 15:37:21",
      "content": "<p>Having the same question myself. Given the extra 1000 images from stage1 test set (with about 350 positives), how much one can get from a retraining? Not to mention the restrictions that have been imposed on the second stage.</p>",
      "rawMarkdown": "Having the same question myself. Given the extra 1000 images from stage1 test set (with about 350 positives), how much one can get from a retraining? Not to mention the restrictions that have been imposed on the second stage.",
      "votes": null
    },
    {
      "id": "410876",
      "postDate": "10/26/2018 19:46:32",
      "content": "<p>Word of caution... If you do go with retraining, make sure you use the same split as before and add 80/20 split of the 1000 to each set. If you re-split the training set with random shuffle, you will get some if the previous training set in your valuation set, and believe me, you don't want that! </p>",
      "rawMarkdown": "Word of caution... If you do go with retraining, make sure you use the same split as before and add 80/20 split of the 1000 to each set. If you re-split the training set with random shuffle, you will get some if the previous training set in your valuation set, and believe me, you don't want that!",
      "votes": null
    },
    {
      "id": "410880",
      "postDate": "10/26/2018 20:17:44",
      "content": "<p>Thanks. I will keep that in mind. I will try my stage 1 model to generate a stage 2 submission later this evening. I will do some training runs with the stage 2 training data...a few days to decide if I will use it as a final submission.</p>",
      "rawMarkdown": "Thanks. I will keep that in mind. I will try my stage 1 model to generate a stage 2 submission later this evening. I will do some training runs with the stage 2 training data...a few days to decide if I will use it as a final submission.",
      "votes": null
    },
    {
      "id": "410887",
      "postDate": "10/26/2018 20:37:00",
      "content": "<p>I'm going to stick with my stage 1 model. The ~330 extra pneumonia cases aren't worth it for me tbh, especially when due to the ambiguity in the original upload requirements I didn't upload all the code required for training but instead just all model weights and code to generate predictions. If I re-trained and somehow managed to win something, I guess I wouldn't be able to prove that I did so with the same code I used to train the originally submitted models.</p>\n\n<p>So fingers crossed for the original models. Good luck all.</p>",
      "rawMarkdown": "I'm going to stick with my stage 1 model. The ~330 extra pneumonia cases aren't worth it for me tbh, especially when due to the ambiguity in the original upload requirements I didn't upload all the code required for training but instead just all model weights and code to generate predictions. If I re-trained and somehow managed to win something, I guess I wouldn't be able to prove that I did so with the same code I used to train the originally submitted models.\n\nSo fingers crossed for the original models. Good luck all.",
      "votes": null
    },
    {
      "id": "411010",
      "postDate": "10/27/2018 07:03:10",
      "content": "<p>@Robin\nDid you use your old model for your current score? It is high enough.</p>",
      "rawMarkdown": "Robin\nDid you use your old model for your current score? It is high enough.",
      "votes": null
    },
    {
      "id": "411034",
      "postDate": "10/27/2018 08:48:03",
      "content": "<p>@Sergey, Yes...that score is based on the exact same model weights as used for my stage 1 final score...... :-)</p>\n\n<p>Is it high enough? Well based on the 1% of test data used for the Public Leaderboard it is. Just a few more days of waiting. I'am sure there will be a lot more great submissions of the other participants coming.\nFor me the answer to my original question is clear...Not to train. I'll leave it with this one and only submission.</p>",
      "rawMarkdown": "Sergey, Yes...that score is based on the exact same model weights as used for my stage 1 final score...... :-)\n\nIs it high enough? Well based on the 1% of test data used for the Public Leaderboard it is. Just a few more days of waiting. I'am sure there will be a lot more great submissions of the other participants coming.\nFor me the answer to my original question is clear...Not to train. I'll leave it with this one and only submission.",
      "votes": null
    },
    {
      "id": "411194",
      "postDate": "10/27/2018 15:43:01",
      "content": "<blockquote>\n  <p>I'm going to stick with my stage 1 model. </p>\n</blockquote>\n\n<p>However you have already 5 submissions. ;)</p>",
      "rawMarkdown": "&gt; I'm going to stick with my stage 1 model. \n\nHowever you have already 5 submissions. ;)",
      "votes": null
    },
    {
      "id": "411250",
      "postDate": "10/27/2018 17:34:44",
      "content": "<p>It may surprise you to find out I'm not going to base my final submission on its performance over 30 images.. </p>",
      "rawMarkdown": "It may surprise you to find out I'm not going to base my final submission on its performance over 30 images..",
      "votes": null
    },
    {
      "id": "413425",
      "postDate": "10/31/2018 23:47:31",
      "content": "<p>As many people had previously commented, the majority of the images are from stage 1. For some reason when I train the model I was using with the new dataset, my score would drop. I am using the weights trained during stage 1.</p>",
      "rawMarkdown": "As many people had previously commented, the majority of the images are from stage 1. For some reason when I train the model I was using with the new dataset, my score would drop. I am using the weights trained during stage 1.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 410532,
      "author_name": "zzz333",
      "author_url": "",
      "post_date": "10/26/2018 08:11:14",
      "content": "<p>I thought 1% is just a joke....</p>",
      "votes": null,
      "replies": [
        {
          "id": 410543,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "10/26/2018 08:31:34",
          "content": "<p>Nope...sorry about that.\nIt says so on the leaderbord.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410561,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/26/2018 09:06:17",
          "content": "<p>Use 80% for training, 20% for validation... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410571,
          "author_name": "zzz333",
          "author_url": "",
          "post_date": "10/26/2018 09:28:11",
          "content": "<p>Maybe it is true,I don't know too~</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410710,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "10/26/2018 14:03:58",
          "content": "<p>Hi Moshel,</p>\n\n<p>Thanks yes I 'am aware of using the 80/20 split for that. My code that was uploaded has to limited automatic tuning of parameters to be usable for stage 2 training data...so I will stick with my model for stage 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410876,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "10/26/2018 19:46:32",
          "content": "<p>Word of caution... If you do go with retraining, make sure you use the same split as before and add 80/20 split of the 1000 to each set. If you re-split the training set with random shuffle, you will get some if the previous training set in your valuation set, and believe me, you don't want that! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410880,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "10/26/2018 20:17:44",
          "content": "<p>Thanks. I will keep that in mind. I will try my stage 1 model to generate a stage 2 submission later this evening. I will do some training runs with the stage 2 training data...a few days to decide if I will use it as a final submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411010,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "10/27/2018 07:03:10",
          "content": "<p>@Robin\nDid you use your old model for your current score? It is high enough.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411034,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "10/27/2018 08:48:03",
          "content": "<p>@Sergey, Yes...that score is based on the exact same model weights as used for my stage 1 final score...... :-)</p>\n\n<p>Is it high enough? Well based on the 1% of test data used for the Public Leaderboard it is. Just a few more days of waiting. I'am sure there will be a lot more great submissions of the other participants coming.\nFor me the answer to my original question is clear...Not to train. I'll leave it with this one and only submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410744,
      "author_name": "xinario",
      "author_url": "",
      "post_date": "10/26/2018 15:37:21",
      "content": "<p>Having the same question myself. Given the extra 1000 images from stage1 test set (with about 350 positives), how much one can get from a retraining? Not to mention the restrictions that have been imposed on the second stage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 410887,
      "author_name": "taindow",
      "author_url": "",
      "post_date": "10/26/2018 20:37:00",
      "content": "<p>I'm going to stick with my stage 1 model. The ~330 extra pneumonia cases aren't worth it for me tbh, especially when due to the ambiguity in the original upload requirements I didn't upload all the code required for training but instead just all model weights and code to generate predictions. If I re-trained and somehow managed to win something, I guess I wouldn't be able to prove that I did so with the same code I used to train the originally submitted models.</p>\n\n<p>So fingers crossed for the original models. Good luck all.</p>",
      "votes": null,
      "replies": [
        {
          "id": 411194,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "10/27/2018 15:43:01",
          "content": "<blockquote>\n  <p>I'm going to stick with my stage 1 model. </p>\n</blockquote>\n\n<p>However you have already 5 submissions. ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411250,
          "author_name": "taindow",
          "author_url": "",
          "post_date": "10/27/2018 17:34:44",
          "content": "<p>It may surprise you to find out I'm not going to base my final submission on its performance over 30 images.. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 413425,
      "author_name": "peeyai",
      "author_url": "",
      "post_date": "10/31/2018 23:47:31",
      "content": "<p>As many people had previously commented, the majority of the images are from stage 1. For some reason when I train the model I was using with the new dataset, my score would drop. I am using the weights trained during stage 1.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "410520": "So with stage 2 data now available it turns out that there is also stage 2 training data. From the original competition rules I concluded that there would be only stage 2 test data....from reading the other discussion topics I'am not the only one confused about that.\n\nAnyway .. it is a given fact that there is stage 2 training data now...so what to do with it?\n\nMy trained model for stage 1 performed already better then I personally expected compared to what the model scored when training locally....so I think just using the model from stage 1 to generate the submission for the stage 2 test data. Also because stage 2 LB only uses 1 % of the test data to give an indication of the model score...so it is a bit of a gamble to retrain with stage 2 data.\n\nCurious to know what you will be doing?",
    "410532": "I thought 1% is just a joke....",
    "410543": "Nope...sorry about that.\nIt says so on the leaderbord.",
    "410561": "Use 80% for training, 20% for validation...",
    "410571": "Maybe it is true,I don't know too~",
    "410710": "Hi Moshel,\n\nThanks yes I 'am aware of using the 80/20 split for that. My code that was uploaded has to limited automatic tuning of parameters to be usable for stage 2 training data...so I will stick with my model for stage 1.",
    "410744": "Having the same question myself. Given the extra 1000 images from stage1 test set (with about 350 positives), how much one can get from a retraining? Not to mention the restrictions that have been imposed on the second stage.",
    "410876": "Word of caution... If you do go with retraining, make sure you use the same split as before and add 80/20 split of the 1000 to each set. If you re-split the training set with random shuffle, you will get some if the previous training set in your valuation set, and believe me, you don't want that!",
    "410880": "Thanks. I will keep that in mind. I will try my stage 1 model to generate a stage 2 submission later this evening. I will do some training runs with the stage 2 training data...a few days to decide if I will use it as a final submission.",
    "410887": "I'm going to stick with my stage 1 model. The ~330 extra pneumonia cases aren't worth it for me tbh, especially when due to the ambiguity in the original upload requirements I didn't upload all the code required for training but instead just all model weights and code to generate predictions. If I re-trained and somehow managed to win something, I guess I wouldn't be able to prove that I did so with the same code I used to train the originally submitted models.\n\nSo fingers crossed for the original models. Good luck all.",
    "411010": "Robin\nDid you use your old model for your current score? It is high enough.",
    "411034": "Sergey, Yes...that score is based on the exact same model weights as used for my stage 1 final score...... :-)\n\nIs it high enough? Well based on the 1% of test data used for the Public Leaderboard it is. Just a few more days of waiting. I'am sure there will be a lot more great submissions of the other participants coming.\nFor me the answer to my original question is clear...Not to train. I'll leave it with this one and only submission.",
    "411194": "&gt; I'm going to stick with my stage 1 model. \n\nHowever you have already 5 submissions. ;)",
    "411250": "It may surprise you to find out I'm not going to base my final submission on its performance over 30 images..",
    "413425": "As many people had previously commented, the majority of the images are from stage 1. For some reason when I train the model I was using with the new dataset, my score would drop. I am using the weights trained during stage 1."
  },
  "source": "meta"
}