{
  "id": 131375,
  "title": "pre-processing images in the test images (private dataset)?",
  "url": "/competitions/bengaliai-cv19/discussion/131375",
  "author_name": "",
  "post_date": "2020-02-19T14:42:19.268250900Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>If we do some image processing before training the model on the images in public dataset, the test image also needs to be pre-processed before testing the images in private dataset, and we are just submitting the kernel for this competition, how will this be done?</p>",
  "messages": [
    {
      "id": "750608",
      "postDate": "02/19/2020 14:42:19",
      "content": "<p>If we do some image processing before training the model on the images in public dataset, the test image also needs to be pre-processed before testing the images in private dataset, and we are just submitting the kernel for this competition, how will this be done?</p>",
      "rawMarkdown": "If we do some image processing before training the model on the images in public dataset, the test image also needs to be pre-processed before testing the images in private dataset, and we are just submitting the kernel for this competition, how will this be done?",
      "votes": null
    },
    {
      "id": "750826",
      "postDate": "02/19/2020 18:09:29",
      "content": "<p>You'll have to process the test images in the kernel, or else confirm that your training on pre-processed images is robust by running an inception using the unprocessed public dataset train images instead of the test set.</p>",
      "rawMarkdown": "You'll have to process the test images in the kernel, or else confirm that your training on pre-processed images is robust by running an inception using the unprocessed public dataset train images instead of the test set.",
      "votes": null
    },
    {
      "id": "751000",
      "postDate": "02/19/2020 22:51:06",
      "content": "<p>When you train offline, you can pre-process your images and save them. Then use the processed images for training. This will save you a lot of time. Remember, to save the preprocessing function. The function should be fast enough to pre-process all the test data from the private set. This part, you have to do live. There is no other way :)</p>",
      "rawMarkdown": "When you train offline, you can pre-process your images and save them. Then use the processed images for training. This will save you a lot of time. Remember, to save the preprocessing function. The function should be fast enough to pre-process all the test data from the private set. This part, you have to do live. There is no other way :)",
      "votes": null
    },
    {
      "id": "751666",
      "postDate": "02/20/2020 11:38:00",
      "content": "<p>Can you tell what do you mean by \" This part, you have to do live.\"?\nI am asking regarding the 'Bengali.AI Handwritten Grapheme Classification', the test data has only 12 images, now back to my question, how does the images in the private dataset (test data) will be pre-processed before (because we don't have access to the private/hidden data) passing it to the model for prediction?</p>",
      "rawMarkdown": "Can you tell what do you mean by \" This part, you have to do live.\"?\nI am asking regarding the 'Bengali.AI Handwritten Grapheme Classification', the test data has only 12 images, now back to my question, how does the images in the private dataset (test data) will be pre-processed before (because we don't have access to the private/hidden data) passing it to the model for prediction?",
      "votes": null
    },
    {
      "id": "751874",
      "postDate": "02/20/2020 15:34:56",
      "content": "<p>when your kernel runs while submitting, public test data will be replaced by private test data. so, do your processing the same way as you do in commit and it will work.</p>",
      "rawMarkdown": "when your kernel runs while submitting, public test data will be replaced by private test data. so, do your processing the same way as you do in commit and it will work.",
      "votes": null
    },
    {
      "id": "751887",
      "postDate": "02/20/2020 16:00:20",
      "content": "<p>The public test has 100,000 images and the private test has 100,000 images. In this competition, both are hidden from us. The 12 images that we can download are probably neither public nor private test. They are there so we can debug our code.</p>",
      "rawMarkdown": "The public test has 100,000 images and the private test has 100,000 images. In this competition, both are hidden from us. The 12 images that we can download are probably neither public nor private test. They are there so we can debug our code.",
      "votes": null
    },
    {
      "id": "751916",
      "postDate": "02/20/2020 16:28:12",
      "content": "<p>Thanks, <a href=\"/abhishek\">@abhishek</a> and <a href=\"/cdeotte\">@cdeotte</a>, but I have one more question, If I do data pre-processing in one notebook (kernel) and then use that data in another notebook (kernel) to train the model, how will the private dataset be handled in that case?</p>",
      "rawMarkdown": "Thanks, @abhishek and @cdeotte, but I have one more question, If I do data pre-processing in one notebook (kernel) and then use that data in another notebook (kernel) to train the model, how will the private dataset be handled in that case?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 750826,
      "author_name": "ljschuster",
      "author_url": "",
      "post_date": "02/19/2020 18:09:29",
      "content": "<p>You'll have to process the test images in the kernel, or else confirm that your training on pre-processed images is robust by running an inception using the unprocessed public dataset train images instead of the test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 751000,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "02/19/2020 22:51:06",
      "content": "<p>When you train offline, you can pre-process your images and save them. Then use the processed images for training. This will save you a lot of time. Remember, to save the preprocessing function. The function should be fast enough to pre-process all the test data from the private set. This part, you have to do live. There is no other way :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 751666,
          "author_name": "amitjslearn",
          "author_url": "",
          "post_date": "02/20/2020 11:38:00",
          "content": "<p>Can you tell what do you mean by \" This part, you have to do live.\"?\nI am asking regarding the 'Bengali.AI Handwritten Grapheme Classification', the test data has only 12 images, now back to my question, how does the images in the private dataset (test data) will be pre-processed before (because we don't have access to the private/hidden data) passing it to the model for prediction?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751874,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "02/20/2020 15:34:56",
          "content": "<p>when your kernel runs while submitting, public test data will be replaced by private test data. so, do your processing the same way as you do in commit and it will work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751887,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/20/2020 16:00:20",
          "content": "<p>The public test has 100,000 images and the private test has 100,000 images. In this competition, both are hidden from us. The 12 images that we can download are probably neither public nor private test. They are there so we can debug our code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751916,
          "author_name": "amitjslearn",
          "author_url": "",
          "post_date": "02/20/2020 16:28:12",
          "content": "<p>Thanks, <a href=\"/abhishek\">@abhishek</a> and <a href=\"/cdeotte\">@cdeotte</a>, but I have one more question, If I do data pre-processing in one notebook (kernel) and then use that data in another notebook (kernel) to train the model, how will the private dataset be handled in that case?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "750608": "If we do some image processing before training the model on the images in public dataset, the test image also needs to be pre-processed before testing the images in private dataset, and we are just submitting the kernel for this competition, how will this be done?",
    "750826": "You'll have to process the test images in the kernel, or else confirm that your training on pre-processed images is robust by running an inception using the unprocessed public dataset train images instead of the test set.",
    "751000": "When you train offline, you can pre-process your images and save them. Then use the processed images for training. This will save you a lot of time. Remember, to save the preprocessing function. The function should be fast enough to pre-process all the test data from the private set. This part, you have to do live. There is no other way :)",
    "751666": "Can you tell what do you mean by \" This part, you have to do live.\"?\nI am asking regarding the 'Bengali.AI Handwritten Grapheme Classification', the test data has only 12 images, now back to my question, how does the images in the private dataset (test data) will be pre-processed before (because we don't have access to the private/hidden data) passing it to the model for prediction?",
    "751874": "when your kernel runs while submitting, public test data will be replaced by private test data. so, do your processing the same way as you do in commit and it will work.",
    "751887": "The public test has 100,000 images and the private test has 100,000 images. In this competition, both are hidden from us. The 12 images that we can download are probably neither public nor private test. They are there so we can debug our code.",
    "751916": "Thanks, @abhishek and @cdeotte, but I have one more question, If I do data pre-processing in one notebook (kernel) and then use that data in another notebook (kernel) to train the model, how will the private dataset be handled in that case?"
  },
  "source": "meta"
}