{
  "id": 123478,
  "title": "Can organizers provide more test images? Please!",
  "url": "/competitions/bengaliai-cv19/discussion/123478",
  "author_name": "",
  "post_date": "2019-12-28T03:24:20.482614400Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Testing on 12 test images is too small to give sufficient confidence to the participants. Can organizers consider exposing more test images in the data provided?</p>",
  "messages": [
    {
      "id": "704818",
      "postDate": "12/28/2019 03:24:20",
      "content": "<p>Testing on 12 test images is too small to give sufficient confidence to the participants. Can organizers consider exposing more test images in the data provided?</p>",
      "rawMarkdown": "Testing on 12 test images is too small to give sufficient confidence to the participants. Can organizers consider exposing more test images in the data provided?",
      "votes": null
    },
    {
      "id": "705087",
      "postDate": "12/28/2019 11:46:51",
      "content": "<p>Hi,\nThe public test set contains around 100,000 samples, you can only access 12 through notebooks, but it gets swapped with the full test set during inference.</p>",
      "rawMarkdown": "Hi,\nThe public test set contains around 100,000 samples, you can only access 12 through notebooks, but it gets swapped with the full test set during inference.",
      "votes": null
    },
    {
      "id": "705091",
      "postDate": "12/28/2019 11:54:27",
      "content": "<p>Appreciate your response <a href=\"/imtiazprio\">@imtiazprio</a> , but dont you think, adding more than 12 images to the test_data provided to us would give strength to decision before submitting the results. The test set swap occurs after submission and that is beyond our(competitors') control.</p>",
      "rawMarkdown": "Appreciate your response @imtiazprio , but dont you think, adding more than 12 images to the test_data provided to us would give strength to decision before submitting the results. The test set swap occurs after submission and that is beyond our(competitors') control.",
      "votes": null
    },
    {
      "id": "705133",
      "postDate": "12/28/2019 13:55:48",
      "content": "<p>You can split training set and test and images not used in training</p>",
      "rawMarkdown": "You can split training set and test and images not used in training",
      "votes": null
    },
    {
      "id": "705568",
      "postDate": "12/29/2019 04:28:52",
      "content": "<p>Did you try cross-validating or maybe keeping a holdout in the training set for validation? I think some of the participants are and they have reported their results <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198\">here</a>. The test set is solely for testing the final models, not validating model performance in any way.</p>\n\n<p>Does this answer your question or are you more worried about the qualitative verification of the test set? In that case since the cross-validation and public test set results are well correlated according to the discussion thread I have linked, you can assume that there is relatively less heterogeneity between the two?</p>",
      "rawMarkdown": "Did you try cross-validating or maybe keeping a holdout in the training set for validation? I think some of the participants are and they have reported their results [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198). The test set is solely for testing the final models, not validating model performance in any way.\n\nDoes this answer your question or are you more worried about the qualitative verification of the test set? In that case since the cross-validation and public test set results are well correlated according to the discussion thread I have linked, you can assume that there is relatively less heterogeneity between the two?",
      "votes": null
    },
    {
      "id": "705574",
      "postDate": "12/29/2019 04:53:28",
      "content": "<p>Sure Ahmed, Thanks</p>",
      "rawMarkdown": "Sure Ahmed, Thanks",
      "votes": null
    },
    {
      "id": "705605",
      "postDate": "12/29/2019 06:11:27",
      "content": "<p>That will eat our gpu quota! Keep it small.</p>",
      "rawMarkdown": "That will eat our gpu quota! Keep it small.",
      "votes": null
    },
    {
      "id": "705627",
      "postDate": "12/29/2019 07:37:19",
      "content": "<p>Just in case if one didn't check that public LB is evaluated based on 47% of data: the provided 12 images are used just to verify that code works, it is not a public part of the test set. Organizers use this approach to fight with PL, I guess, and also save our GPU time.</p>",
      "rawMarkdown": "Just in case if one didn't check that public LB is evaluated based on 47% of data: the provided 12 images are used just to verify that code works, it is not a public part of the test set. Organizers use this approach to fight with PL, I guess, and also save our GPU time.",
      "votes": null
    },
    {
      "id": "705635",
      "postDate": "12/29/2019 07:56:39",
      "content": "<p>Noted @lafoss</p>",
      "rawMarkdown": "Noted @lafoss",
      "votes": null
    },
    {
      "id": "706496",
      "postDate": "12/30/2019 13:19:05",
      "content": "<p>Noted <a href=\"/hanjoonchoe\">@hanjoonchoe</a> </p>",
      "rawMarkdown": "Noted @hanjoonchoe",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 705087,
      "author_name": "imtiazprio",
      "author_url": "",
      "post_date": "12/28/2019 11:46:51",
      "content": "<p>Hi,\nThe public test set contains around 100,000 samples, you can only access 12 through notebooks, but it gets swapped with the full test set during inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 705091,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "12/28/2019 11:54:27",
          "content": "<p>Appreciate your response <a href=\"/imtiazprio\">@imtiazprio</a> , but dont you think, adding more than 12 images to the test_data provided to us would give strength to decision before submitting the results. The test set swap occurs after submission and that is beyond our(competitors') control.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705568,
          "author_name": "imtiazprio",
          "author_url": "",
          "post_date": "12/29/2019 04:28:52",
          "content": "<p>Did you try cross-validating or maybe keeping a holdout in the training set for validation? I think some of the participants are and they have reported their results <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198\">here</a>. The test set is solely for testing the final models, not validating model performance in any way.</p>\n\n<p>Does this answer your question or are you more worried about the qualitative verification of the test set? In that case since the cross-validation and public test set results are well correlated according to the discussion thread I have linked, you can assume that there is relatively less heterogeneity between the two?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705574,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "12/29/2019 04:53:28",
          "content": "<p>Sure Ahmed, Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705133,
      "author_name": "macarrony00",
      "author_url": "",
      "post_date": "12/28/2019 13:55:48",
      "content": "<p>You can split training set and test and images not used in training</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 705605,
      "author_name": "hanjoonchoe",
      "author_url": "",
      "post_date": "12/29/2019 06:11:27",
      "content": "<p>That will eat our gpu quota! Keep it small.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706496,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "12/30/2019 13:19:05",
          "content": "<p>Noted <a href=\"/hanjoonchoe\">@hanjoonchoe</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705627,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "12/29/2019 07:37:19",
      "content": "<p>Just in case if one didn't check that public LB is evaluated based on 47% of data: the provided 12 images are used just to verify that code works, it is not a public part of the test set. Organizers use this approach to fight with PL, I guess, and also save our GPU time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 705635,
          "author_name": "chandraroy",
          "author_url": "",
          "post_date": "12/29/2019 07:56:39",
          "content": "<p>Noted @lafoss</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "704818": "Testing on 12 test images is too small to give sufficient confidence to the participants. Can organizers consider exposing more test images in the data provided?",
    "705087": "Hi,\nThe public test set contains around 100,000 samples, you can only access 12 through notebooks, but it gets swapped with the full test set during inference.",
    "705091": "Appreciate your response @imtiazprio , but dont you think, adding more than 12 images to the test_data provided to us would give strength to decision before submitting the results. The test set swap occurs after submission and that is beyond our(competitors') control.",
    "705133": "You can split training set and test and images not used in training",
    "705568": "Did you try cross-validating or maybe keeping a holdout in the training set for validation? I think some of the participants are and they have reported their results [here](https://www.kaggle.com/c/bengaliai-cv19/discussion/123198). The test set is solely for testing the final models, not validating model performance in any way.\n\nDoes this answer your question or are you more worried about the qualitative verification of the test set? In that case since the cross-validation and public test set results are well correlated according to the discussion thread I have linked, you can assume that there is relatively less heterogeneity between the two?",
    "705574": "Sure Ahmed, Thanks",
    "705605": "That will eat our gpu quota! Keep it small.",
    "705627": "Just in case if one didn't check that public LB is evaluated based on 47% of data: the provided 12 images are used just to verify that code works, it is not a public part of the test set. Organizers use this approach to fight with PL, I guess, and also save our GPU time.",
    "705635": "Noted @lafoss",
    "706496": "Noted @hanjoonchoe"
  },
  "source": "meta"
}