{
  "id": 417133,
  "title": "The public dataset is the training data! How to test real model robustness?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417133",
  "author_name": "",
  "post_date": "2023-06-14T10:22:36.998002900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As it stands the test data is a split of the training data and the current lb probably contains some over-fitting.<br>\nOn a portion of the validation set that I manually segmented, the best result was only about 0.6, which made me very anxious. Does anyone know of any additional data to test the robustness of the model?</p>",
  "messages": [
    {
      "id": "2302082",
      "postDate": "06/14/2023 10:22:36",
      "content": "<p>As it stands the test data is a split of the training data and the current lb probably contains some over-fitting.<br>\nOn a portion of the validation set that I manually segmented, the best result was only about 0.6, which made me very anxious. Does anyone know of any additional data to test the robustness of the model?</p>",
      "rawMarkdown": "As it stands the test data is a split of the training data and the current lb probably contains some over-fitting.\nOn a portion of the validation set that I manually segmented, the best result was only about 0.6, which made me very anxious. Does anyone know of any additional data to test the robustness of the model?",
      "votes": null
    },
    {
      "id": "2302145",
      "postDate": "06/14/2023 11:15:09",
      "content": "<p>You can use the standard evaluation metrics such as: Accuracy, Precision, Recall, F1 score. Try to repeat the evaluation multiple times.</p>",
      "rawMarkdown": "You can use the standard evaluation metrics such as: Accuracy, Precision, Recall, F1 score. Try to repeat the evaluation multiple times.",
      "votes": null
    },
    {
      "id": "2302270",
      "postDate": "06/14/2023 12:42:07",
      "content": "<p>The public dataset is not the training data. Those two pieces (a and b) are just dummy inputs for your notebook for the commit part.<br>\nThe real test set is processed during the submission.</p>",
      "rawMarkdown": "The public dataset is not the training data. Those two pieces (a and b) are just dummy inputs for your notebook for the commit part.\nThe real test set is processed during the submission.",
      "votes": null
    },
    {
      "id": "2302442",
      "postDate": "06/14/2023 14:57:46",
      "content": "<p>Non-public public datasets😂</p>",
      "rawMarkdown": "Non-public public datasets😂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2302145,
      "author_name": "gianetan",
      "author_url": "",
      "post_date": "06/14/2023 11:15:09",
      "content": "<p>You can use the standard evaluation metrics such as: Accuracy, Precision, Recall, F1 score. Try to repeat the evaluation multiple times.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2302270,
      "author_name": "igorkrashenyi",
      "author_url": "",
      "post_date": "06/14/2023 12:42:07",
      "content": "<p>The public dataset is not the training data. Those two pieces (a and b) are just dummy inputs for your notebook for the commit part.<br>\nThe real test set is processed during the submission.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2302442,
          "author_name": "jimmyisme1",
          "author_url": "",
          "post_date": "06/14/2023 14:57:46",
          "content": "<p>Non-public public datasets😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2302082": "As it stands the test data is a split of the training data and the current lb probably contains some over-fitting.\nOn a portion of the validation set that I manually segmented, the best result was only about 0.6, which made me very anxious. Does anyone know of any additional data to test the robustness of the model?",
    "2302145": "You can use the standard evaluation metrics such as: Accuracy, Precision, Recall, F1 score. Try to repeat the evaluation multiple times.",
    "2302270": "The public dataset is not the training data. Those two pieces (a and b) are just dummy inputs for your notebook for the commit part.\nThe real test set is processed during the submission.",
    "2302442": "Non-public public datasets😂"
  },
  "source": "meta"
}