{
  "id": 436123,
  "title": "More robust validation needed?",
  "url": "/competitions/predict-ai-model-runtime/discussion/436123",
  "author_name": "",
  "post_date": "2023-09-01T06:55:04.883300200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Playing with hyperparams my notebooks Val score increases like crazy but seems in overfiting and this doesn't generalize to lb<br>\n<a href=\"https://www.kaggle.com/code/jainam213/0-97val0-12lb\" target=\"_blank\">https://www.kaggle.com/code/jainam213/0-97val0-12lb</a><br>\nWe need a more robust validation dataset to correlate cv and lb closely </p>",
  "messages": [
    {
      "id": "2418136",
      "postDate": "09/01/2023 06:55:04",
      "content": "<p>Playing with hyperparams my notebooks Val score increases like crazy but seems in overfiting and this doesn't generalize to lb<br>\n<a href=\"https://www.kaggle.com/code/jainam213/0-97val0-12lb\" target=\"_blank\">https://www.kaggle.com/code/jainam213/0-97val0-12lb</a><br>\nWe need a more robust validation dataset to correlate cv and lb closely </p>",
      "rawMarkdown": "Playing with hyperparams my notebooks Val score increases like crazy but seems in overfiting and this doesn't generalize to lb\nhttps://www.kaggle.com/code/jainam213/0-97val0-12lb\nWe need a more robust validation dataset to correlate cv and lb closely",
      "votes": null
    },
    {
      "id": "2419099",
      "postDate": "09/01/2023 17:22:12",
      "content": "<p>Thank you for pointing this out. I skimmed through your notebook, and it looks like you use the final evaluation metric for validation score. The top-k score for the tile collection is definitely not suitable for validation because your model can just get lucky and happen to pick one good config in the top k candidates and have high validation score. In our experiments, we use different validation metrics. In particular, we're using MSE or MAPE for tile collection, and OPA (<a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric\" target=\"_blank\">https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric</a>) for layout collections. You may want to try those.</p>",
      "rawMarkdown": "Thank you for pointing this out. I skimmed through your notebook, and it looks like you use the final evaluation metric for validation score. The top-k score for the tile collection is definitely not suitable for validation because your model can just get lucky and happen to pick one good config in the top k candidates and have high validation score. In our experiments, we use different validation metrics. In particular, we're using MSE or MAPE for tile collection, and OPA (https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric) for layout collections. You may want to try those.",
      "votes": null
    },
    {
      "id": "2419144",
      "postDate": "09/01/2023 17:58:57",
      "content": "<p>Oooh Thanks a ton</p>",
      "rawMarkdown": "Oooh Thanks a ton",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2419099,
      "author_name": "mangpophothilimthana",
      "author_url": "",
      "post_date": "09/01/2023 17:22:12",
      "content": "<p>Thank you for pointing this out. I skimmed through your notebook, and it looks like you use the final evaluation metric for validation score. The top-k score for the tile collection is definitely not suitable for validation because your model can just get lucky and happen to pick one good config in the top k candidates and have high validation score. In our experiments, we use different validation metrics. In particular, we're using MSE or MAPE for tile collection, and OPA (<a href=\"https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric\" target=\"_blank\">https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric</a>) for layout collections. You may want to try those.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2419144,
          "author_name": "jainam213",
          "author_url": "",
          "post_date": "09/01/2023 17:58:57",
          "content": "<p>Oooh Thanks a ton</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2418136": "Playing with hyperparams my notebooks Val score increases like crazy but seems in overfiting and this doesn't generalize to lb\nhttps://www.kaggle.com/code/jainam213/0-97val0-12lb\nWe need a more robust validation dataset to correlate cv and lb closely",
    "2419099": "Thank you for pointing this out. I skimmed through your notebook, and it looks like you use the final evaluation metric for validation score. The top-k score for the tile collection is definitely not suitable for validation because your model can just get lucky and happen to pick one good config in the top k candidates and have high validation score. In our experiments, we use different validation metrics. In particular, we're using MSE or MAPE for tile collection, and OPA (https://www.tensorflow.org/ranking/api_docs/python/tfr/keras/metrics/OPAMetric) for layout collections. You may want to try those.",
    "2419144": "Oooh Thanks a ton"
  },
  "source": "meta"
}