{
  "id": 340945,
  "title": "difference between result on given data and test data",
  "url": "/competitions/amex-default-prediction/discussion/340945",
  "author_name": "",
  "post_date": "2022-07-31T16:04:24.279179500Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,<br>\nI divide the given dataset between train data and validation data. I fit my model, and find 0.92 with the function given for the challenge to evaluate the model on validation dataset (not use to train).</p>\n<p>But when i do my submissions i get 0.4 score.</p>\n<p>The only hypothesis i have it is that test data are totaly different than given data.</p>\n<p>but lot of poeple have higher score. That mean that they have not the probleme.</p>\n<p>am i the only one in this situation?</p>",
  "messages": [
    {
      "id": "1878784",
      "postDate": "07/31/2022 16:04:24",
      "content": "<p>Hi,<br>\nI divide the given dataset between train data and validation data. I fit my model, and find 0.92 with the function given for the challenge to evaluate the model on validation dataset (not use to train).</p>\n<p>But when i do my submissions i get 0.4 score.</p>\n<p>The only hypothesis i have it is that test data are totaly different than given data.</p>\n<p>but lot of poeple have higher score. That mean that they have not the probleme.</p>\n<p>am i the only one in this situation?</p>",
      "rawMarkdown": "Hi,\nI divide the given dataset between train data and validation data. I fit my model, and find 0.92 with the function given for the challenge to evaluate the model on validation dataset (not use to train).\n\n But when i do my submissions i get 0.4 score.\n\nThe only hypothesis i have it is that test data are totaly different than given data.\n\nbut lot of poeple have higher score. That mean that they have not the probleme.\n\nam i the only one in this situation?",
      "votes": null
    },
    {
      "id": "1878790",
      "postDate": "07/31/2022 16:10:33",
      "content": "<p>You have data leak.</p>\n<p>Yes test set is different but not such drastically.</p>\n<p>You need to recheck your pipeline.</p>",
      "rawMarkdown": "You have data leak.\n\nYes test set is different but not such drastically.\n\nYou need to recheck your pipeline.",
      "votes": null
    },
    {
      "id": "1878869",
      "postDate": "07/31/2022 17:21:49",
      "content": "<p>If your validation score is &gt; 0.800, something's wrong with your train-val split. Make sure that your data has only one row per customer before you split. If you have more than one row per customer, you'll get the same customer in train and validation, and that would amount to a data leak.</p>",
      "rawMarkdown": "If your validation score is > 0.800, something's wrong with your train-val split. Make sure that your data has only one row per customer before you split. If you have more than one row per customer, you'll get the same customer in train and validation, and that would amount to a data leak.",
      "votes": null
    },
    {
      "id": "1879117",
      "postDate": "07/31/2022 21:21:27",
      "content": "<p>my validation score is 0.80049, and it still matches the LB well. I would say, \"if the validation score looks not plausible instead of &gt;0.800\".</p>",
      "rawMarkdown": "my validation score is 0.80049, and it still matches the LB well. I would say, \"if the validation score looks not plausible instead of >0.800\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1878790,
      "author_name": "kyakovlev",
      "author_url": "",
      "post_date": "07/31/2022 16:10:33",
      "content": "<p>You have data leak.</p>\n<p>Yes test set is different but not such drastically.</p>\n<p>You need to recheck your pipeline.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1878869,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "07/31/2022 17:21:49",
      "content": "<p>If your validation score is &gt; 0.800, something's wrong with your train-val split. Make sure that your data has only one row per customer before you split. If you have more than one row per customer, you'll get the same customer in train and validation, and that would amount to a data leak.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1879117,
          "author_name": "meli19",
          "author_url": "",
          "post_date": "07/31/2022 21:21:27",
          "content": "<p>my validation score is 0.80049, and it still matches the LB well. I would say, \"if the validation score looks not plausible instead of &gt;0.800\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1878784": "Hi,\nI divide the given dataset between train data and validation data. I fit my model, and find 0.92 with the function given for the challenge to evaluate the model on validation dataset (not use to train).\n\n But when i do my submissions i get 0.4 score.\n\nThe only hypothesis i have it is that test data are totaly different than given data.\n\nbut lot of poeple have higher score. That mean that they have not the probleme.\n\nam i the only one in this situation?",
    "1878790": "You have data leak.\n\nYes test set is different but not such drastically.\n\nYou need to recheck your pipeline.",
    "1878869": "If your validation score is > 0.800, something's wrong with your train-val split. Make sure that your data has only one row per customer before you split. If you have more than one row per customer, you'll get the same customer in train and validation, and that would amount to a data leak.",
    "1879117": "my validation score is 0.80049, and it still matches the LB well. I would say, \"if the validation score looks not plausible instead of >0.800\"."
  },
  "source": "meta"
}