{
  "id": 507559,
  "title": "Is it possible to use 'df_test' for learning?🤔",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/507559",
  "author_name": "SeungWoo-Seo",
  "post_date": "2024-05-26T10:46:51.939000",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>If you look at the script1.py code in the metric hack that currently tops the leaderboard, df_train and df_test are concatenated into one and used as training data for the model. Doesn't this method violate the contest rules? I think this is my first time seeing df_test used for learning.</p>",
  "messages": [
    {
      "id": 2837188,
      "postDate": "2024-05-26T10:46:51.940Z",
      "content": "<p>If you look at the script1.py code in the metric hack that currently tops the leaderboard, df_train and df_test are concatenated into one and used as training data for the model. Doesn't this method violate the contest rules? I think this is my first time seeing df_test used for learning.</p>",
      "rawMarkdown": "If you look at the script1.py code in the metric hack that currently tops the leaderboard, df_train and df_test are concatenated into one and used as training data for the model. Doesn't this method violate the contest rules? I think this is my first time seeing df_test used for learning.",
      "votes": 3
    },
    {
      "id": 2838588,
      "postDate": "2024-05-27T05:59:04.677Z",
      "content": "<p>No, it's not a violation, online training is allowed with any data you are provided with. In this particular situation you probably confused the real target which is not there in the test data, with the made-up pseudo-target used to train a model for the metric hack.</p>",
      "rawMarkdown": "No, it's not a violation, online training is allowed with any data you are provided with. In this particular situation you probably confused the real target which is not there in the test data, with the made-up pseudo-target used to train a model for the metric hack.",
      "replies": [
        {
          "id": 2838847,
          "postDate": "2024-05-27T08:28:00.127Z",
          "content": "<ol>\n<li>COMPETITION ENTRY. in Rules<br>\n\"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"</li>\n</ol>\n<p>Doesn’t this mean that using test data through labeling is impossible??</p>",
          "rawMarkdown": "5. COMPETITION ENTRY. in Rules\n\"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"\n\nDoesn’t this mean that using test data through labeling is impossible??",
          "replies": [
            {
              "id": 2838871,
              "postDate": "2024-05-27T08:40:41.663Z",
              "content": "<p>That's an interesting find, however I don't think \"hand labeling or human prediction\" fits in this case… it may be labeling, but is it hand labeling?</p>",
              "rawMarkdown": "That's an interesting find, however I don't think \"hand labeling or human prediction\" fits in this case... it may be labeling, but is it hand labeling?"
            },
            {
              "id": 2838881,
              "postDate": "2024-05-27T08:51:00.950Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8950082%2F0a4eb10101125c820a5efe756704d38c%2Fscript1.png?generation=1716799695078387&amp;alt=media\" alt=\"script1.py\"></p>\n<p>I don't know exactly what hand labeling or human prediction means here, but I thought it applies here.</p>",
              "rawMarkdown": "![script1.py](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8950082%2F0a4eb10101125c820a5efe756704d38c%2Fscript1.png?generation=1716799695078387&alt=media)\n\nI don't know exactly what hand labeling or human prediction means here, but I thought it applies here."
            },
            {
              "id": 2838902,
              "postDate": "2024-05-27T09:03:44.620Z",
              "content": "<p>Yes, that's the fragment where it takes place, but is it manual labeling? I think \"hand labeling or human prediction\" is when you go example by example and assign a label. The above one on the other hand is a kind of algorithmic processing, not like a look-up table at all. I may be wrong, though. I'm curious what our hosts would say.. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> (if not too busy at this time)</p>",
              "rawMarkdown": "Yes, that's the fragment where it takes place, but is it manual labeling? I think \"hand labeling or human prediction\" is when you go example by example and assign a label. The above one on the other hand is a kind of algorithmic processing, not like a look-up table at all. I may be wrong, though. I'm curious what our hosts would say.. @jetakow (if not too busy at this time)"
            }
          ]
        }
      ]
    },
    {
      "id": 2838197,
      "postDate": "2024-05-26T22:44:42.643Z",
      "content": "<p>If which columns are actually filled in changes over time and/or people wish to pseudo label, I’m not sure how those approaches would be considered cheating. The former I didn’t really have much time to investigate and the latter didn’t work well for me.</p>",
      "rawMarkdown": "If which columns are actually filled in changes over time and/or people wish to pseudo label, I’m not sure how those approaches would be considered cheating. The former I didn’t really have much time to investigate and the latter didn’t work well for me.",
      "replies": [
        {
          "id": 2838855,
          "postDate": "2024-05-27T08:30:51.090Z",
          "content": "<p>According to my understanding of the script1 file, the test data is labeled as 1 and concated with the train data. According to my experiment results, if test data is excluded, the performance drops significantly to 0.4xx.</p>",
          "rawMarkdown": "According to my understanding of the script1 file, the test data is labeled as 1 and concated with the train data. According to my experiment results, if test data is excluded, the performance drops significantly to 0.4xx."
        }
      ]
    },
    {
      "id": 2837578,
      "postDate": "2024-05-26T15:22:32.383Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2838859,
          "postDate": "2024-05-27T08:33:49.277Z",
          "content": "<p>I feel the same way…</p>",
          "rawMarkdown": "I feel the same way..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2838588,
      "author_name": "loh-maa",
      "author_url": "",
      "post_date": "2024-05-27T05:59:04.677000",
      "content": "<p>No, it's not a violation, online training is allowed with any data you are provided with. In this particular situation you probably confused the real target which is not there in the test data, with the made-up pseudo-target used to train a model for the metric hack.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2838847,
          "author_name": "SeungWoo-Seo",
          "author_url": "",
          "post_date": "2024-05-27T08:28:00.127000",
          "content": "<ol>\n<li>COMPETITION ENTRY. in Rules<br>\n\"Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\"</li>\n</ol>\n<p>Doesn’t this mean that using test data through labeling is impossible??</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2838871,
              "author_name": "loh-maa",
              "author_url": "",
              "post_date": "2024-05-27T08:40:41.663000",
              "content": "<p>That's an interesting find, however I don't think \"hand labeling or human prediction\" fits in this case… it may be labeling, but is it hand labeling?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2838881,
              "author_name": "SeungWoo-Seo",
              "author_url": "",
              "post_date": "2024-05-27T08:51:00.950000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8950082%2F0a4eb10101125c820a5efe756704d38c%2Fscript1.png?generation=1716799695078387&amp;alt=media\" alt=\"script1.py\"></p>\n<p>I don't know exactly what hand labeling or human prediction means here, but I thought it applies here.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2838902,
              "author_name": "loh-maa",
              "author_url": "",
              "post_date": "2024-05-27T09:03:44.620000",
              "content": "<p>Yes, that's the fragment where it takes place, but is it manual labeling? I think \"hand labeling or human prediction\" is when you go example by example and assign a label. The above one on the other hand is a kind of algorithmic processing, not like a look-up table at all. I may be wrong, though. I'm curious what our hosts would say.. <a href=\"https://www.kaggle.com/jetakow\" target=\"_blank\">@jetakow</a> (if not too busy at this time)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2838197,
      "author_name": "Rob Freeman",
      "author_url": "",
      "post_date": "2024-05-26T22:44:42.643000",
      "content": "<p>If which columns are actually filled in changes over time and/or people wish to pseudo label, I’m not sure how those approaches would be considered cheating. The former I didn’t really have much time to investigate and the latter didn’t work well for me.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2838855,
          "author_name": "SeungWoo-Seo",
          "author_url": "",
          "post_date": "2024-05-27T08:30:51.090000",
          "content": "<p>According to my understanding of the script1 file, the test data is labeled as 1 and concated with the train data. According to my experiment results, if test data is excluded, the performance drops significantly to 0.4xx.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2837578,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-26T15:22:32.383000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2838859,
          "author_name": "SeungWoo-Seo",
          "author_url": "",
          "post_date": "2024-05-27T08:33:49.277000",
          "content": "<p>I feel the same way…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2837188": "If you look at the script1.py code in the metric hack that currently tops the leaderboard, df_train and df_test are concatenated into one and used as training data for the model. Doesn't this method violate the contest rules? I think this is my first time seeing df_test used for learning.",
    "2838588": "No, it's not a violation, online training is allowed with any data you are provided with. In this particular situation you probably confused the real target which is not there in the test data, with the made-up pseudo-target used to train a model for the metric hack.",
    "2838197": "If which columns are actually filled in changes over time and/or people wish to pseudo label, I’m not sure how those approaches would be considered cheating. The former I didn’t really have much time to investigate and the latter didn’t work well for me.",
    "2837578": ""
  }
}