{
  "id": 453760,
  "title": "usage of data, e.g. train on \"valid+train\" or pretrain on \"test\"?",
  "url": "/competitions/predict-ai-model-runtime/discussion/453760",
  "author_name": "",
  "post_date": "2023-11-07T15:50:53.243671700Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>since this is an academic competition, i was wondering do we have to strictly follow the data split given, i.e. training is only allowed on \"train split\".</p>\n<p>Or, are we allowed to:</p>\n<ol>\n<li>train on \"valid+train splits\" </li>\n<li>pretrain on \"test split\"?</li>\n<li>pesudo-label on \"test split\"?</li>\n</ol>",
  "messages": [
    {
      "id": "2516292",
      "postDate": "11/07/2023 15:50:53",
      "content": "<p>since this is an academic competition, i was wondering do we have to strictly follow the data split given, i.e. training is only allowed on \"train split\".</p>\n<p>Or, are we allowed to:</p>\n<ol>\n<li>train on \"valid+train splits\" </li>\n<li>pretrain on \"test split\"?</li>\n<li>pesudo-label on \"test split\"?</li>\n</ol>",
      "rawMarkdown": "since this is an academic competition, i was wondering do we have to strictly follow the data split given, i.e. training is only allowed on \"train split\".\n\nOr, are we allowed to:\n1. train on \"valid+train splits\" \n2. pretrain on \"test split\"?\n3. pesudo-label on \"test split\"?",
      "votes": null
    },
    {
      "id": "2517373",
      "postDate": "11/08/2023 12:35:01",
      "content": "<h2>EDIT: Yes, you can use the data however you want!</h2>\n<p>The common trend in ML is to train only on training set but use the validation set for selection of hyperparameters. I see very few papers also using the validation for training (i.e., obtaining gradients), but they would be at the risk of making their results not comparable to other papers. Nonetheless, for this competition, you can use anything you'd like.</p>",
      "rawMarkdown": "EDIT: Yes, you can use the data however you want!\n--------------------------------------------------\nThe common trend in ML is to train only on training set but use the validation set for selection of hyperparameters. I see very few papers also using the validation for training (i.e., obtaining gradients), but they would be at the risk of making their results not comparable to other papers. Nonetheless, for this competition, you can use anything you'd like.",
      "votes": null
    },
    {
      "id": "2517558",
      "postDate": "11/08/2023 15:16:20",
      "content": "<p>\"see very few papers  ….  they would be at the risk of making their results not comparable to other papers.\"</p>\n<p>thanks for the verification, i understand.</p>\n<p>actually for almost all kaggle competitions, we are allowed to use the test data in unsupervised way.(that is why i asked the is question). Kaggle competitions are actually different from academic benchmarks.</p>\n<p>you want want to pin this thread to avoid confusion later.</p>\n<hr>\n<p>note: if there are other rules, you may also want to carify, e.g.<br>\ntrain = nlp+xla (join dataset to train a model)<br>\ntrain = xla:default + xla:random<br>\n…</p>\n<p>using test results from one split  to post/pre process results from another test split<br>\n(e.g. xla:default:test + xla:random:test)</p>",
      "rawMarkdown": "\"see very few papers  ....  they would be at the risk of making their results not comparable to other papers.\"\n\nthanks for the verification, i understand.\n\nactually for almost all kaggle competitions, we are allowed to use the test data in unsupervised way.(that is why i asked the is question). Kaggle competitions are actually different from academic benchmarks.\n\nyou want want to pin this thread to avoid confusion later.\n\n---\n\nnote: if there are other rules, you may also want to carify, e.g.\ntrain = nlp+xla (join dataset to train a model)\ntrain = xla:default + xla:random\n...\n\nusing test results from one split  to post/pre process results from another test split\n(e.g. xla:default:test + xla:random:test)",
      "votes": null
    },
    {
      "id": "2517685",
      "postDate": "11/08/2023 17:20:38",
      "content": "<p>At this point in the competition, we don't want to impose any additional rules, so I think we will let the participants do whatever they want with the data. However, keep in mind that using validation data for training may risk overfitting the model, and it may hurt the test scores. I'm not sure if using test data in an unsupervised fashion would help much because the amount of test data is not that large compared to training data.</p>",
      "rawMarkdown": "At this point in the competition, we don't want to impose any additional rules, so I think we will let the participants do whatever they want with the data. However, keep in mind that using validation data for training may risk overfitting the model, and it may hurt the test scores. I'm not sure if using test data in an unsupervised fashion would help much because the amount of test data is not that large compared to training data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2517373,
      "author_name": "samihaija",
      "author_url": "",
      "post_date": "11/08/2023 12:35:01",
      "content": "<h2>EDIT: Yes, you can use the data however you want!</h2>\n<p>The common trend in ML is to train only on training set but use the validation set for selection of hyperparameters. I see very few papers also using the validation for training (i.e., obtaining gradients), but they would be at the risk of making their results not comparable to other papers. Nonetheless, for this competition, you can use anything you'd like.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2517558,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/08/2023 15:16:20",
          "content": "<p>\"see very few papers  ….  they would be at the risk of making their results not comparable to other papers.\"</p>\n<p>thanks for the verification, i understand.</p>\n<p>actually for almost all kaggle competitions, we are allowed to use the test data in unsupervised way.(that is why i asked the is question). Kaggle competitions are actually different from academic benchmarks.</p>\n<p>you want want to pin this thread to avoid confusion later.</p>\n<hr>\n<p>note: if there are other rules, you may also want to carify, e.g.<br>\ntrain = nlp+xla (join dataset to train a model)<br>\ntrain = xla:default + xla:random<br>\n…</p>\n<p>using test results from one split  to post/pre process results from another test split<br>\n(e.g. xla:default:test + xla:random:test)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2517685,
              "author_name": "mangpophothilimthana",
              "author_url": "",
              "post_date": "11/08/2023 17:20:38",
              "content": "<p>At this point in the competition, we don't want to impose any additional rules, so I think we will let the participants do whatever they want with the data. However, keep in mind that using validation data for training may risk overfitting the model, and it may hurt the test scores. I'm not sure if using test data in an unsupervised fashion would help much because the amount of test data is not that large compared to training data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2516292": "since this is an academic competition, i was wondering do we have to strictly follow the data split given, i.e. training is only allowed on \"train split\".\n\nOr, are we allowed to:\n1. train on \"valid+train splits\" \n2. pretrain on \"test split\"?\n3. pesudo-label on \"test split\"?",
    "2517373": "EDIT: Yes, you can use the data however you want!\n--------------------------------------------------\nThe common trend in ML is to train only on training set but use the validation set for selection of hyperparameters. I see very few papers also using the validation for training (i.e., obtaining gradients), but they would be at the risk of making their results not comparable to other papers. Nonetheless, for this competition, you can use anything you'd like.",
    "2517558": "\"see very few papers  ....  they would be at the risk of making their results not comparable to other papers.\"\n\nthanks for the verification, i understand.\n\nactually for almost all kaggle competitions, we are allowed to use the test data in unsupervised way.(that is why i asked the is question). Kaggle competitions are actually different from academic benchmarks.\n\nyou want want to pin this thread to avoid confusion later.\n\n---\n\nnote: if there are other rules, you may also want to carify, e.g.\ntrain = nlp+xla (join dataset to train a model)\ntrain = xla:default + xla:random\n...\n\nusing test results from one split  to post/pre process results from another test split\n(e.g. xla:default:test + xla:random:test)",
    "2517685": "At this point in the competition, we don't want to impose any additional rules, so I think we will let the participants do whatever they want with the data. However, keep in mind that using validation data for training may risk overfitting the model, and it may hurt the test scores. I'm not sure if using test data in an unsupervised fashion would help much because the amount of test data is not that large compared to training data."
  },
  "source": "meta"
}