{
  "id": 470875,
  "title": "Do ensemble solutions satisfy the competition's code requirements?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470875",
  "author_name": "",
  "post_date": "2024-01-26T01:37:12.651088Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm curious if ensemble solutions are valid entries into the competition. Specifically, when using trained models from other public notebooks. The code requirements state the notebook run time must be less than 9 hours. Are the run times of imported notebook models excluded from this limitation?</p>\n<p>The code requirements also states \"Freely &amp; publicly available external data is allowed, including pre-trained models.\" When a Kaggler publishes a public notebook, does the trained model within it (of within a dataset) qualify as a freely and publicly available pre-trained model?</p>\n<p>If you have experience with similar competitions, please share your insights.</p>",
  "messages": [
    {
      "id": "2620195",
      "postDate": "01/26/2024 01:37:12",
      "content": "<p>I'm curious if ensemble solutions are valid entries into the competition. Specifically, when using trained models from other public notebooks. The code requirements state the notebook run time must be less than 9 hours. Are the run times of imported notebook models excluded from this limitation?</p>\n<p>The code requirements also states \"Freely &amp; publicly available external data is allowed, including pre-trained models.\" When a Kaggler publishes a public notebook, does the trained model within it (of within a dataset) qualify as a freely and publicly available pre-trained model?</p>\n<p>If you have experience with similar competitions, please share your insights.</p>",
      "rawMarkdown": "I'm curious if ensemble solutions are valid entries into the competition. Specifically, when using trained models from other public notebooks. The code requirements state the notebook run time must be less than 9 hours. Are the run times of imported notebook models excluded from this limitation?\n\nThe code requirements also states \"Freely & publicly available external data is allowed, including pre-trained models.\" When a Kaggler publishes a public notebook, does the trained model within it (of within a dataset) qualify as a freely and publicly available pre-trained model?\n\nIf you have experience with similar competitions, please share your insights.",
      "votes": null
    },
    {
      "id": "2620207",
      "postDate": "01/26/2024 01:56:39",
      "content": "<p>I'm also curious whether this time limit refers to 'training + inference.' If the answer is yes, then training all the models in a single notebook in order to ensemble afterwards, may not be feasible due to kaggle memory constraints.</p>",
      "rawMarkdown": "I'm also curious whether this time limit refers to 'training + inference.' If the answer is yes, then training all the models in a single notebook in order to ensemble afterwards, may not be feasible due to kaggle memory constraints.",
      "votes": null
    },
    {
      "id": "2620252",
      "postDate": "01/26/2024 02:47:14",
      "content": "<p>In code competitions, kagglers usually divide notebook into training part and inference part. Then submit the inference notebook to save time. So this time limit may refer to inference.</p>",
      "rawMarkdown": "In code competitions, kagglers usually divide notebook into training part and inference part. Then submit the inference notebook to save time. So this time limit may refer to inference.",
      "votes": null
    },
    {
      "id": "2620275",
      "postDate": "01/26/2024 03:17:21",
      "content": "<p>The time limit only refers to inference. We can train models offline for as long as we want. Then upload the model weights to Kaggle dataset. Then our inference notebook loads the models and must run less than 9 hours when submitted to Kaggle test LB data.</p>",
      "rawMarkdown": "The time limit only refers to inference. We can train models offline for as long as we want. Then upload the model weights to Kaggle dataset. Then our inference notebook loads the models and must run less than 9 hours when submitted to Kaggle test LB data.",
      "votes": null
    },
    {
      "id": "2621042",
      "postDate": "01/26/2024 14:41:40",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Does it matter if the dataset is private or public?</p>",
      "rawMarkdown": "cdeotte Does it matter if the dataset is private or public?",
      "votes": null
    },
    {
      "id": "2621055",
      "postDate": "01/26/2024 14:54:16",
      "content": "<p>No. Each of our solution model weights will be private. This is what happens in every Kaggle code competition.</p>\n<p>When the competition talks about \"datasets must be public\", they are not referring to trained model weights. They are referring to any data we use to train our models. For example, let's say that I find a dataset on the internet which contains EEG waveforms and I use that dataset to train my models (in addition to the dataset that Kaggle has provided). Then that public internet dataset needs to be available to everyone. (And i don't need to advertise nor make a copy available. Regarding datasets, it is enough if it is possible for everyone to search and find it themselves and download it themselves).</p>",
      "rawMarkdown": "No. Each of our solution model weights will be private. This is what happens in every Kaggle code competition.\n\nWhen the competition talks about \"datasets must be public\", they are not referring to trained model weights. They are referring to any data we use to train our models. For example, let's say that I find a dataset on the internet which contains EEG waveforms and I use that dataset to train my models (in addition to the dataset that Kaggle has provided). Then that public internet dataset needs to be available to everyone. (And i don't need to advertise nor make a copy available. Regarding datasets, it is enough if it is possible for everyone to search and find it themselves and download it themselves).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2620207,
      "author_name": "yantxx",
      "author_url": "",
      "post_date": "01/26/2024 01:56:39",
      "content": "<p>I'm also curious whether this time limit refers to 'training + inference.' If the answer is yes, then training all the models in a single notebook in order to ensemble afterwards, may not be feasible due to kaggle memory constraints.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2620252,
          "author_name": "zrongchu",
          "author_url": "",
          "post_date": "01/26/2024 02:47:14",
          "content": "<p>In code competitions, kagglers usually divide notebook into training part and inference part. Then submit the inference notebook to save time. So this time limit may refer to inference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2620275,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/26/2024 03:17:21",
      "content": "<p>The time limit only refers to inference. We can train models offline for as long as we want. Then upload the model weights to Kaggle dataset. Then our inference notebook loads the models and must run less than 9 hours when submitted to Kaggle test LB data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2621042,
          "author_name": "seanbearden",
          "author_url": "",
          "post_date": "01/26/2024 14:41:40",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Does it matter if the dataset is private or public?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2621055,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/26/2024 14:54:16",
              "content": "<p>No. Each of our solution model weights will be private. This is what happens in every Kaggle code competition.</p>\n<p>When the competition talks about \"datasets must be public\", they are not referring to trained model weights. They are referring to any data we use to train our models. For example, let's say that I find a dataset on the internet which contains EEG waveforms and I use that dataset to train my models (in addition to the dataset that Kaggle has provided). Then that public internet dataset needs to be available to everyone. (And i don't need to advertise nor make a copy available. Regarding datasets, it is enough if it is possible for everyone to search and find it themselves and download it themselves).</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2620195": "I'm curious if ensemble solutions are valid entries into the competition. Specifically, when using trained models from other public notebooks. The code requirements state the notebook run time must be less than 9 hours. Are the run times of imported notebook models excluded from this limitation?\n\nThe code requirements also states \"Freely & publicly available external data is allowed, including pre-trained models.\" When a Kaggler publishes a public notebook, does the trained model within it (of within a dataset) qualify as a freely and publicly available pre-trained model?\n\nIf you have experience with similar competitions, please share your insights.",
    "2620207": "I'm also curious whether this time limit refers to 'training + inference.' If the answer is yes, then training all the models in a single notebook in order to ensemble afterwards, may not be feasible due to kaggle memory constraints.",
    "2620252": "In code competitions, kagglers usually divide notebook into training part and inference part. Then submit the inference notebook to save time. So this time limit may refer to inference.",
    "2620275": "The time limit only refers to inference. We can train models offline for as long as we want. Then upload the model weights to Kaggle dataset. Then our inference notebook loads the models and must run less than 9 hours when submitted to Kaggle test LB data.",
    "2621042": "cdeotte Does it matter if the dataset is private or public?",
    "2621055": "No. Each of our solution model weights will be private. This is what happens in every Kaggle code competition.\n\nWhen the competition talks about \"datasets must be public\", they are not referring to trained model weights. They are referring to any data we use to train our models. For example, let's say that I find a dataset on the internet which contains EEG waveforms and I use that dataset to train my models (in addition to the dataset that Kaggle has provided). Then that public internet dataset needs to be available to everyone. (And i don't need to advertise nor make a copy available. Regarding datasets, it is enough if it is possible for everyone to search and find it themselves and download it themselves)."
  },
  "source": "meta"
}