{
  "id": 189555,
  "title": "Is feature preprocessing external data?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189555",
  "author_name": "",
  "post_date": "2020-10-07T22:13:03.765292900Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Sorry if this is already answered somewhere here or in a previous competition, but I haven't been able to find a definitive explanation of what is and isn't considered external data that needs to be publicly available to be used.</p>\n<p>In particular I want to know about 2 cases:</p>\n<ol>\n<li><p>Is a preprocessed feature set external data? E.g. if I derive features solely from the competition data, can I import that as a private dataset to save time or would it need to be publicly shared?</p></li>\n<li><p>Is a pre-trained model imported to make predictions external data? This seems like a weird one, because any pre-trained model parameters should be useless without the exactly correct feature set to accompany them. It seems unnecessary to have these be public because they're essentially private no matter what.</p></li>\n</ol>\n<p>From a glance at the 2019 DS Bowl it looks like pre-trained models are more of a norm than preprocessed feature sets, but I'm really unsure. If the spirit of a code competition is making efficiency a big focus, it seems to me like both preprocessed feature sets and pre-trained models go against that. </p>",
  "messages": [
    {
      "id": "1041726",
      "postDate": "10/07/2020 22:13:03",
      "content": "<p>Sorry if this is already answered somewhere here or in a previous competition, but I haven't been able to find a definitive explanation of what is and isn't considered external data that needs to be publicly available to be used.</p>\n<p>In particular I want to know about 2 cases:</p>\n<ol>\n<li><p>Is a preprocessed feature set external data? E.g. if I derive features solely from the competition data, can I import that as a private dataset to save time or would it need to be publicly shared?</p></li>\n<li><p>Is a pre-trained model imported to make predictions external data? This seems like a weird one, because any pre-trained model parameters should be useless without the exactly correct feature set to accompany them. It seems unnecessary to have these be public because they're essentially private no matter what.</p></li>\n</ol>\n<p>From a glance at the 2019 DS Bowl it looks like pre-trained models are more of a norm than preprocessed feature sets, but I'm really unsure. If the spirit of a code competition is making efficiency a big focus, it seems to me like both preprocessed feature sets and pre-trained models go against that. </p>",
      "rawMarkdown": "Sorry if this is already answered somewhere here or in a previous competition, but I haven't been able to find a definitive explanation of what is and isn't considered external data that needs to be publicly available to be used.\n\nIn particular I want to know about 2 cases:\n\n1. Is a preprocessed feature set external data? E.g. if I derive features solely from the competition data, can I import that as a private dataset to save time or would it need to be publicly shared?\n\n2. Is a pre-trained model imported to make predictions external data? This seems like a weird one, because any pre-trained model parameters should be useless without the exactly correct feature set to accompany them. It seems unnecessary to have these be public because they're essentially private no matter what.\n\nFrom a glance at the 2019 DS Bowl it looks like pre-trained models are more of a norm than preprocessed feature sets, but I'm really unsure. If the spirit of a code competition is making efficiency a big focus, it seems to me like both preprocessed feature sets and pre-trained models go against that.",
      "votes": null
    },
    {
      "id": "1041784",
      "postDate": "10/07/2020 23:07:43",
      "content": "<p>It's a moot question since you don't have direct access to the full test set.</p>",
      "rawMarkdown": "It's a moot question since you don't have direct access to the full test set.",
      "votes": null
    },
    {
      "id": "1041796",
      "postDate": "10/07/2020 23:13:04",
      "content": "<p>Don't worry Joe, you can build features and models and upload them as private datasets. Regardless of the whole performance aspect, it makes our life much easier. Having to go through the whole training pipeline each time we submit a kernel would be too painful, as least from my point of view.</p>",
      "rawMarkdown": "Don't worry Joe, you can build features and models and upload them as private datasets. Regardless of the whole performance aspect, it makes our life much easier. Having to go through the whole training pipeline each time we submit a kernel would be too painful, as least from my point of view.",
      "votes": null
    },
    {
      "id": "1041850",
      "postDate": "10/08/2020 00:07:02",
      "content": "<p>Sure, but there's still a lot of feature preparation that can be done with the training data in advance, then appended to the test data batch by batch. For example if I calculate a bunch of statistics about the questions across all the training data I could reasonably use those same statistics as test features. Maybe Ideally I'm updating them live as I process the test data, but even so it's still faster to preload those features from a static file rather than compute them again and again at runtime. </p>\n<p>My normal workflow would be to modularize feature engineering as much as possible into different scripts because it makes it easier to maintain good code and more quickly iterate on model selection. So I'd like to also do that here, but it adds extra overhead if I have to eventually translate everything into an all FE happens at runtime type of situation.</p>",
      "rawMarkdown": "Sure, but there's still a lot of feature preparation that can be done with the training data in advance, then appended to the test data batch by batch. For example if I calculate a bunch of statistics about the questions across all the training data I could reasonably use those same statistics as test features. Maybe Ideally I'm updating them live as I process the test data, but even so it's still faster to preload those features from a static file rather than compute them again and again at runtime. \n\nMy normal workflow would be to modularize feature engineering as much as possible into different scripts because it makes it easier to maintain good code and more quickly iterate on model selection. So I'd like to also do that here, but it adds extra overhead if I have to eventually translate everything into an all FE happens at runtime type of situation.",
      "votes": null
    },
    {
      "id": "1042950",
      "postDate": "10/08/2020 15:16:50",
      "content": "<p>You can upload and use the results of your own work, just like you would for a model trained offline.</p>",
      "rawMarkdown": "You can upload and use the results of your own work, just like you would for a model trained offline.",
      "votes": null
    },
    {
      "id": "1042980",
      "postDate": "10/08/2020 15:37:43",
      "content": "<p>Thanks, that helps!</p>",
      "rawMarkdown": "Thanks, that helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041784,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/07/2020 23:07:43",
      "content": "<p>It's a moot question since you don't have direct access to the full test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041850,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/08/2020 00:07:02",
          "content": "<p>Sure, but there's still a lot of feature preparation that can be done with the training data in advance, then appended to the test data batch by batch. For example if I calculate a bunch of statistics about the questions across all the training data I could reasonably use those same statistics as test features. Maybe Ideally I'm updating them live as I process the test data, but even so it's still faster to preload those features from a static file rather than compute them again and again at runtime. </p>\n<p>My normal workflow would be to modularize feature engineering as much as possible into different scripts because it makes it easier to maintain good code and more quickly iterate on model selection. So I'd like to also do that here, but it adds extra overhead if I have to eventually translate everything into an all FE happens at runtime type of situation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042950,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "10/08/2020 15:16:50",
          "content": "<p>You can upload and use the results of your own work, just like you would for a model trained offline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042980,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/08/2020 15:37:43",
          "content": "<p>Thanks, that helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041796,
      "author_name": "maxhalford",
      "author_url": "",
      "post_date": "10/07/2020 23:13:04",
      "content": "<p>Don't worry Joe, you can build features and models and upload them as private datasets. Regardless of the whole performance aspect, it makes our life much easier. Having to go through the whole training pipeline each time we submit a kernel would be too painful, as least from my point of view.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1041726": "Sorry if this is already answered somewhere here or in a previous competition, but I haven't been able to find a definitive explanation of what is and isn't considered external data that needs to be publicly available to be used.\n\nIn particular I want to know about 2 cases:\n\n1. Is a preprocessed feature set external data? E.g. if I derive features solely from the competition data, can I import that as a private dataset to save time or would it need to be publicly shared?\n\n2. Is a pre-trained model imported to make predictions external data? This seems like a weird one, because any pre-trained model parameters should be useless without the exactly correct feature set to accompany them. It seems unnecessary to have these be public because they're essentially private no matter what.\n\nFrom a glance at the 2019 DS Bowl it looks like pre-trained models are more of a norm than preprocessed feature sets, but I'm really unsure. If the spirit of a code competition is making efficiency a big focus, it seems to me like both preprocessed feature sets and pre-trained models go against that.",
    "1041784": "It's a moot question since you don't have direct access to the full test set.",
    "1041796": "Don't worry Joe, you can build features and models and upload them as private datasets. Regardless of the whole performance aspect, it makes our life much easier. Having to go through the whole training pipeline each time we submit a kernel would be too painful, as least from my point of view.",
    "1041850": "Sure, but there's still a lot of feature preparation that can be done with the training data in advance, then appended to the test data batch by batch. For example if I calculate a bunch of statistics about the questions across all the training data I could reasonably use those same statistics as test features. Maybe Ideally I'm updating them live as I process the test data, but even so it's still faster to preload those features from a static file rather than compute them again and again at runtime. \n\nMy normal workflow would be to modularize feature engineering as much as possible into different scripts because it makes it easier to maintain good code and more quickly iterate on model selection. So I'd like to also do that here, but it adds extra overhead if I have to eventually translate everything into an all FE happens at runtime type of situation.",
    "1042950": "You can upload and use the results of your own work, just like you would for a model trained offline.",
    "1042980": "Thanks, that helps!"
  },
  "source": "meta"
}