{
  "id": 74499,
  "title": "feature computation time?",
  "url": "/competitions/PLAsTiCC-2018/discussion/74499",
  "author_name": "",
  "post_date": "2018-12-12T22:48:03.837059500Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have submitted only once but im curious to know what is the processing time for generating features for everyone for train and test.</p>\n\n<p>for training, it takes me around 2mins and a local CV of 0.58 (not submitted yet)\ntest time: still optimizing.</p>",
  "messages": [
    {
      "id": "437990",
      "postDate": "12/12/2018 22:48:03",
      "content": "<p>I have submitted only once but im curious to know what is the processing time for generating features for everyone for train and test.</p>\n\n<p>for training, it takes me around 2mins and a local CV of 0.58 (not submitted yet)\ntest time: still optimizing.</p>",
      "rawMarkdown": "I have submitted only once but im curious to know what is the processing time for generating features for everyone for train and test.\n\nfor training, it takes me around 2mins and a local CV of 0.58 (not submitted yet)\ntest time: still optimizing.",
      "votes": null
    },
    {
      "id": "437992",
      "postDate": "12/12/2018 22:57:30",
      "content": "<p>I am running roughly 1 minute for the training set and 10 hours for the test set.</p>",
      "rawMarkdown": "I am running roughly 1 minute for the training set and 10 hours for the test set.",
      "votes": null
    },
    {
      "id": "437997",
      "postDate": "12/12/2018 23:16:47",
      "content": "<p>features calculated in 1 minute in the training set, take around 6-7 hours to compute and save with kaggle kernels. it's good to incrementally add features to your model in this competition, but as we are reaching the end I would say to trust your CV adding features and bear in mind that some features can stop being useful(and start being harmful) after the addition of stronger ones transmitting equivalent information to your model.</p>",
      "rawMarkdown": "features calculated in 1 minute in the training set, take around 6-7 hours to compute and save with kaggle kernels. it's good to incrementally add features to your model in this competition, but as we are reaching the end I would say to trust your CV adding features and bear in mind that some features can stop being useful(and start being harmful) after the addition of stronger ones transmitting equivalent information to your model.",
      "votes": null
    },
    {
      "id": "438008",
      "postDate": "12/12/2018 23:41:08",
      "content": "<p>depends, periods can take 2 weeks on 24 CPUs, while easy ones can take 8 min -- see <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827</a></p>",
      "rawMarkdown": "depends, periods can take 2 weeks on 24 CPUs, while easy ones can take 8 min -- see https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827",
      "votes": null
    },
    {
      "id": "438067",
      "postDate": "12/13/2018 03:44:30",
      "content": "<p>The feature set in my kernel runs on local desktop take 13s for training set, the whole feature generation and prediction is about 230 min. It takes ~6 hours+ as kaggle's kernel.</p>",
      "rawMarkdown": "The feature set in my kernel runs on local desktop take 13s for training set, the whole feature generation and prediction is about 230 min. It takes ~6 hours+ as kaggle's kernel.",
      "votes": null
    },
    {
      "id": "438109",
      "postDate": "12/13/2018 05:36:06",
      "content": "<p>Almost two weeks</p>",
      "rawMarkdown": "Almost two weeks",
      "votes": null
    },
    {
      "id": "438113",
      "postDate": "12/13/2018 05:52:38",
      "content": "<p>8-10 minutes for a bag (n approx. = 100) of basic stats-based features on the test set using BigQuery, no further preprocessing required. For the more complicated features, you may need to run them longer (even more than a day), but I accelerated one feature extractor from a recommended package 20x using a simple @numba.jit trick. You can always find an interesting feature extractor, go to its source code, and find ways to speed it up, either by vectorizing or compiling via numba/cython. </p>",
      "rawMarkdown": "8-10 minutes for a bag (n approx. = 100) of basic stats-based features on the test set using BigQuery, no further preprocessing required. For the more complicated features, you may need to run them longer (even more than a day), but I accelerated one feature extractor from a recommended package 20x using a simple @numba.jit trick. You can always find an interesting feature extractor, go to its source code, and find ways to speed it up, either by vectorizing or compiling via numba/cython.",
      "votes": null
    },
    {
      "id": "438115",
      "postDate": "12/13/2018 05:59:22",
      "content": "<p>No Kidding! 😂 </p>",
      "rawMarkdown": "No Kidding! 😂",
      "votes": null
    },
    {
      "id": "438132",
      "postDate": "12/13/2018 06:37:19",
      "content": "<p>Of course we reuse cached versions.  With that, best model submission computation (lgb) takes under one hour on my machine.</p>",
      "rawMarkdown": "Of course we reuse cached versions.  With that, best model submission computation (lgb) takes under one hour on my machine.",
      "votes": null
    },
    {
      "id": "438550",
      "postDate": "12/13/2018 21:41:25",
      "content": "<p>Takes around 16 hours from scratch on a quad core laptop. ~150 features are used</p>",
      "rawMarkdown": "Takes around 16 hours from scratch on a quad core laptop. ~150 features are used",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 437992,
      "author_name": "molenr",
      "author_url": "",
      "post_date": "12/12/2018 22:57:30",
      "content": "<p>I am running roughly 1 minute for the training set and 10 hours for the test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437997,
      "author_name": "iprapas",
      "author_url": "",
      "post_date": "12/12/2018 23:16:47",
      "content": "<p>features calculated in 1 minute in the training set, take around 6-7 hours to compute and save with kaggle kernels. it's good to incrementally add features to your model in this competition, but as we are reaching the end I would say to trust your CV adding features and bear in mind that some features can stop being useful(and start being harmful) after the addition of stronger ones transmitting equivalent information to your model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438008,
      "author_name": "blondinka",
      "author_url": "",
      "post_date": "12/12/2018 23:41:08",
      "content": "<p>depends, periods can take 2 weeks on 24 CPUs, while easy ones can take 8 min -- see <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438067,
      "author_name": "cttsai",
      "author_url": "",
      "post_date": "12/13/2018 03:44:30",
      "content": "<p>The feature set in my kernel runs on local desktop take 13s for training set, the whole feature generation and prediction is about 230 min. It takes ~6 hours+ as kaggle's kernel.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438109,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "12/13/2018 05:36:06",
      "content": "<p>Almost two weeks</p>",
      "votes": null,
      "replies": [
        {
          "id": 438115,
          "author_name": "jiweiliu",
          "author_url": "",
          "post_date": "12/13/2018 05:59:22",
          "content": "<p>No Kidding! 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438132,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/13/2018 06:37:19",
          "content": "<p>Of course we reuse cached versions.  With that, best model submission computation (lgb) takes under one hour on my machine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438113,
      "author_name": "mithrillion",
      "author_url": "",
      "post_date": "12/13/2018 05:52:38",
      "content": "<p>8-10 minutes for a bag (n approx. = 100) of basic stats-based features on the test set using BigQuery, no further preprocessing required. For the more complicated features, you may need to run them longer (even more than a day), but I accelerated one feature extractor from a recommended package 20x using a simple @numba.jit trick. You can always find an interesting feature extractor, go to its source code, and find ways to speed it up, either by vectorizing or compiling via numba/cython. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438550,
      "author_name": "ihsansecer",
      "author_url": "",
      "post_date": "12/13/2018 21:41:25",
      "content": "<p>Takes around 16 hours from scratch on a quad core laptop. ~150 features are used</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "437990": "I have submitted only once but im curious to know what is the processing time for generating features for everyone for train and test.\n\nfor training, it takes me around 2mins and a local CV of 0.58 (not submitted yet)\ntest time: still optimizing.",
    "437992": "I am running roughly 1 minute for the training set and 10 hours for the test set.",
    "437997": "features calculated in 1 minute in the training set, take around 6-7 hours to compute and save with kaggle kernels. it's good to incrementally add features to your model in this competition, but as we are reaching the end I would say to trust your CV adding features and bear in mind that some features can stop being useful(and start being harmful) after the addition of stronger ones transmitting equivalent information to your model.",
    "438008": "depends, periods can take 2 weeks on 24 CPUs, while easy ones can take 8 min -- see https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71827",
    "438067": "The feature set in my kernel runs on local desktop take 13s for training set, the whole feature generation and prediction is about 230 min. It takes ~6 hours+ as kaggle's kernel.",
    "438109": "Almost two weeks",
    "438113": "8-10 minutes for a bag (n approx. = 100) of basic stats-based features on the test set using BigQuery, no further preprocessing required. For the more complicated features, you may need to run them longer (even more than a day), but I accelerated one feature extractor from a recommended package 20x using a simple @numba.jit trick. You can always find an interesting feature extractor, go to its source code, and find ways to speed it up, either by vectorizing or compiling via numba/cython.",
    "438115": "No Kidding! 😂",
    "438132": "Of course we reuse cached versions.  With that, best model submission computation (lgb) takes under one hour on my machine.",
    "438550": "Takes around 16 hours from scratch on a quad core laptop. ~150 features are used"
  },
  "source": "meta"
}