{
  "id": 403107,
  "title": "How do you check if a new feature is useful efficiently ?",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/403107",
  "author_name": "",
  "post_date": "2023-04-21T07:15:28.193029400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>HI all ,</p>\n<p>I wonder how do you check if a new feature is useful efficiently.<br>\nIn the competition  , It takes me ~2days to train a model from scratch to baseline score and it is unacceptable.<br>\nany efficient way(except gpu++) to speed up the flow ?</p>",
  "messages": [
    {
      "id": "2229235",
      "postDate": "04/21/2023 07:15:28",
      "content": "<p>HI all ,</p>\n<p>I wonder how do you check if a new feature is useful efficiently.<br>\nIn the competition  , It takes me ~2days to train a model from scratch to baseline score and it is unacceptable.<br>\nany efficient way(except gpu++) to speed up the flow ?</p>",
      "rawMarkdown": "HI all ,\n \nI wonder how do you check if a new feature is useful efficiently.\nIn the competition  , It takes me ~2days to train a model from scratch to baseline score and it is unacceptable.\nany efficient way(except gpu++) to speed up the flow ?",
      "votes": null
    },
    {
      "id": "2229248",
      "postDate": "04/21/2023 07:28:14",
      "content": "<p>The gold standard would be to use CV but it would have been unfeasible with this dataset as you mentioned</p>\n<p>In this competition, I used the profile of the train loss as a proxy. If train loss is falling fast or with a more preferable profile, then it may give you an indication within a few hours rather than days. You often see this in research papers with large datasets, where they compare train loss curves between optimisers/schedulers etc.</p>\n<p>But I think the false positive rate of this method is high. I had quite a few cases where I thought train loss looked \"better\" but this didn't necessarily translate to a better valid loss or metric</p>",
      "rawMarkdown": "The gold standard would be to use CV but it would have been unfeasible with this dataset as you mentioned\n\nIn this competition, I used the profile of the train loss as a proxy. If train loss is falling fast or with a more preferable profile, then it may give you an indication within a few hours rather than days. You often see this in research papers with large datasets, where they compare train loss curves between optimisers/schedulers etc.\n\nBut I think the false positive rate of this method is high. I had quite a few cases where I thought train loss looked \"better\" but this didn't necessarily translate to a better valid loss or metric",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2229248,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "04/21/2023 07:28:14",
      "content": "<p>The gold standard would be to use CV but it would have been unfeasible with this dataset as you mentioned</p>\n<p>In this competition, I used the profile of the train loss as a proxy. If train loss is falling fast or with a more preferable profile, then it may give you an indication within a few hours rather than days. You often see this in research papers with large datasets, where they compare train loss curves between optimisers/schedulers etc.</p>\n<p>But I think the false positive rate of this method is high. I had quite a few cases where I thought train loss looked \"better\" but this didn't necessarily translate to a better valid loss or metric</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2229235": "HI all ,\n \nI wonder how do you check if a new feature is useful efficiently.\nIn the competition  , It takes me ~2days to train a model from scratch to baseline score and it is unacceptable.\nany efficient way(except gpu++) to speed up the flow ?",
    "2229248": "The gold standard would be to use CV but it would have been unfeasible with this dataset as you mentioned\n\nIn this competition, I used the profile of the train loss as a proxy. If train loss is falling fast or with a more preferable profile, then it may give you an indication within a few hours rather than days. You often see this in research papers with large datasets, where they compare train loss curves between optimisers/schedulers etc.\n\nBut I think the false positive rate of this method is high. I had quite a few cases where I thought train loss looked \"better\" but this didn't necessarily translate to a better valid loss or metric"
  },
  "source": "meta"
}