{
  "id": 76962,
  "title": "Word to the wise",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/76962",
  "author_name": "",
  "post_date": "2019-01-08T10:05:32.626866100Z",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n\n<p>I just wanted to give a quick tip by saying that a lot of conventional wisdom might have to be dropped here. For example I've found that dropping the <code>phase</code> variable gives me a boost on the leaderboard but decreases my local score. I've also found that using a learning rate superior to 1 with LightGBM works better for the leaderboard as well as for my local score.</p>\n\n<p>These two insights might be specific to my pipeline so take them with a grain of salt. The idea I want to convey is that because the training and test sets are so different and so small, you have to think outside the box.</p>\n\n<p>By the way, as I said I'm using LightGBM with a small set of custom features. I'm really curious to know if people around me on the leaderboard are using neural networks.</p>",
  "messages": [
    {
      "id": "452170",
      "postDate": "01/08/2019 10:05:32",
      "content": "<p>Hello everyone,</p>\n\n<p>I just wanted to give a quick tip by saying that a lot of conventional wisdom might have to be dropped here. For example I've found that dropping the <code>phase</code> variable gives me a boost on the leaderboard but decreases my local score. I've also found that using a learning rate superior to 1 with LightGBM works better for the leaderboard as well as for my local score.</p>\n\n<p>These two insights might be specific to my pipeline so take them with a grain of salt. The idea I want to convey is that because the training and test sets are so different and so small, you have to think outside the box.</p>\n\n<p>By the way, as I said I'm using LightGBM with a small set of custom features. I'm really curious to know if people around me on the leaderboard are using neural networks.</p>",
      "rawMarkdown": "Hello everyone,\n\nI just wanted to give a quick tip by saying that a lot of conventional wisdom might have to be dropped here. For example I've found that dropping the `phase` variable gives me a boost on the leaderboard but decreases my local score. I've also found that using a learning rate superior to 1 with LightGBM works better for the leaderboard as well as for my local score.\n\nThese two insights might be specific to my pipeline so take them with a grain of salt. The idea I want to convey is that because the training and test sets are so different and so small, you have to think outside the box.\n\nBy the way, as I said I'm using LightGBM with a small set of custom features. I'm really curious to know if people around me on the leaderboard are using neural networks.",
      "votes": null
    },
    {
      "id": "452203",
      "postDate": "01/08/2019 10:55:22",
      "content": "<p>I'm curious what aspect you are referring to in \"because the training and test sets are so different\", since the organiser said they have similar distributions. At least that's how I interpreted it.</p>",
      "rawMarkdown": "I'm curious what aspect you are referring to in \"because the training and test sets are so different\", since the organiser said they have similar distributions. At least that's how I interpreted it.",
      "votes": null
    },
    {
      "id": "452210",
      "postDate": "01/08/2019 11:10:42",
      "content": "<p><a href=\"/its7171\">@its7171</a> showed they are different\n<a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">https://www.kaggle.com/its7171/train-vs-test-analisys</a></p>",
      "rawMarkdown": "its7171 showed they are different\nhttps://www.kaggle.com/its7171/train-vs-test-analisys",
      "votes": null
    },
    {
      "id": "452215",
      "postDate": "01/08/2019 11:15:18",
      "content": "<p>Indeed I was referring to that thread.</p>",
      "rawMarkdown": "Indeed I was referring to that thread.",
      "votes": null
    },
    {
      "id": "452269",
      "postDate": "01/08/2019 13:29:51",
      "content": "<p>ive been using lgbm but local mcc has been wildly different to lb score by a factor of 0.3 have you had this problem as well?</p>",
      "rawMarkdown": "ive been using lgbm but local mcc has been wildly different to lb score by a factor of 0.3 have you had this problem as well?",
      "votes": null
    },
    {
      "id": "452319",
      "postDate": "01/08/2019 15:06:58",
      "content": "<p>By \"factor\" I assume you mean \"difference\". I have a difference of around 0.1 and it gets lower the better the score. </p>",
      "rawMarkdown": "By \"factor\" I assume you mean \"difference\". I have a difference of around 0.1 and it gets lower the better the score.",
      "votes": null
    },
    {
      "id": "452327",
      "postDate": "01/08/2019 15:20:50",
      "content": "<blockquote>\n  <p>By the way, as I said I'm using LightGBM with a small set of custom features.</p>\n</blockquote>\n\n<p>Is Plasticc experience useful here? ;)</p>",
      "rawMarkdown": "&gt; By the way, as I said I'm using LightGBM with a small set of custom features.\n\nIs Plasticc experience useful here? ;)",
      "votes": null
    },
    {
      "id": "452359",
      "postDate": "01/08/2019 16:07:17",
      "content": "<p>To be honest no! I've mostly been re-implementing Tomas Vantuch's paper.</p>",
      "rawMarkdown": "To be honest no! I've mostly been re-implementing Tomas Vantuch's paper.",
      "votes": null
    },
    {
      "id": "452391",
      "postDate": "01/08/2019 17:08:22",
      "content": "<p>this paper: A novel approach of partial discharges detection in a real environment ?</p>\n\n<p>edit: nevermind, I guess it is the paper shared by Tomas in his welcome thread.</p>",
      "rawMarkdown": "this paper: A novel approach of partial discharges detection in a real environment ?\n\nedit: nevermind, I guess it is the paper shared by Tomas in his welcome thread.",
      "votes": null
    },
    {
      "id": "452396",
      "postDate": "01/08/2019 17:13:18",
      "content": "<p>Yes that one, and I'm also using his PhD dissertation in parallel as it contains some complementary details</p>",
      "rawMarkdown": "Yes that one, and I'm also using his PhD dissertation in parallel as it contains some complementary details",
      "votes": null
    },
    {
      "id": "452419",
      "postDate": "01/08/2019 17:46:30",
      "content": "<p>The distributions might be similar. However, the signals might be drastically different. If we think about it from an operational point of view, then it makes sense because a grid with an equal amount of failure to working condition is horrible and if the failure signal would be the same, then it could be reduced to similar errors. In such a case it would be rather stupid not find a solution to very similar reasons.</p>",
      "rawMarkdown": "The distributions might be similar. However, the signals might be drastically different. If we think about it from an operational point of view, then it makes sense because a grid with an equal amount of failure to working condition is horrible and if the failure signal would be the same, then it could be reduced to similar errors. In such a case it would be rather stupid not find a solution to very similar reasons.",
      "votes": null
    },
    {
      "id": "452448",
      "postDate": "01/08/2019 18:31:21",
      "content": "<p>I bet Mike is using NNs ;)</p>",
      "rawMarkdown": "I bet Mike is using NNs ;)",
      "votes": null
    },
    {
      "id": "453447",
      "postDate": "01/10/2019 08:00:05",
      "content": "<p>Could well be! Also at first @ONODERA's team name was something like \"300 layer benchmark\", he changed since then.</p>",
      "rawMarkdown": "Could well be! Also at first @ONODERA's team name was something like \"300 layer benchmark\", he changed since then.",
      "votes": null
    },
    {
      "id": "470820",
      "postDate": "02/13/2019 16:09:33",
      "content": "<p>Thanks for sharing Max!</p>",
      "rawMarkdown": "Thanks for sharing Max!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 452203,
      "author_name": "sangxia",
      "author_url": "",
      "post_date": "01/08/2019 10:55:22",
      "content": "<p>I'm curious what aspect you are referring to in \"because the training and test sets are so different\", since the organiser said they have similar distributions. At least that's how I interpreted it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 452210,
          "author_name": "corner200",
          "author_url": "",
          "post_date": "01/08/2019 11:10:42",
          "content": "<p><a href=\"/its7171\">@its7171</a> showed they are different\n<a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">https://www.kaggle.com/its7171/train-vs-test-analisys</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452215,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "01/08/2019 11:15:18",
          "content": "<p>Indeed I was referring to that thread.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452419,
          "author_name": "simonwenkel",
          "author_url": "",
          "post_date": "01/08/2019 17:46:30",
          "content": "<p>The distributions might be similar. However, the signals might be drastically different. If we think about it from an operational point of view, then it makes sense because a grid with an equal amount of failure to working condition is horrible and if the failure signal would be the same, then it could be reduced to similar errors. In such a case it would be rather stupid not find a solution to very similar reasons.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452269,
      "author_name": "mab270",
      "author_url": "",
      "post_date": "01/08/2019 13:29:51",
      "content": "<p>ive been using lgbm but local mcc has been wildly different to lb score by a factor of 0.3 have you had this problem as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 452319,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "01/08/2019 15:06:58",
          "content": "<p>By \"factor\" I assume you mean \"difference\". I have a difference of around 0.1 and it gets lower the better the score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452327,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/08/2019 15:20:50",
      "content": "<blockquote>\n  <p>By the way, as I said I'm using LightGBM with a small set of custom features.</p>\n</blockquote>\n\n<p>Is Plasticc experience useful here? ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 452359,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "01/08/2019 16:07:17",
          "content": "<p>To be honest no! I've mostly been re-implementing Tomas Vantuch's paper.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452391,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "01/08/2019 17:08:22",
          "content": "<p>this paper: A novel approach of partial discharges detection in a real environment ?</p>\n\n<p>edit: nevermind, I guess it is the paper shared by Tomas in his welcome thread.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452396,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "01/08/2019 17:13:18",
          "content": "<p>Yes that one, and I'm also using his PhD dissertation in parallel as it contains some complementary details</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452448,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/08/2019 18:31:21",
      "content": "<p>I bet Mike is using NNs ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 453447,
          "author_name": "maxhalford",
          "author_url": "",
          "post_date": "01/10/2019 08:00:05",
          "content": "<p>Could well be! Also at first @ONODERA's team name was something like \"300 layer benchmark\", he changed since then.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 470820,
      "author_name": "mtodisco10",
      "author_url": "",
      "post_date": "02/13/2019 16:09:33",
      "content": "<p>Thanks for sharing Max!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "452170": "Hello everyone,\n\nI just wanted to give a quick tip by saying that a lot of conventional wisdom might have to be dropped here. For example I've found that dropping the `phase` variable gives me a boost on the leaderboard but decreases my local score. I've also found that using a learning rate superior to 1 with LightGBM works better for the leaderboard as well as for my local score.\n\nThese two insights might be specific to my pipeline so take them with a grain of salt. The idea I want to convey is that because the training and test sets are so different and so small, you have to think outside the box.\n\nBy the way, as I said I'm using LightGBM with a small set of custom features. I'm really curious to know if people around me on the leaderboard are using neural networks.",
    "452203": "I'm curious what aspect you are referring to in \"because the training and test sets are so different\", since the organiser said they have similar distributions. At least that's how I interpreted it.",
    "452210": "its7171 showed they are different\nhttps://www.kaggle.com/its7171/train-vs-test-analisys",
    "452215": "Indeed I was referring to that thread.",
    "452269": "ive been using lgbm but local mcc has been wildly different to lb score by a factor of 0.3 have you had this problem as well?",
    "452319": "By \"factor\" I assume you mean \"difference\". I have a difference of around 0.1 and it gets lower the better the score.",
    "452327": "&gt; By the way, as I said I'm using LightGBM with a small set of custom features.\n\nIs Plasticc experience useful here? ;)",
    "452359": "To be honest no! I've mostly been re-implementing Tomas Vantuch's paper.",
    "452391": "this paper: A novel approach of partial discharges detection in a real environment ?\n\nedit: nevermind, I guess it is the paper shared by Tomas in his welcome thread.",
    "452396": "Yes that one, and I'm also using his PhD dissertation in parallel as it contains some complementary details",
    "452419": "The distributions might be similar. However, the signals might be drastically different. If we think about it from an operational point of view, then it makes sense because a grid with an equal amount of failure to working condition is horrible and if the failure signal would be the same, then it could be reduced to similar errors. In such a case it would be rather stupid not find a solution to very similar reasons.",
    "452448": "I bet Mike is using NNs ;)",
    "453447": "Could well be! Also at first @ONODERA's team name was something like \"300 layer benchmark\", he changed since then.",
    "470820": "Thanks for sharing Max!"
  },
  "source": "meta"
}