{
  "id": 81001,
  "title": "local cv 0.9+ LB 0.2（non DL)... why are there so big difference?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/81001",
  "author_name": "little_snail",
  "post_date": "2019-02-18T14:50:42.763000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>When I merge 3-phase data into one line, local cv could reach 0.9+. However, LB is still bad. \nI have checked my code, and I think there is no leakage. Have you met similar problem?</p>\n\n<pre><code>def threePhasesConcat(features, meta):    \ntemp = pd.merge(left=features, right=meta[['signal_id', 'id_measurement', 'phase']], on='signal_id', how='left')\ntemp =  temp.drop(['signal_id'], axis=1)\n\ntemp = temp.set_index(['id_measurement', 'phase']).unstack('phase')\ntempNp = temp.values\n\nmeasure_features = pd.DataFrame(tempNp)\nmeasure_features['id_measurement'] = temp.index\n\nreturn pd.merge(left=meta, right=measure_features, how='left', on='id_measurement')\n</code></pre>",
  "messages": [
    {
      "id": 473808,
      "postDate": "2019-02-18T14:50:42.763Z",
      "content": "<p>When I merge 3-phase data into one line, local cv could reach 0.9+. However, LB is still bad. \nI have checked my code, and I think there is no leakage. Have you met similar problem?</p>\n\n<pre><code>def threePhasesConcat(features, meta):    \ntemp = pd.merge(left=features, right=meta[['signal_id', 'id_measurement', 'phase']], on='signal_id', how='left')\ntemp =  temp.drop(['signal_id'], axis=1)\n\ntemp = temp.set_index(['id_measurement', 'phase']).unstack('phase')\ntempNp = temp.values\n\nmeasure_features = pd.DataFrame(tempNp)\nmeasure_features['id_measurement'] = temp.index\n\nreturn pd.merge(left=meta, right=measure_features, how='left', on='id_measurement')\n</code></pre>",
      "rawMarkdown": "When I merge 3-phase data into one line, local cv could reach 0.9+. However, LB is still bad. \nI have checked my code, and I think there is no leakage. Have you met similar problem?\n\n    def threePhasesConcat(features, meta):    \n    temp = pd.merge(left=features, right=meta[['signal_id', 'id_measurement', 'phase']], on='signal_id', how='left')\n    temp =  temp.drop(['signal_id'], axis=1)\n\n    temp = temp.set_index(['id_measurement', 'phase']).unstack('phase')\n    tempNp = temp.values\n\n    measure_features = pd.DataFrame(tempNp)\n    measure_features['id_measurement'] = temp.index\n    \n    return pd.merge(left=meta, right=measure_features, how='left', on='id_measurement')\n",
      "votes": 1
    },
    {
      "id": 474209,
      "postDate": "2019-02-19T04:33:38.440Z",
      "content": "<p>Have you tried adversarial validation to compare distribution of train and test set ? Also, hoping that you are not using \"id_measurement\" as a feature. If you are using some decision tree, check <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166</a>.</p>",
      "rawMarkdown": "Have you tried adversarial validation to compare distribution of train and test set ? Also, hoping that you are not using \"id_measurement\" as a feature. If you are using some decision tree, check https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166.",
      "votes": 2,
      "replies": [
        {
          "id": 474214,
          "postDate": "2019-02-19T04:48:55.053Z",
          "content": "<p>Thank you very much. I will try \"adversarial validation\"</p>",
          "rawMarkdown": "Thank you very much. I will try \"adversarial validation\""
        },
        {
          "id": 476113,
          "postDate": "2019-02-21T16:17:02.980Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 476112,
      "postDate": "2019-02-21T16:13:48.737Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 474209,
      "author_name": "HarshitMehta",
      "author_url": "",
      "post_date": "2019-02-19T04:33:38.440000",
      "content": "<p>Have you tried adversarial validation to compare distribution of train and test set ? Also, hoping that you are not using \"id_measurement\" as a feature. If you are using some decision tree, check <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166</a>.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 474214,
          "author_name": "little_snail",
          "author_url": "",
          "post_date": "2019-02-19T04:48:55.053000",
          "content": "<p>Thank you very much. I will try \"adversarial validation\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 476113,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-21T16:17:02.980000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 476112,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T16:13:48.737000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "473808": "When I merge 3-phase data into one line, local cv could reach 0.9+. However, LB is still bad. \nI have checked my code, and I think there is no leakage. Have you met similar problem?\n\n    def threePhasesConcat(features, meta):    \n    temp = pd.merge(left=features, right=meta[['signal_id', 'id_measurement', 'phase']], on='signal_id', how='left')\n    temp =  temp.drop(['signal_id'], axis=1)\n\n    temp = temp.set_index(['id_measurement', 'phase']).unstack('phase')\n    tempNp = temp.values\n\n    measure_features = pd.DataFrame(tempNp)\n    measure_features['id_measurement'] = temp.index\n    \n    return pd.merge(left=meta, right=measure_features, how='left', on='id_measurement')\n",
    "474209": "Have you tried adversarial validation to compare distribution of train and test set ? Also, hoping that you are not using \"id_measurement\" as a feature. If you are using some decision tree, check https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166.",
    "476112": ""
  }
}