{
  "id": 85279,
  "title": "Extremely simple four-dim feature, works very well for local CV but...",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85279",
  "author_name": "yu4u",
  "post_date": "2019-03-22T16:47:06.028000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats to everyone who received any medal!\nI enjoyed trying many approaches but I could not find out fundamental differences between train and test signals.</p>\n\n<p>Here let me introduce my simple feature for LSTM that uses almost raw signals, and is a sequence of only four-dimensional features.</p>\n\n<ol>\n<li>Extract high frequency signal F by subtracting median filtered signal from original signal</li>\n<li>Select 1000 points out of 800,000 measurement from F so that selected points have top 1000 largest absolute values in any phase</li>\n<li>Use 1000 sequence of (interval, sig) as an input to LSTM model. Interval is an interval between two successive selected points and sig is three-dimensional feature from 3-phase signal F. Thus, resulting feature is a 1000 sequence of four-dimensional features.</li>\n</ol>\n\n<p>This very simple feature + 0.694 kernel LSTM stably achieves ~0.79 local CV (but not so good for public LB; relatively works for private LB).\nI really love this approach due to its simplicity but I could not fill the gap between train and test accuracy (I could not find significant differences between train and test features).</p>\n\n<p>Did anyone try similar approach or have any ideas to make this model work better for test dataset...?</p>\n\n<p>```\ndef extract_feature(x):\n    # x.shape = (3, 800000)\n    med_x = scipy.signal.medfilt(x, kernel_size=(1, 31))\n    diff = x - med_x\n    y = np.argsort(-(abs(diff).max(axis=0)))[:1000]\n    z = sorted(y)\n    prev_idx = -1\n    results = []</p>\n\n<pre><code>for i in z:\n    results.append([i - prev_idx, *x[:, i]])\n    prev_idx = i\n\nreturn np.array(results)\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": 496846,
      "postDate": "2019-03-22T16:47:06.030Z",
      "content": "<p>Congrats to everyone who received any medal!\nI enjoyed trying many approaches but I could not find out fundamental differences between train and test signals.</p>\n\n<p>Here let me introduce my simple feature for LSTM that uses almost raw signals, and is a sequence of only four-dimensional features.</p>\n\n<ol>\n<li>Extract high frequency signal F by subtracting median filtered signal from original signal</li>\n<li>Select 1000 points out of 800,000 measurement from F so that selected points have top 1000 largest absolute values in any phase</li>\n<li>Use 1000 sequence of (interval, sig) as an input to LSTM model. Interval is an interval between two successive selected points and sig is three-dimensional feature from 3-phase signal F. Thus, resulting feature is a 1000 sequence of four-dimensional features.</li>\n</ol>\n\n<p>This very simple feature + 0.694 kernel LSTM stably achieves ~0.79 local CV (but not so good for public LB; relatively works for private LB).\nI really love this approach due to its simplicity but I could not fill the gap between train and test accuracy (I could not find significant differences between train and test features).</p>\n\n<p>Did anyone try similar approach or have any ideas to make this model work better for test dataset...?</p>\n\n<p>```\ndef extract_feature(x):\n    # x.shape = (3, 800000)\n    med_x = scipy.signal.medfilt(x, kernel_size=(1, 31))\n    diff = x - med_x\n    y = np.argsort(-(abs(diff).max(axis=0)))[:1000]\n    z = sorted(y)\n    prev_idx = -1\n    results = []</p>\n\n<pre><code>for i in z:\n    results.append([i - prev_idx, *x[:, i]])\n    prev_idx = i\n\nreturn np.array(results)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Congrats to everyone who received any medal!\nI enjoyed trying many approaches but I could not find out fundamental differences between train and test signals.\n\nHere let me introduce my simple feature for LSTM that uses almost raw signals, and is a sequence of only four-dimensional features.\n\n1. Extract high frequency signal F by subtracting median filtered signal from original signal\n2. Select 1000 points out of 800,000 measurement from F so that selected points have top 1000 largest absolute values in any phase\n3. Use 1000 sequence of (interval, sig) as an input to LSTM model. Interval is an interval between two successive selected points and sig is three-dimensional feature from 3-phase signal F. Thus, resulting feature is a 1000 sequence of four-dimensional features.\n\nThis very simple feature + 0.694 kernel LSTM stably achieves ~0.79 local CV (but not so good for public LB; relatively works for private LB).\nI really love this approach due to its simplicity but I could not fill the gap between train and test accuracy (I could not find significant differences between train and test features).\n\nDid anyone try similar approach or have any ideas to make this model work better for test dataset...?\n\n```\ndef extract_feature(x):\n    # x.shape = (3, 800000)\n    med_x = scipy.signal.medfilt(x, kernel_size=(1, 31))\n    diff = x - med_x\n    y = np.argsort(-(abs(diff).max(axis=0)))[:1000]\n    z = sorted(y)\n    prev_idx = -1\n    results = []\n\n    for i in z:\n        results.append([i - prev_idx, *x[:, i]])\n        prev_idx = i\n\n    return np.array(results)\n```",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "496846": "Congrats to everyone who received any medal!\nI enjoyed trying many approaches but I could not find out fundamental differences between train and test signals.\n\nHere let me introduce my simple feature for LSTM that uses almost raw signals, and is a sequence of only four-dimensional features.\n\n1. Extract high frequency signal F by subtracting median filtered signal from original signal\n2. Select 1000 points out of 800,000 measurement from F so that selected points have top 1000 largest absolute values in any phase\n3. Use 1000 sequence of (interval, sig) as an input to LSTM model. Interval is an interval between two successive selected points and sig is three-dimensional feature from 3-phase signal F. Thus, resulting feature is a 1000 sequence of four-dimensional features.\n\nThis very simple feature + 0.694 kernel LSTM stably achieves ~0.79 local CV (but not so good for public LB; relatively works for private LB).\nI really love this approach due to its simplicity but I could not fill the gap between train and test accuracy (I could not find significant differences between train and test features).\n\nDid anyone try similar approach or have any ideas to make this model work better for test dataset...?\n\n```\ndef extract_feature(x):\n    # x.shape = (3, 800000)\n    med_x = scipy.signal.medfilt(x, kernel_size=(1, 31))\n    diff = x - med_x\n    y = np.argsort(-(abs(diff).max(axis=0)))[:1000]\n    z = sorted(y)\n    prev_idx = -1\n    results = []\n\n    for i in z:\n        results.append([i - prev_idx, *x[:, i]])\n        prev_idx = i\n\n    return np.array(results)\n```"
  }
}