{
  "id": 94459,
  "title": "3rd place memo",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/94459",
  "author_name": "hmdhmd",
  "post_date": "2019-06-04T16:17:23.172000",
  "votes": 33,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to organizers for hosting such an interesting competition, and congratulations to all the top teams.</p>\n\n<p>Already mentioned in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction\">discussion</a> , we expected test from p4677.\nSo we can roughly estimate the test ttf.\nBased on this information, We adjusted the ttf of train.\nGreen line(I will call it best y) is 12.625.</p>\n\n<p>By training based on the above adjusted ttf, you can get a good private score.\nFor example, RNN with only 6 features and simple GRU, you can get 2.33 on privateLB.\nActually, lightgbm with many features is slightly better, so we mainly used it.</p>\n\n<p>How to find best y?\nFor example, consider the situation:\n・train data : 1~11 earthquakes of train\n・test data : 12,13,14,15 earthquakes of train\nThe best value can be calculated as follows.</p>\n\n<pre><code>max_ttf_list = [8.828100, 8.566000, 14.751800, 9.459500]\nchunk_length_list = [33988602, 32976890, 56791029, 36417529]\nbest_score = None\nfor y in np.linspace(5,15,1000): # search range\n    error_list = []\n    segment_list = []\n    for max_ttf, chunk_length in zip(max_ttf_list, chunk_length_list):\n        slope = max_ttf/chunk_length*150000\n        segument_num = chunk_length//150000\n        ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n        error = abs(np.array(ttf_true)-np.array(ttf_adjusted)).sum()\n        error_list.append(error)\n        segment_list.append(segument_num)\n    score = sum(error_list)/sum(segment_list)\n    if best_score is None:\n        best_score = score\n    elif best_score &gt; score:\n        best_score = score\n        best_y = y\nprint(best_y)\n</code></pre>\n\n<p>If training by default ttf, MAE of test data is 2.251.\nIf training by best y and modified ttf, MAE of test data is 1.948.</p>\n\n<p>It may be said that this competition is a leak, but it was interesting to think about creating a better model on that premise.</p>\n\n<p>We would like to say thanks to everyone.\nThat's all. Thank you for your reading.</p>",
  "messages": [
    {
      "id": 543607,
      "postDate": "2019-06-04T16:17:23.173Z",
      "content": "<p>Thanks to organizers for hosting such an interesting competition, and congratulations to all the top teams.</p>\n\n<p>Already mentioned in the <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction\">discussion</a> , we expected test from p4677.\nSo we can roughly estimate the test ttf.\nBased on this information, We adjusted the ttf of train.\nGreen line(I will call it best y) is 12.625.</p>\n\n<p>By training based on the above adjusted ttf, you can get a good private score.\nFor example, RNN with only 6 features and simple GRU, you can get 2.33 on privateLB.\nActually, lightgbm with many features is slightly better, so we mainly used it.</p>\n\n<p>How to find best y?\nFor example, consider the situation:\n・train data : 1~11 earthquakes of train\n・test data : 12,13,14,15 earthquakes of train\nThe best value can be calculated as follows.</p>\n\n<pre><code>max_ttf_list = [8.828100, 8.566000, 14.751800, 9.459500]\nchunk_length_list = [33988602, 32976890, 56791029, 36417529]\nbest_score = None\nfor y in np.linspace(5,15,1000): # search range\n    error_list = []\n    segment_list = []\n    for max_ttf, chunk_length in zip(max_ttf_list, chunk_length_list):\n        slope = max_ttf/chunk_length*150000\n        segument_num = chunk_length//150000\n        ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n        error = abs(np.array(ttf_true)-np.array(ttf_adjusted)).sum()\n        error_list.append(error)\n        segment_list.append(segument_num)\n    score = sum(error_list)/sum(segment_list)\n    if best_score is None:\n        best_score = score\n    elif best_score &gt; score:\n        best_score = score\n        best_y = y\nprint(best_y)\n</code></pre>\n\n<p>If training by default ttf, MAE of test data is 2.251.\nIf training by best y and modified ttf, MAE of test data is 1.948.</p>\n\n<p>It may be said that this competition is a leak, but it was interesting to think about creating a better model on that premise.</p>\n\n<p>We would like to say thanks to everyone.\nThat's all. Thank you for your reading.</p>",
      "rawMarkdown": "Thanks to organizers for hosting such an interesting competition, and congratulations to all the top teams.\n\nAlready mentioned in the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction) , we expected test from p4677.\nSo we can roughly estimate the test ttf.\nBased on this information, We adjusted the ttf of train.\nGreen line(I will call it best y) is 12.625.\n\nBy training based on the above adjusted ttf, you can get a good private score.\nFor example, RNN with only 6 features and simple GRU, you can get 2.33 on privateLB.\nActually, lightgbm with many features is slightly better, so we mainly used it.\n\nHow to find best y?\nFor example, consider the situation:\n・train data : 1~11 earthquakes of train\n・test data : 12,13,14,15 earthquakes of train\nThe best value can be calculated as follows.\n\n\n    max_ttf_list = [8.828100, 8.566000, 14.751800, 9.459500]\n    chunk_length_list = [33988602, 32976890, 56791029, 36417529]\n    best_score = None\n    for y in np.linspace(5,15,1000): # search range\n        error_list = []\n        segment_list = []\n        for max_ttf, chunk_length in zip(max_ttf_list, chunk_length_list):\n            slope = max_ttf/chunk_length*150000\n            segument_num = chunk_length//150000\n            ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n            ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n            error = abs(np.array(ttf_true)-np.array(ttf_adjusted)).sum()\n            error_list.append(error)\n            segment_list.append(segument_num)\n        score = sum(error_list)/sum(segment_list)\n        if best_score is None:\n            best_score = score\n        elif best_score &gt; score:\n            best_score = score\n            best_y = y\n    print(best_y)\n\n\nIf training by default ttf, MAE of test data is 2.251.\nIf training by best y and modified ttf, MAE of test data is 1.948.\n\nIt may be said that this competition is a leak, but it was interesting to think about creating a better model on that premise.\n\nWe would like to say thanks to everyone.\nThat's all. Thank you for your reading.",
      "votes": 33
    },
    {
      "id": 548985,
      "postDate": "2019-06-10T07:01:25.780Z",
      "content": "<p>Congrats! Can you show the hyperparams of lightgbm? I suppose we tend to cause overfitting with many features. If you have any tips, please let me know.</p>",
      "rawMarkdown": "Congrats! Can you show the hyperparams of lightgbm? I suppose we tend to cause overfitting with many features. If you have any tips, please let me know.",
      "votes": 1,
      "replies": [
        {
          "id": 549323,
          "postDate": "2019-06-10T14:46:55.937Z",
          "content": "<p>Thank you for your reply.\nTo be honestly, we did not spend much time on tuning parameter.\nMy teammate currypurin used <a href=\"https://github.com/pfnet/optuna\">optuna</a> and shared with me this param.</p>\n\n<pre><code>param = {\n    'num_leaves': 37,\n    'objective': 'regression',\n    'max_depth':9,\n    'learning_rate': 0.01,\n    'boosting': 'gbdt',\n    'feature_fraction': 0.311473528266664,\n    'bagging_freq': 14,\n    'bagging_fraction': 0.4190138426177889,\n    'bagging_seed': 42,\n    'metric': 'mae',\n    'lambda_l1': 0.0003661365327691201,\n    'lambda_l2': 0.2855479018074398,\n    'verbosity': -1,\n    'nthread': -1,\n    'random_state': 0,\n    'min_data_in_leaf': 40\n}\n</code></pre>\n\n<p>Usually ,  when data size is small and have many features,  I start with conservative parameters.\n(And this is fast, you can spend much time with feature engineering and so on)</p>\n\n<pre><code>param = {\n    \"num_leaves\": 20,\n    \"max_depth\": 4,\n    \"bagging_freq\": 1,\n    \"bagging_fraction\": 0.7,\n    \"feature_fraction\": 0.7,\n    ...\n}\n</code></pre>",
          "rawMarkdown": "Thank you for your reply.\nTo be honestly, we did not spend much time on tuning parameter.\nMy teammate currypurin used [optuna](https://github.com/pfnet/optuna) and shared with me this param.\n\n    param = {\n        'num_leaves': 37,\n        'objective': 'regression',\n        'max_depth':9,\n        'learning_rate': 0.01,\n        'boosting': 'gbdt',\n        'feature_fraction': 0.311473528266664,\n        'bagging_freq': 14,\n        'bagging_fraction': 0.4190138426177889,\n        'bagging_seed': 42,\n        'metric': 'mae',\n        'lambda_l1': 0.0003661365327691201,\n        'lambda_l2': 0.2855479018074398,\n        'verbosity': -1,\n        'nthread': -1,\n        'random_state': 0,\n        'min_data_in_leaf': 40\n    }\n\nUsually ,  when data size is small and have many features,  I start with conservative parameters.\n(And this is fast, you can spend much time with feature engineering and so on)\n\n    param = {\n        \"num_leaves\": 20,\n        \"max_depth\": 4,\n        \"bagging_freq\": 1,\n        \"bagging_fraction\": 0.7,\n        \"feature_fraction\": 0.7,\n        ...\n    }\n",
          "votes": 1
        },
        {
          "id": 549356,
          "postDate": "2019-06-10T15:33:28.570Z",
          "content": "<p>Thank you for showing params and tips! I got a little surprised to see you use 'regression' not 'gamma' as a 'objective'.</p>",
          "rawMarkdown": "Thank you for showing params and tips! I got a little surprised to see you use 'regression' not 'gamma' as a 'objective'.",
          "votes": 1
        }
      ]
    },
    {
      "id": 547318,
      "postDate": "2019-06-07T15:09:48.253Z",
      "content": "<p>This seems like a numerical approach to approximate:\n<code>best_y = np.median(np.hstack([np.repeat(cycle_len, int(100*cycle_len)) for cycle_len in chunk_length_list]))</code></p>",
      "rawMarkdown": "This seems like a numerical approach to approximate:\n`best_y = np.median(np.hstack([np.repeat(cycle_len, int(100*cycle_len)) for cycle_len in chunk_length_list]))`",
      "votes": 1,
      "replies": [
        {
          "id": 548037,
          "postDate": "2019-06-08T16:56:16.963Z",
          "content": "<p>Thank you for your reply.\nIt's cool.\nMy code is lengthy...</p>",
          "rawMarkdown": "Thank you for your reply.\nIt's cool.\nMy code is lengthy...",
          "votes": 1
        },
        {
          "id": 548392,
          "postDate": "2019-06-09T10:10:23.443Z",
          "content": "<p>Your approach is certainly cool too! It's great to be able to approach this from multiple angles :D</p>",
          "rawMarkdown": "Your approach is certainly cool too! It's great to be able to approach this from multiple angles :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 543747,
      "postDate": "2019-06-04T19:01:14.820Z",
      "content": "<p>oh nice!!</p>",
      "rawMarkdown": "oh nice!!",
      "votes": 1
    },
    {
      "id": 543685,
      "postDate": "2019-06-04T17:41:19.847Z",
      "content": "<p>Very interesting approach. Tried a bit of target modifications briefly but never occurred to me set EQs to the same TTF or do it the way you did. Seems brilliant! Congrats! </p>",
      "rawMarkdown": "Very interesting approach. Tried a bit of target modifications briefly but never occurred to me set EQs to the same TTF or do it the way you did. Seems brilliant! Congrats! ",
      "votes": 1
    },
    {
      "id": 543619,
      "postDate": "2019-06-04T16:28:40.903Z",
      "content": "<p>Thanks for sharing, and congrats on the result!</p>",
      "rawMarkdown": "Thanks for sharing, and congrats on the result!",
      "votes": 1
    },
    {
      "id": 565843,
      "postDate": "2019-07-01T12:54:16.553Z",
      "content": "<p>Congrats and thanks for share! But I dont really understand you solution and have some questions, what is Green line or best y?mean ttf of validate set? why<code>\n ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n</code>\nhow to use best y?</p>",
      "rawMarkdown": "Congrats and thanks for share! But I dont really understand you solution and have some questions, what is Green line or best y?mean ttf of validate set? why```\n ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n```\nhow to use best y?"
    },
    {
      "id": 547313,
      "postDate": "2019-06-07T14:50:05.860Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 544280,
      "postDate": "2019-06-05T10:55:25.940Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 543669,
      "postDate": "2019-06-04T17:21:53.663Z",
      "content": "<p>Great idea ! thanks !</p>",
      "rawMarkdown": "Great idea ! thanks !",
      "votes": 1
    },
    {
      "id": 543629,
      "postDate": "2019-06-04T16:38:39.327Z",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 548985,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2019-06-10T07:01:25.780000",
      "content": "<p>Congrats! Can you show the hyperparams of lightgbm? I suppose we tend to cause overfitting with many features. If you have any tips, please let me know.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 549323,
          "author_name": "hmdhmd",
          "author_url": "",
          "post_date": "2019-06-10T14:46:55.937000",
          "content": "<p>Thank you for your reply.\nTo be honestly, we did not spend much time on tuning parameter.\nMy teammate currypurin used <a href=\"https://github.com/pfnet/optuna\">optuna</a> and shared with me this param.</p>\n\n<pre><code>param = {\n    'num_leaves': 37,\n    'objective': 'regression',\n    'max_depth':9,\n    'learning_rate': 0.01,\n    'boosting': 'gbdt',\n    'feature_fraction': 0.311473528266664,\n    'bagging_freq': 14,\n    'bagging_fraction': 0.4190138426177889,\n    'bagging_seed': 42,\n    'metric': 'mae',\n    'lambda_l1': 0.0003661365327691201,\n    'lambda_l2': 0.2855479018074398,\n    'verbosity': -1,\n    'nthread': -1,\n    'random_state': 0,\n    'min_data_in_leaf': 40\n}\n</code></pre>\n\n<p>Usually ,  when data size is small and have many features,  I start with conservative parameters.\n(And this is fast, you can spend much time with feature engineering and so on)</p>\n\n<pre><code>param = {\n    \"num_leaves\": 20,\n    \"max_depth\": 4,\n    \"bagging_freq\": 1,\n    \"bagging_fraction\": 0.7,\n    \"feature_fraction\": 0.7,\n    ...\n}\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 549356,
          "author_name": "u++",
          "author_url": "",
          "post_date": "2019-06-10T15:33:28.570000",
          "content": "<p>Thank you for showing params and tips! I got a little surprised to see you use 'regression' not 'gamma' as a 'objective'.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 547318,
      "author_name": "Tom Van de Wiele",
      "author_url": "",
      "post_date": "2019-06-07T15:09:48.253000",
      "content": "<p>This seems like a numerical approach to approximate:\n<code>best_y = np.median(np.hstack([np.repeat(cycle_len, int(100*cycle_len)) for cycle_len in chunk_length_list]))</code></p>",
      "votes": 1,
      "replies": [
        {
          "id": 548037,
          "author_name": "hmdhmd",
          "author_url": "",
          "post_date": "2019-06-08T16:56:16.963000",
          "content": "<p>Thank you for your reply.\nIt's cool.\nMy code is lengthy...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 548392,
          "author_name": "Tom Van de Wiele",
          "author_url": "",
          "post_date": "2019-06-09T10:10:23.443000",
          "content": "<p>Your approach is certainly cool too! It's great to be able to approach this from multiple angles :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 543747,
      "author_name": "Tasfik Rahman",
      "author_url": "",
      "post_date": "2019-06-04T19:01:14.820000",
      "content": "<p>oh nice!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543685,
      "author_name": "Tim Yee",
      "author_url": "",
      "post_date": "2019-06-04T17:41:19.847000",
      "content": "<p>Very interesting approach. Tried a bit of target modifications briefly but never occurred to me set EQs to the same TTF or do it the way you did. Seems brilliant! Congrats! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543619,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-06-04T16:28:40.903000",
      "content": "<p>Thanks for sharing, and congrats on the result!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 565843,
      "author_name": "Kyle",
      "author_url": "",
      "post_date": "2019-07-01T12:54:16.553000",
      "content": "<p>Congrats and thanks for share! But I dont really understand you solution and have some questions, what is Green line or best y?mean ttf of validate set? why<code>\n ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n</code>\nhow to use best y?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 547313,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-07T14:50:05.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 544280,
      "author_name": "Timmmmmms",
      "author_url": "",
      "post_date": "2019-06-05T10:55:25.940000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543669,
      "author_name": "Antoine",
      "author_url": "",
      "post_date": "2019-06-04T17:21:53.663000",
      "content": "<p>Great idea ! thanks !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 543629,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-06-04T16:38:39.327000",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "543607": "Thanks to organizers for hosting such an interesting competition, and congratulations to all the top teams.\n\nAlready mentioned in the [discussion](https://www.kaggle.com/c/LANL-Earthquake-Prediction) , we expected test from p4677.\nSo we can roughly estimate the test ttf.\nBased on this information, We adjusted the ttf of train.\nGreen line(I will call it best y) is 12.625.\n\nBy training based on the above adjusted ttf, you can get a good private score.\nFor example, RNN with only 6 features and simple GRU, you can get 2.33 on privateLB.\nActually, lightgbm with many features is slightly better, so we mainly used it.\n\nHow to find best y?\nFor example, consider the situation:\n・train data : 1~11 earthquakes of train\n・test data : 12,13,14,15 earthquakes of train\nThe best value can be calculated as follows.\n\n\n    max_ttf_list = [8.828100, 8.566000, 14.751800, 9.459500]\n    chunk_length_list = [33988602, 32976890, 56791029, 36417529]\n    best_score = None\n    for y in np.linspace(5,15,1000): # search range\n        error_list = []\n        segment_list = []\n        for max_ttf, chunk_length in zip(max_ttf_list, chunk_length_list):\n            slope = max_ttf/chunk_length*150000\n            segument_num = chunk_length//150000\n            ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n            ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n            error = abs(np.array(ttf_true)-np.array(ttf_adjusted)).sum()\n            error_list.append(error)\n            segment_list.append(segument_num)\n        score = sum(error_list)/sum(segment_list)\n        if best_score is None:\n            best_score = score\n        elif best_score &gt; score:\n            best_score = score\n            best_y = y\n    print(best_y)\n\n\nIf training by default ttf, MAE of test data is 2.251.\nIf training by best y and modified ttf, MAE of test data is 1.948.\n\nIt may be said that this competition is a leak, but it was interesting to think about creating a better model on that premise.\n\nWe would like to say thanks to everyone.\nThat's all. Thank you for your reading.",
    "548985": "Congrats! Can you show the hyperparams of lightgbm? I suppose we tend to cause overfitting with many features. If you have any tips, please let me know.",
    "547318": "This seems like a numerical approach to approximate:\n`best_y = np.median(np.hstack([np.repeat(cycle_len, int(100*cycle_len)) for cycle_len in chunk_length_list]))`",
    "543747": "oh nice!!",
    "543685": "Very interesting approach. Tried a bit of target modifications briefly but never occurred to me set EQs to the same TTF or do it the way you did. Seems brilliant! Congrats! ",
    "543619": "Thanks for sharing, and congrats on the result!",
    "565843": "Congrats and thanks for share! But I dont really understand you solution and have some questions, what is Green line or best y?mean ttf of validate set? why```\n ttf_true = [max_ttf-slope*i for i in range(segument_num)]\n        ttf_adjusted = [y-(y/segument_num)*i for i in range(segument_num)]\n```\nhow to use best y?",
    "547313": "",
    "544280": "Thanks for sharing",
    "543669": "Great idea ! thanks !",
    "543629": "Congrats and thanks for sharing."
  }
}