{
  "id": 262195,
  "title": "6th place solution(monnu part)",
  "url": "/competitions/google-smartphone-decimeter-challenge/discussion/262195",
  "author_name": "",
  "post_date": "2021-08-06T02:00:07.737591900Z",
  "votes": 19,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I would like to thank the host for hosting such a wonderful competition.<br>\nAlso thanks to my three excellent teammates.<br>\nWithout them, I would never have been able to achieve these results.</p>\n<p>Here is a brief description of my part.<br>\nMy teammate <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">shimacos</a>'s part is <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">here</a></p>\n<h3>Overview</h3>\n<ul>\n<li><p>My part is divided into seven sections.</p></li>\n<li><p>Most of the processing is done for each three areas shown in the notebook below.<br>\n<a href=\"https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\" target=\"_blank\">https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label</a></p></li>\n<li><p>A lot of the credit for my part goes to my other teammates.<br>\n(In particular, shimacos's ADR contributed greatly to my model as well.)</p></li>\n</ul>\n<h3>1. Phone Mean</h3>\n<p>This is almost same as  public kernel.<br>\nBy interpolating, I also used the prediction of the phone with a slight shift in timestamp.</p>\n<h3>2. Remove noise while stopping</h3>\n<p>I searched the stopping point and filled in the prediction with the median of near timestamp predictions.<br>\nWhen the distance from the previous position is less than threshold, I process this as stopping.</p>\n<h3>3. Repeat \"phone mean\" and \"Remove noise while stopping\"</h3>\n<p>By repeating \"phone mean\" and \"removing noise while stopped\" several times, the score improved.<br>\nI don't know much about why the repetition was effective, but I expect it was due to the smoothing effect.<br>\nThe number of repetitions was optimized for each area.</p>\n<h3>4. Predict relative positions with LightGBM</h3>\n<p>I predict the relative latDeg and LngDeg to the Ground Truth  in each timestamp.<br>\nFor example, the following features were effective.</p>\n<ul>\n<li>Relative lat,lng and dist of Neighnor Groud Truth(in train, except for that collection)</li>\n<li>other phone relative lat or lng of same timestamp</li>\n<li>lat, lng of previous or next timestamp (etc..)</li>\n</ul>\n<p>After predicting the relative position, I corrected the solution using it.</p>\n<h3>5. Snap to Grid  using cost optimization(downtown only)</h3>\n<p>This was mostly <a href=\"https://www.kaggle.com/go5kuramubon\" target=\"_blank\">pao</a>'s work, with a few changes on my part.<br>\nWe used linearly interpolated ground truth for grid points. <br>\nWe evaluated the cost of each grid point and used the greedy method.<br>\nThere are four points to evaluate the cost.</p>\n<ul>\n<li>we want to snap to nearest ground truth.(simplest )</li>\n<li>The distance to the previous grid point should be as close as possible to the predicted relative distance (using LGBM speed prediction).</li>\n<li>The distance of the previous snap to grid and the distance of the current snap to grid are equal.</li>\n<li>Use the same collection as the grid point used for the previous snap to grid.</li>\n</ul>\n<h3>6. Smoothing of AccumulatedDeltaRange</h3>\n<p>This is my teammate shimacos's findings.This worked also well for my model.<br>\nHere are the details.<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904</a></p>\n<h3>7. Ensembles in different baselines</h3>\n<p>I did 1-6 for several types of baselines.<br>\nBy ensembling them at the end, CV has been improved.<br>\nI used the following three kinds of baselines.</p>\n<ul>\n<li>original baseline + Kalman filter</li>\n<li>WLSADR + position shift(this is also my teammate pao`s work)</li>\n<li>WLSADR + gauss smoothing</li>\n</ul>\n<h3>Result</h3>\n<p>After \"Ensembles in different baselines\", best score of my part is this.</p>\n<ul>\n<li>CV:2.36026, Public:3.745, Private:2.336(Public and Private were checked on latesub)</li>\n</ul>\n<p>By blending it with other team members' models and adding post-processing, we got the final score.</p>",
  "messages": [
    {
      "id": "1453762",
      "postDate": "08/06/2021 02:00:07",
      "content": "<p>First of all, I would like to thank the host for hosting such a wonderful competition.<br>\nAlso thanks to my three excellent teammates.<br>\nWithout them, I would never have been able to achieve these results.</p>\n<p>Here is a brief description of my part.<br>\nMy teammate <a href=\"https://www.kaggle.com/shimacos\" target=\"_blank\">shimacos</a>'s part is <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">here</a></p>\n<h3>Overview</h3>\n<ul>\n<li><p>My part is divided into seven sections.</p></li>\n<li><p>Most of the processing is done for each three areas shown in the notebook below.<br>\n<a href=\"https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\" target=\"_blank\">https://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label</a></p></li>\n<li><p>A lot of the credit for my part goes to my other teammates.<br>\n(In particular, shimacos's ADR contributed greatly to my model as well.)</p></li>\n</ul>\n<h3>1. Phone Mean</h3>\n<p>This is almost same as  public kernel.<br>\nBy interpolating, I also used the prediction of the phone with a slight shift in timestamp.</p>\n<h3>2. Remove noise while stopping</h3>\n<p>I searched the stopping point and filled in the prediction with the median of near timestamp predictions.<br>\nWhen the distance from the previous position is less than threshold, I process this as stopping.</p>\n<h3>3. Repeat \"phone mean\" and \"Remove noise while stopping\"</h3>\n<p>By repeating \"phone mean\" and \"removing noise while stopped\" several times, the score improved.<br>\nI don't know much about why the repetition was effective, but I expect it was due to the smoothing effect.<br>\nThe number of repetitions was optimized for each area.</p>\n<h3>4. Predict relative positions with LightGBM</h3>\n<p>I predict the relative latDeg and LngDeg to the Ground Truth  in each timestamp.<br>\nFor example, the following features were effective.</p>\n<ul>\n<li>Relative lat,lng and dist of Neighnor Groud Truth(in train, except for that collection)</li>\n<li>other phone relative lat or lng of same timestamp</li>\n<li>lat, lng of previous or next timestamp (etc..)</li>\n</ul>\n<p>After predicting the relative position, I corrected the solution using it.</p>\n<h3>5. Snap to Grid  using cost optimization(downtown only)</h3>\n<p>This was mostly <a href=\"https://www.kaggle.com/go5kuramubon\" target=\"_blank\">pao</a>'s work, with a few changes on my part.<br>\nWe used linearly interpolated ground truth for grid points. <br>\nWe evaluated the cost of each grid point and used the greedy method.<br>\nThere are four points to evaluate the cost.</p>\n<ul>\n<li>we want to snap to nearest ground truth.(simplest )</li>\n<li>The distance to the previous grid point should be as close as possible to the predicted relative distance (using LGBM speed prediction).</li>\n<li>The distance of the previous snap to grid and the distance of the current snap to grid are equal.</li>\n<li>Use the same collection as the grid point used for the previous snap to grid.</li>\n</ul>\n<h3>6. Smoothing of AccumulatedDeltaRange</h3>\n<p>This is my teammate shimacos's findings.This worked also well for my model.<br>\nHere are the details.<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904</a></p>\n<h3>7. Ensembles in different baselines</h3>\n<p>I did 1-6 for several types of baselines.<br>\nBy ensembling them at the end, CV has been improved.<br>\nI used the following three kinds of baselines.</p>\n<ul>\n<li>original baseline + Kalman filter</li>\n<li>WLSADR + position shift(this is also my teammate pao`s work)</li>\n<li>WLSADR + gauss smoothing</li>\n</ul>\n<h3>Result</h3>\n<p>After \"Ensembles in different baselines\", best score of my part is this.</p>\n<ul>\n<li>CV:2.36026, Public:3.745, Private:2.336(Public and Private were checked on latesub)</li>\n</ul>\n<p>By blending it with other team members' models and adding post-processing, we got the final score.</p>",
      "rawMarkdown": "First of all, I would like to thank the host for hosting such a wonderful competition.\nAlso thanks to my three excellent teammates.\nWithout them, I would never have been able to achieve these results.\n\nHere is a brief description of my part.\nMy teammate [shimacos](https://www.kaggle.com/shimacos)'s part is [here](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904)\n### Overview\n* My part is divided into seven sections.\n\n* Most of the processing is done for each three areas shown in the notebook below.\nhttps://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\n\n* A lot of the credit for my part goes to my other teammates.\n(In particular, shimacos's ADR contributed greatly to my model as well.)\n\n### 1. Phone Mean\nThis is almost same as  public kernel.\nBy interpolating, I also used the prediction of the phone with a slight shift in timestamp.\n\n### 2. Remove noise while stopping\nI searched the stopping point and filled in the prediction with the median of near timestamp predictions.\nWhen the distance from the previous position is less than threshold, I process this as stopping.\n    \n### 3. Repeat \"phone mean\" and \"Remove noise while stopping\"\nBy repeating \"phone mean\" and \"removing noise while stopped\" several times, the score improved.\nI don't know much about why the repetition was effective, but I expect it was due to the smoothing effect.\nThe number of repetitions was optimized for each area.\n\n### 4. Predict relative positions with LightGBM \nI predict the relative latDeg and LngDeg to the Ground Truth  in each timestamp.\nFor example, the following features were effective.\n* Relative lat,lng and dist of Neighnor Groud Truth(in train, except for that collection)\n* other phone relative lat or lng of same timestamp\n* lat, lng of previous or next timestamp (etc..)\n\nAfter predicting the relative position, I corrected the solution using it.\n\n### 5. Snap to Grid  using cost optimization(downtown only)\nThis was mostly [pao](https://www.kaggle.com/go5kuramubon)'s work, with a few changes on my part.\nWe used linearly interpolated ground truth for grid points. \nWe evaluated the cost of each grid point and used the greedy method.\nThere are four points to evaluate the cost.\n* we want to snap to nearest ground truth.(simplest )\n* The distance to the previous grid point should be as close as possible to the predicted relative distance (using LGBM speed prediction).\n* The distance of the previous snap to grid and the distance of the current snap to grid are equal.\n* Use the same collection as the grid point used for the previous snap to grid.\n\n### 6. Smoothing of AccumulatedDeltaRange\nThis is my teammate shimacos's findings.This worked also well for my model.\nHere are the details.\nhttps://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\n\n### 7. Ensembles in different baselines\nI did 1-6 for several types of baselines.\nBy ensembling them at the end, CV has been improved.\nI used the following three kinds of baselines.\n* original baseline + Kalman filter\n* WLSADR + position shift(this is also my teammate pao`s work)\n* WLSADR + gauss smoothing\n        \n### Result \nAfter \"Ensembles in different baselines\", best score of my part is this.\n* CV:2.36026, Public:3.745, Private:2.336(Public and Private were checked on latesub)\n\nBy blending it with other team members' models and adding post-processing, we got the final score.",
      "votes": null
    },
    {
      "id": "1454771",
      "postDate": "08/06/2021 10:15:04",
      "content": "<p>Congrats Competition Master, and It was fun to work with you.</p>\n<p>Our almost solution is written by <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">Shimacos</a> and monnu, so I will add some other point especially ensemble phase.</p>\n<h2>Large different point remove in ensemble phase</h2>\n<p>In ensemble phase, some point has large different between each model prediction, in that case one of them have possible to missing prediction.<br>\nSo in that pattern, I will remove this point and interpolate it from previous and next point.<br>\nIt makes a little improvement</p>\n<h2>Speed</h2>\n<p>Model has good point and bad point at not only area but also speed.<br>\nOne model has good in high speed, but not good in slow or stop.<br>\nSo I will change weight per speed which is predicted by LightGBM.<br>\nIt makes also a little improvement.</p>",
      "rawMarkdown": "Congrats Competition Master, and It was fun to work with you.\n\nOur almost solution is written by [Shimacos](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904) and monnu, so I will add some other point especially ensemble phase.\n\n## Large different point remove in ensemble phase\n\nIn ensemble phase, some point has large different between each model prediction, in that case one of them have possible to missing prediction.\nSo in that pattern, I will remove this point and interpolate it from previous and next point.\nIt makes a little improvement\n\n## Speed\n\nModel has good point and bad point at not only area but also speed.\nOne model has good in high speed, but not good in slow or stop.\nSo I will change weight per speed which is predicted by LightGBM.\nIt makes also a little improvement.",
      "votes": null
    },
    {
      "id": "1456516",
      "postDate": "08/07/2021 00:34:11",
      "content": "<p>Congratulations on Gold Medal and promotion to Competition Master.<br>\nI have some questions about \"4. Predict relative positions with LightGBM\".I am happy you to answer if you have time.</p>\n<ul>\n<li>Did this process work well for \"highway route\" ? <br>\nWe had a similar process, but it did not work for \"highway route\". (So we had a big shake down.)</li>\n<li>How did you create the CV ?  Is it a division by collectionName?</li>\n</ul>",
      "rawMarkdown": "Congratulations on Gold Medal and promotion to Competition Master.\nI have some questions about \"4. Predict relative positions with LightGBM\".I am happy you to answer if you have time.\n\n- Did this process work well for \"highway route\" ? \n  We had a similar process, but it did not work for \"highway route\". (So we had a big shake down.)\n- How did you create the CV ?  Is it a division by collectionName?",
      "votes": null
    },
    {
      "id": "1461085",
      "postDate": "08/09/2021 07:56:25",
      "content": "<p>Thanks for your comments and questions!<br>\nI will answer in order.</p>\n<p>Q1. <br>\nIn my case, the CV improved on the highway as well, though only slightly.<br>\nThe CV(highway only) before and after this process are shown below.<br>\n2.163→2.082</p>\n<p>Q2.<br>\nYes. I used GroupKFold to build the CV(group = collectionName)</p>\n<p>BTW, You kept the top positions during the competition and also published useful kernels and discussions. <br>\nI think you are definitely one of the great contributors to this competition.</p>",
      "rawMarkdown": "Thanks for your comments and questions!\nI will answer in order.\n\nQ1. \nIn my case, the CV improved on the highway as well, though only slightly.\nThe CV(highway only) before and after this process are shown below.\n2.163→2.082\n\nQ2.\nYes. I used GroupKFold to build the CV(group = collectionName)\n\nBTW, You kept the top positions during the competition and also published useful kernels and discussions. \nI think you are definitely one of the great contributors to this competition.",
      "votes": null
    },
    {
      "id": "1462613",
      "postDate": "08/09/2021 23:55:48",
      "content": "<p>Thank you for answer to questions.<br>\nIt's great to see the effect on the highway.<br>\nCompared to our model, I think your model has better feature set.</p>\n<p>Also,Thanks for your kinds words.<br>\nI was depressed by the shake down, but your words have cheered me up.</p>",
      "rawMarkdown": "Thank you for answer to questions.\nIt's great to see the effect on the highway.\nCompared to our model, I think your model has better feature set.\n\nAlso,Thanks for your kinds words.\nI was depressed by the shake down, but your words have cheered me up.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1454771,
      "author_name": "go5kuramubon",
      "author_url": "",
      "post_date": "08/06/2021 10:15:04",
      "content": "<p>Congrats Competition Master, and It was fun to work with you.</p>\n<p>Our almost solution is written by <a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\" target=\"_blank\">Shimacos</a> and monnu, so I will add some other point especially ensemble phase.</p>\n<h2>Large different point remove in ensemble phase</h2>\n<p>In ensemble phase, some point has large different between each model prediction, in that case one of them have possible to missing prediction.<br>\nSo in that pattern, I will remove this point and interpolate it from previous and next point.<br>\nIt makes a little improvement</p>\n<h2>Speed</h2>\n<p>Model has good point and bad point at not only area but also speed.<br>\nOne model has good in high speed, but not good in slow or stop.<br>\nSo I will change weight per speed which is predicted by LightGBM.<br>\nIt makes also a little improvement.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1456516,
      "author_name": "dehokanta",
      "author_url": "",
      "post_date": "08/07/2021 00:34:11",
      "content": "<p>Congratulations on Gold Medal and promotion to Competition Master.<br>\nI have some questions about \"4. Predict relative positions with LightGBM\".I am happy you to answer if you have time.</p>\n<ul>\n<li>Did this process work well for \"highway route\" ? <br>\nWe had a similar process, but it did not work for \"highway route\". (So we had a big shake down.)</li>\n<li>How did you create the CV ?  Is it a division by collectionName?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1461085,
          "author_name": "fuumin621",
          "author_url": "",
          "post_date": "08/09/2021 07:56:25",
          "content": "<p>Thanks for your comments and questions!<br>\nI will answer in order.</p>\n<p>Q1. <br>\nIn my case, the CV improved on the highway as well, though only slightly.<br>\nThe CV(highway only) before and after this process are shown below.<br>\n2.163→2.082</p>\n<p>Q2.<br>\nYes. I used GroupKFold to build the CV(group = collectionName)</p>\n<p>BTW, You kept the top positions during the competition and also published useful kernels and discussions. <br>\nI think you are definitely one of the great contributors to this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1462613,
          "author_name": "dehokanta",
          "author_url": "",
          "post_date": "08/09/2021 23:55:48",
          "content": "<p>Thank you for answer to questions.<br>\nIt's great to see the effect on the highway.<br>\nCompared to our model, I think your model has better feature set.</p>\n<p>Also,Thanks for your kinds words.<br>\nI was depressed by the shake down, but your words have cheered me up.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1453762": "First of all, I would like to thank the host for hosting such a wonderful competition.\nAlso thanks to my three excellent teammates.\nWithout them, I would never have been able to achieve these results.\n\nHere is a brief description of my part.\nMy teammate [shimacos](https://www.kaggle.com/shimacos)'s part is [here](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904)\n### Overview\n* My part is divided into seven sections.\n\n* Most of the processing is done for each three areas shown in the notebook below.\nhttps://www.kaggle.com/columbia2131/area-knn-prediction-train-hand-label\n\n* A lot of the credit for my part goes to my other teammates.\n(In particular, shimacos's ADR contributed greatly to my model as well.)\n\n### 1. Phone Mean\nThis is almost same as  public kernel.\nBy interpolating, I also used the prediction of the phone with a slight shift in timestamp.\n\n### 2. Remove noise while stopping\nI searched the stopping point and filled in the prediction with the median of near timestamp predictions.\nWhen the distance from the previous position is less than threshold, I process this as stopping.\n    \n### 3. Repeat \"phone mean\" and \"Remove noise while stopping\"\nBy repeating \"phone mean\" and \"removing noise while stopped\" several times, the score improved.\nI don't know much about why the repetition was effective, but I expect it was due to the smoothing effect.\nThe number of repetitions was optimized for each area.\n\n### 4. Predict relative positions with LightGBM \nI predict the relative latDeg and LngDeg to the Ground Truth  in each timestamp.\nFor example, the following features were effective.\n* Relative lat,lng and dist of Neighnor Groud Truth(in train, except for that collection)\n* other phone relative lat or lng of same timestamp\n* lat, lng of previous or next timestamp (etc..)\n\nAfter predicting the relative position, I corrected the solution using it.\n\n### 5. Snap to Grid  using cost optimization(downtown only)\nThis was mostly [pao](https://www.kaggle.com/go5kuramubon)'s work, with a few changes on my part.\nWe used linearly interpolated ground truth for grid points. \nWe evaluated the cost of each grid point and used the greedy method.\nThere are four points to evaluate the cost.\n* we want to snap to nearest ground truth.(simplest )\n* The distance to the previous grid point should be as close as possible to the predicted relative distance (using LGBM speed prediction).\n* The distance of the previous snap to grid and the distance of the current snap to grid are equal.\n* Use the same collection as the grid point used for the previous snap to grid.\n\n### 6. Smoothing of AccumulatedDeltaRange\nThis is my teammate shimacos's findings.This worked also well for my model.\nHere are the details.\nhttps://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904\n\n### 7. Ensembles in different baselines\nI did 1-6 for several types of baselines.\nBy ensembling them at the end, CV has been improved.\nI used the following three kinds of baselines.\n* original baseline + Kalman filter\n* WLSADR + position shift(this is also my teammate pao`s work)\n* WLSADR + gauss smoothing\n        \n### Result \nAfter \"Ensembles in different baselines\", best score of my part is this.\n* CV:2.36026, Public:3.745, Private:2.336(Public and Private were checked on latesub)\n\nBy blending it with other team members' models and adding post-processing, we got the final score.",
    "1454771": "Congrats Competition Master, and It was fun to work with you.\n\nOur almost solution is written by [Shimacos](https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/261904) and monnu, so I will add some other point especially ensemble phase.\n\n## Large different point remove in ensemble phase\n\nIn ensemble phase, some point has large different between each model prediction, in that case one of them have possible to missing prediction.\nSo in that pattern, I will remove this point and interpolate it from previous and next point.\nIt makes a little improvement\n\n## Speed\n\nModel has good point and bad point at not only area but also speed.\nOne model has good in high speed, but not good in slow or stop.\nSo I will change weight per speed which is predicted by LightGBM.\nIt makes also a little improvement.",
    "1456516": "Congratulations on Gold Medal and promotion to Competition Master.\nI have some questions about \"4. Predict relative positions with LightGBM\".I am happy you to answer if you have time.\n\n- Did this process work well for \"highway route\" ? \n  We had a similar process, but it did not work for \"highway route\". (So we had a big shake down.)\n- How did you create the CV ?  Is it a division by collectionName?",
    "1461085": "Thanks for your comments and questions!\nI will answer in order.\n\nQ1. \nIn my case, the CV improved on the highway as well, though only slightly.\nThe CV(highway only) before and after this process are shown below.\n2.163→2.082\n\nQ2.\nYes. I used GroupKFold to build the CV(group = collectionName)\n\nBTW, You kept the top positions during the competition and also published useful kernels and discussions. \nI think you are definitely one of the great contributors to this competition.",
    "1462613": "Thank you for answer to questions.\nIt's great to see the effect on the highway.\nCompared to our model, I think your model has better feature set.\n\nAlso,Thanks for your kinds words.\nI was depressed by the shake down, but your words have cheered me up."
  },
  "source": "meta"
}