{
  "id": 263663,
  "title": "10 days bronze (63 place solution)",
  "url": "/competitions/google-smartphone-decimeter-challenge/writeups/victorasso-10-days-bronze-63-place-solution",
  "author_name": "",
  "post_date": "2021-08-10T01:07:27.303253800Z",
  "votes": 9,
  "comment_count": 1,
  "views": 0,
  "content": "<p><strong>tl;dr: New baseline from raw data &gt; Recalculate all pipelines on the best ensemble &gt; Snap SJC to ground true</strong></p>\n<p>Hello everyone, i wanted to share the steps i took to get the bronze medal, i only started working on this competition 10 days before the deadline and mostly used work from the community to get my final score, so i'll also leave my special thanks to the ones who shared the knowledge i used.</p>\n<p><a href=\"https://www.kaggle.com/jeongyoonlee\" target=\"_blank\">@jeongyoonlee</a> for this great baseline<br>\n<a href=\"https://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu\" target=\"_blank\">https://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu</a></p>\n<p><a href=\"https://www.kaggle.com/bpetrb\" target=\"_blank\">@bpetrb</a> for this amazing notebook<br>\n<a href=\"https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output\" target=\"_blank\">https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output</a></p>\n<p><a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a> for his clustering technique<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160</a></p>\n<p><a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a> for this great notebook<br>\n<a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output\" target=\"_blank\">https://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output</a></p>\n<p><a href=\"https://www.kaggle.com/tensorchoko\" target=\"_blank\">@tensorchoko</a> for her multioutput regressor, i believe she removed the notebook since i can't find it anymore</p>\n<p><a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a> for this amazing ensemble notebook<br>\n<a href=\"https://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling\" target=\"_blank\">https://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling</a></p>\n<p>The combined output from those notebooks plus a little original code granted me the medal, this is how it was done:</p>\n<h1>1 - Get a better baseline</h1>\n<p>On the first day of my research when i was trying to understand the data, i could see people talking about the <strong>base_locations_test.csv</strong> file was a good starting point and doing many post-processing techniques on it, so if i could find a better baseline, i could just apply the same public techniques people were doing and i would probably end up with a better result, so i spend some time trying to find someone who had calculated the baseline again, i end up finding a fork of <a href=\"https://www.kaggle.com/jeongyoonlee\" target=\"_blank\">@jeongyoonlee</a> code and after a few tweaks, i ended up with a <strong>slightly better LB score</strong> than the raw file, so in theory that could give me a little edge.</p>\n<h1>2 - Finding a tweakable pipeline</h1>\n<p><a href=\"https://www.kaggle.com/bpetrb\" target=\"_blank\">@bpetrb</a> provided us with a pipeline that were strongly connected with some hyperparameters (i actually got a fork of it), after a night cycling trough a few, i found out that the guy did a good job setting the parameters, <strong>i couldn't improve his score</strong> by changing the parameters alone.</p>\n<h1>3 - Clustering the data and setting hyperparameters</h1>\n<p>After that failed attempt i found <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a> comment explaining how he <strong>clustered the datasets</strong>, i used his logic 'as-is' and tried a new combinations of hyperparameters for each cluster, after a while i found a <strong>better set of parameters</strong> for the highway cluster, since that one was one of the biggest, the improvement was substantial.</p>\n<h1>4 - Finding a good ensembling</h1>\n<p>At the time, <a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a> had a very good ensemble pipeline that were using the same post-processing that i improved before <em>(step 3)</em> as one of its inputs, so i figured that if i already had a better baseline, maybe i could <strong>recalculate the same inputs</strong> that she was using and combine with the best set of hyperparameters that i had.</p>\n<h1>5 - Checking the outputs and snapping to grid</h1>\n<p>During all steps i was constantly checking the outputs on the map, the SJC paths were always very messy and out of the road, so reading the forums, i found that i could implement a **snap to grid **process *(something that i saw before on the indoor competition)* and people were also talking about here, but i couldn't find any public notebook that could do this the way i wanted to, so <strong>i had to code this part</strong>. <br>\nThis is basically the only original code on my solution, since i was out of time <em>(2 days left)</em>, i decided to implement this <strong>without interpolations and only on SJC</strong>, using the provided ground true as grid points (they are dense on SJC), it ended up increasing my score by a good amount, the table below is showing the change it made, so i was confortable submitting that as my final submission.</p>\n<p>Now, looking at the private dataset, this little change ended up giving me a much needed <strong>0.1 on LB</strong>, that is almost the exact difference between my position and the 100 place.</p>\n<table>\n<thead>\n<tr>\n<th>Without snap</th>\n<th>With snap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://64.media.tumblr.com/9076c2d39c15a6484eaaea07d726debe/357c1ccef231f98b-59/s1280x1920/7363d07e30fc4f92e379c9f46f0180dd80b3ddce.png\" alt=\"\"></td>\n<td><img src=\"https://64.media.tumblr.com/bd202b8827e0e531e75c4b909b37e055/ff22d3ad0772e7db-77/s1280x1920/95b14d336214fa6d1d8a48cd9fcd04bebd94e1a9.png\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<p>So that was my final submission, i tried to interpolate the ground true points to create more snapping targets, but i just ran out of time, luckily the public and private LB ended up improving together, so my best private score was indeed my best public one.</p>\n<p>Thank you everyone to read until the end, i'm sorry if i didn't gave credit to someone who was forked before the guys i forked, if so, please comment, i'll gladly extend my thanks to you too!</p>",
  "messages": [
    {
      "id": "1462694",
      "postDate": "08/10/2021 01:07:27",
      "content": "<p><strong>tl;dr: New baseline from raw data &gt; Recalculate all pipelines on the best ensemble &gt; Snap SJC to ground true</strong></p>\n<p>Hello everyone, i wanted to share the steps i took to get the bronze medal, i only started working on this competition 10 days before the deadline and mostly used work from the community to get my final score, so i'll also leave my special thanks to the ones who shared the knowledge i used.</p>\n<p><a href=\"https://www.kaggle.com/jeongyoonlee\" target=\"_blank\">@jeongyoonlee</a> for this great baseline<br>\n<a href=\"https://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu\" target=\"_blank\">https://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu</a></p>\n<p><a href=\"https://www.kaggle.com/bpetrb\" target=\"_blank\">@bpetrb</a> for this amazing notebook<br>\n<a href=\"https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output\" target=\"_blank\">https://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output</a></p>\n<p><a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a> for his clustering technique<br>\n<a href=\"https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160\" target=\"_blank\">https://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160</a></p>\n<p><a href=\"https://www.kaggle.com/t88take\" target=\"_blank\">@t88take</a> for this great notebook<br>\n<a href=\"https://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output\" target=\"_blank\">https://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output</a></p>\n<p><a href=\"https://www.kaggle.com/tensorchoko\" target=\"_blank\">@tensorchoko</a> for her multioutput regressor, i believe she removed the notebook since i can't find it anymore</p>\n<p><a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a> for this amazing ensemble notebook<br>\n<a href=\"https://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling\" target=\"_blank\">https://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling</a></p>\n<p>The combined output from those notebooks plus a little original code granted me the medal, this is how it was done:</p>\n<h1>1 - Get a better baseline</h1>\n<p>On the first day of my research when i was trying to understand the data, i could see people talking about the <strong>base_locations_test.csv</strong> file was a good starting point and doing many post-processing techniques on it, so if i could find a better baseline, i could just apply the same public techniques people were doing and i would probably end up with a better result, so i spend some time trying to find someone who had calculated the baseline again, i end up finding a fork of <a href=\"https://www.kaggle.com/jeongyoonlee\" target=\"_blank\">@jeongyoonlee</a> code and after a few tweaks, i ended up with a <strong>slightly better LB score</strong> than the raw file, so in theory that could give me a little edge.</p>\n<h1>2 - Finding a tweakable pipeline</h1>\n<p><a href=\"https://www.kaggle.com/bpetrb\" target=\"_blank\">@bpetrb</a> provided us with a pipeline that were strongly connected with some hyperparameters (i actually got a fork of it), after a night cycling trough a few, i found out that the guy did a good job setting the parameters, <strong>i couldn't improve his score</strong> by changing the parameters alone.</p>\n<h1>3 - Clustering the data and setting hyperparameters</h1>\n<p>After that failed attempt i found <a href=\"https://www.kaggle.com/taroz1461\" target=\"_blank\">@taroz1461</a> comment explaining how he <strong>clustered the datasets</strong>, i used his logic 'as-is' and tried a new combinations of hyperparameters for each cluster, after a while i found a <strong>better set of parameters</strong> for the highway cluster, since that one was one of the biggest, the improvement was substantial.</p>\n<h1>4 - Finding a good ensembling</h1>\n<p>At the time, <a href=\"https://www.kaggle.com/somayyehgholami\" target=\"_blank\">@somayyehgholami</a> had a very good ensemble pipeline that were using the same post-processing that i improved before <em>(step 3)</em> as one of its inputs, so i figured that if i already had a better baseline, maybe i could <strong>recalculate the same inputs</strong> that she was using and combine with the best set of hyperparameters that i had.</p>\n<h1>5 - Checking the outputs and snapping to grid</h1>\n<p>During all steps i was constantly checking the outputs on the map, the SJC paths were always very messy and out of the road, so reading the forums, i found that i could implement a **snap to grid **process *(something that i saw before on the indoor competition)* and people were also talking about here, but i couldn't find any public notebook that could do this the way i wanted to, so <strong>i had to code this part</strong>. <br>\nThis is basically the only original code on my solution, since i was out of time <em>(2 days left)</em>, i decided to implement this <strong>without interpolations and only on SJC</strong>, using the provided ground true as grid points (they are dense on SJC), it ended up increasing my score by a good amount, the table below is showing the change it made, so i was confortable submitting that as my final submission.</p>\n<p>Now, looking at the private dataset, this little change ended up giving me a much needed <strong>0.1 on LB</strong>, that is almost the exact difference between my position and the 100 place.</p>\n<table>\n<thead>\n<tr>\n<th>Without snap</th>\n<th>With snap</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://64.media.tumblr.com/9076c2d39c15a6484eaaea07d726debe/357c1ccef231f98b-59/s1280x1920/7363d07e30fc4f92e379c9f46f0180dd80b3ddce.png\" alt=\"\"></td>\n<td><img src=\"https://64.media.tumblr.com/bd202b8827e0e531e75c4b909b37e055/ff22d3ad0772e7db-77/s1280x1920/95b14d336214fa6d1d8a48cd9fcd04bebd94e1a9.png\" alt=\"\"></td>\n</tr>\n</tbody>\n</table>\n<p>So that was my final submission, i tried to interpolate the ground true points to create more snapping targets, but i just ran out of time, luckily the public and private LB ended up improving together, so my best private score was indeed my best public one.</p>\n<p>Thank you everyone to read until the end, i'm sorry if i didn't gave credit to someone who was forked before the guys i forked, if so, please comment, i'll gladly extend my thanks to you too!</p>",
      "rawMarkdown": "**tl;dr: New baseline from raw data > Recalculate all pipelines on the best ensemble > Snap SJC to ground true**\n\nHello everyone, i wanted to share the steps i took to get the bronze medal, i only started working on this competition 10 days before the deadline and mostly used work from the community to get my final score, so i'll also leave my special thanks to the ones who shared the knowledge i used.\n\n@jeongyoonlee for this great baseline\nhttps://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu\n\n@bpetrb for this amazing notebook\nhttps://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output\n\n@taroz1461 for his clustering technique\nhttps://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160\n\n@t88take for this great notebook\nhttps://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output\n\n@tensorchoko for her multioutput regressor, i believe she removed the notebook since i can't find it anymore\n\n@somayyehgholami for this amazing ensemble notebook\nhttps://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling\n\nThe combined output from those notebooks plus a little original code granted me the medal, this is how it was done:\n\n# 1 - Get a better baseline\nOn the first day of my research when i was trying to understand the data, i could see people talking about the **base_locations_test.csv** file was a good starting point and doing many post-processing techniques on it, so if i could find a better baseline, i could just apply the same public techniques people were doing and i would probably end up with a better result, so i spend some time trying to find someone who had calculated the baseline again, i end up finding a fork of @jeongyoonlee code and after a few tweaks, i ended up with a **slightly better LB score** than the raw file, so in theory that could give me a little edge.\n\n# 2 - Finding a tweakable pipeline\n@bpetrb provided us with a pipeline that were strongly connected with some hyperparameters (i actually got a fork of it), after a night cycling trough a few, i found out that the guy did a good job setting the parameters, **i couldn't improve his score** by changing the parameters alone.\n\n# 3 - Clustering the data and setting hyperparameters\nAfter that failed attempt i found @taroz1461 comment explaining how he **clustered the datasets**, i used his logic 'as-is' and tried a new combinations of hyperparameters for each cluster, after a while i found a **better set of parameters** for the highway cluster, since that one was one of the biggest, the improvement was substantial.\n\n# 4 - Finding a good ensembling\nAt the time, @somayyehgholami had a very good ensemble pipeline that were using the same post-processing that i improved before *(step 3)* as one of its inputs, so i figured that if i already had a better baseline, maybe i could **recalculate the same inputs** that she was using and combine with the best set of hyperparameters that i had.\n\n# 5 - Checking the outputs and snapping to grid\nDuring all steps i was constantly checking the outputs on the map, the SJC paths were always very messy and out of the road, so reading the forums, i found that i could implement a **snap to grid **process *(something that i saw before on the indoor competition)* and people were also talking about here, but i couldn't find any public notebook that could do this the way i wanted to, so **i had to code this part**. \nThis is basically the only original code on my solution, since i was out of time *(2 days left)*, i decided to implement this **without interpolations and only on SJC**, using the provided ground true as grid points (they are dense on SJC), it ended up increasing my score by a good amount, the table below is showing the change it made, so i was confortable submitting that as my final submission.\n\nNow, looking at the private dataset, this little change ended up giving me a much needed **0.1 on LB**, that is almost the exact difference between my position and the 100 place.\n\n| Without snap | With snap |\n| --- | --- |\n| ![](https://64.media.tumblr.com/9076c2d39c15a6484eaaea07d726debe/357c1ccef231f98b-59/s1280x1920/7363d07e30fc4f92e379c9f46f0180dd80b3ddce.png) | ![](https://64.media.tumblr.com/bd202b8827e0e531e75c4b909b37e055/ff22d3ad0772e7db-77/s1280x1920/95b14d336214fa6d1d8a48cd9fcd04bebd94e1a9.png) |\n\n\nSo that was my final submission, i tried to interpolate the ground true points to create more snapping targets, but i just ran out of time, luckily the public and private LB ended up improving together, so my best private score was indeed my best public one.\n\nThank you everyone to read until the end, i'm sorry if i didn't gave credit to someone who was forked before the guys i forked, if so, please comment, i'll gladly extend my thanks to you too!",
      "votes": null
    },
    {
      "id": "1462974",
      "postDate": "08/10/2021 03:58:50",
      "content": "<p>happy to  you a lot. I was not good score. 😂</p>",
      "rawMarkdown": "happy to  you a lot. I was not good score. 😂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1462974,
      "author_name": "tensorchoko",
      "author_url": "",
      "post_date": "08/10/2021 03:58:50",
      "content": "<p>happy to  you a lot. I was not good score. 😂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1462694": "**tl;dr: New baseline from raw data > Recalculate all pipelines on the best ensemble > Snap SJC to ground true**\n\nHello everyone, i wanted to share the steps i took to get the bronze medal, i only started working on this competition 10 days before the deadline and mostly used work from the community to get my final score, so i'll also leave my special thanks to the ones who shared the knowledge i used.\n\n@jeongyoonlee for this great baseline\nhttps://www.kaggle.com/jeongyoonlee/google-smartphone-decimeter-eda-keras-tpu\n\n@bpetrb for this amazing notebook\nhttps://www.kaggle.com/bpetrb/adaptive-gauss-phone-mean/output\n\n@taroz1461 for his clustering technique\nhttps://www.kaggle.com/c/google-smartphone-decimeter-challenge/discussion/245160\n\n@t88take for this great notebook\nhttps://www.kaggle.com/t88take/gsdc-phones-mean-prediction/output\n\n@tensorchoko for her multioutput regressor, i believe she removed the notebook since i can't find it anymore\n\n@somayyehgholami for this amazing ensemble notebook\nhttps://www.kaggle.com/somayyehgholami/gsdc-smart-ensembling\n\nThe combined output from those notebooks plus a little original code granted me the medal, this is how it was done:\n\n# 1 - Get a better baseline\nOn the first day of my research when i was trying to understand the data, i could see people talking about the **base_locations_test.csv** file was a good starting point and doing many post-processing techniques on it, so if i could find a better baseline, i could just apply the same public techniques people were doing and i would probably end up with a better result, so i spend some time trying to find someone who had calculated the baseline again, i end up finding a fork of @jeongyoonlee code and after a few tweaks, i ended up with a **slightly better LB score** than the raw file, so in theory that could give me a little edge.\n\n# 2 - Finding a tweakable pipeline\n@bpetrb provided us with a pipeline that were strongly connected with some hyperparameters (i actually got a fork of it), after a night cycling trough a few, i found out that the guy did a good job setting the parameters, **i couldn't improve his score** by changing the parameters alone.\n\n# 3 - Clustering the data and setting hyperparameters\nAfter that failed attempt i found @taroz1461 comment explaining how he **clustered the datasets**, i used his logic 'as-is' and tried a new combinations of hyperparameters for each cluster, after a while i found a **better set of parameters** for the highway cluster, since that one was one of the biggest, the improvement was substantial.\n\n# 4 - Finding a good ensembling\nAt the time, @somayyehgholami had a very good ensemble pipeline that were using the same post-processing that i improved before *(step 3)* as one of its inputs, so i figured that if i already had a better baseline, maybe i could **recalculate the same inputs** that she was using and combine with the best set of hyperparameters that i had.\n\n# 5 - Checking the outputs and snapping to grid\nDuring all steps i was constantly checking the outputs on the map, the SJC paths were always very messy and out of the road, so reading the forums, i found that i could implement a **snap to grid **process *(something that i saw before on the indoor competition)* and people were also talking about here, but i couldn't find any public notebook that could do this the way i wanted to, so **i had to code this part**. \nThis is basically the only original code on my solution, since i was out of time *(2 days left)*, i decided to implement this **without interpolations and only on SJC**, using the provided ground true as grid points (they are dense on SJC), it ended up increasing my score by a good amount, the table below is showing the change it made, so i was confortable submitting that as my final submission.\n\nNow, looking at the private dataset, this little change ended up giving me a much needed **0.1 on LB**, that is almost the exact difference between my position and the 100 place.\n\n| Without snap | With snap |\n| --- | --- |\n| ![](https://64.media.tumblr.com/9076c2d39c15a6484eaaea07d726debe/357c1ccef231f98b-59/s1280x1920/7363d07e30fc4f92e379c9f46f0180dd80b3ddce.png) | ![](https://64.media.tumblr.com/bd202b8827e0e531e75c4b909b37e055/ff22d3ad0772e7db-77/s1280x1920/95b14d336214fa6d1d8a48cd9fcd04bebd94e1a9.png) |\n\n\nSo that was my final submission, i tried to interpolate the ground true points to create more snapping targets, but i just ran out of time, luckily the public and private LB ended up improving together, so my best private score was indeed my best public one.\n\nThank you everyone to read until the end, i'm sorry if i didn't gave credit to someone who was forked before the guys i forked, if so, please comment, i'll gladly extend my thanks to you too!",
    "1462974": "happy to  you a lot. I was not good score. 😂"
  },
  "source": "meta"
}