{
  "id": 240077,
  "title": "[17th] 🔥 PERFECT cv 🔥  & improve step by step [中英]",
  "url": "/competitions/indoor-location-navigation/writeups/max2020-17th-perfect-cv-improve-step-by-step",
  "author_name": "",
  "post_date": "2021-05-18T13:55:28.843Z",
  "votes": 11,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>This competition is very good, thanks to XYZ10 Technology, thanks to the organizer of kaggle, I love kaggle so much. Completing this competition I think is a very good growth for me. I have learned (reviewed) a lot of things, including: tensorfloor, pytorch, multi-threading skills, code specifications, standardized experimental records, and the importance of cv . I haven't made any breakthroughs in algorithms and features. The only thing I haven't seen anyone discuss is the construction method of my cv. I will share it later.</p>\n<h1>step by step</h1>\n<p>Now let’s take a look at my main improvement points in the past month:</p>\n<table>\n<thead>\n<tr>\n<th>Mainline</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Lstm public kernel is used with simple parameter adjustment, also use cost min+snap2grid</td>\n<td>4.72</td>\n</tr>\n<tr>\n<td>A fake test is constructed by simulating the characteristics of test data</td>\n<td></td>\n</tr>\n<tr>\n<td>Interpolate the wifi timestamp of the training set, the test set is still the timestamp of the waypoint</td>\n<td>4.45</td>\n</tr>\n<tr>\n<td>Drop out wifi log with an interval greater than 5 seconds</td>\n<td>4.20</td>\n</tr>\n<tr>\n<td>By find a timestamp very close to the train, fix start and end point</td>\n<td>4.18</td>\n</tr>\n<tr>\n<td>The training set is unchanged, the test set is also the wifi timestamp, and then de-interpolated to get the waypoints</td>\n<td>4.16</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and finally do cost+snap again</td>\n<td>4.0x-390</td>\n</tr>\n<tr>\n<td>4 fold use fake test strategy</td>\n<td>3.82</td>\n</tr>\n<tr>\n<td>More careful fix start and end points</td>\n<td>3.67</td>\n</tr>\n<tr>\n<td>Duild delta modeling of wifi timestamp granularity improves the result of wifi cost min</td>\n<td>3.46</td>\n</tr>\n<tr>\n<td>Duild delta modeling of waypoint timestamp granularity improves the result of wp cost min</td>\n<td>3.36</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and twice cost+snap+fix</td>\n<td>3.23</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and three times cost+snap+fix</td>\n<td>3.19</td>\n</tr>\n</tbody>\n</table>\n<h1>fake test</h1>\n<p>When I did simple path group Nfold for the first time, there was no improvement. I realized that the distribution of the training set and the test set was inconsistent. I conducted a series of probes and found that the test set has some characteristics that should be artificially limited. <strong>The txt size of the test set&gt;=2M (look in linux), the time span&gt;=60s, and the number of path points&gt;=5</strong>. In the first half of the month, i also selected 626 paths as fake tests using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. In the last week, I used 4-fold cross-validation. I also made the above restrictions to make my cv more consistent. Later, I checked the score of private, which also conformed to the same trend.</p>\n<h1>Failed attempt</h1>\n<ol>\n<li>Directly use png image features as end-to-end input</li>\n<li>Use pytorch to reproduce the network structure and lb score</li>\n<li>Use MLP to replace lstm that is not a time series</li>\n<li>Use all site data</li>\n<li>Try sequential RNN, but it does not work well and it is difficult to converge <br>\n. . . . . .</li>\n</ol>\n<h1>I have a question</h1>\n<p>I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .<br>\n<strong>If you have any comments on this. I will be very grateful.</strong></p>\n<hr>\n<hr>\n<p>大家好，</p>\n<p>这个比赛非常棒，感谢十域科技，感谢kaggle主办方，我太爱kaggle了。完成这个比赛我觉得是对我个人的一次非常好的成长，我学到了（复习了）非常多的东西，包括：tensorfloor、pytorch、多线程技巧、代码规范、规范的实验记录、cv的重要性。我没有在算法和特征上有什么突破，唯一有一点我没有看到有人讨论的是我cv的构建方法。后面我将会分享到它。</p>\n<h1>step by step</h1>\n<p>现在先看一下我这一个月来的主要提升点：</p>\n<table>\n<thead>\n<tr>\n<th>主线</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>使用了 lstm public kernel，进行简单调参，使用cost min+snap2grid</td>\n<td>4.72</td>\n</tr>\n<tr>\n<td>模拟test data 的特点构造了一个 fake test</td>\n<td></td>\n</tr>\n<tr>\n<td>对训练集wifi时间戳进行插值，测试集还是路径点的时间戳</td>\n<td>4.45</td>\n</tr>\n<tr>\n<td>过滤掉间隔大于5秒的wifi</td>\n<td>4.20</td>\n</tr>\n<tr>\n<td>寻找时间非常接近的点，修复开始结束点</td>\n<td>4.18</td>\n</tr>\n<tr>\n<td>训练集不变，测试集也是wifi时间戳，再反插值得到路径点</td>\n<td>4.16</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，最后再嵌套一次cost+snap</td>\n<td>4.0x-390</td>\n</tr>\n<tr>\n<td>引入4折 fake test</td>\n<td>3.82</td>\n</tr>\n<tr>\n<td>更小心的修复开始和结束点</td>\n<td>3.67</td>\n</tr>\n<tr>\n<td>对wifi时间戳粒度的delta建模提高wifi cost min的结果</td>\n<td>3.46</td>\n</tr>\n<tr>\n<td>对wp时间戳粒度误差建模提高cost min</td>\n<td>3.36</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套两次cost+snap+fix</td>\n<td>3.23</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套三次cost+snap+fix</td>\n<td>3.19</td>\n</tr>\n</tbody>\n</table>\n<h1>fake test</h1>\n<p>第一次做Nfold的时候没有提升，我意识到训练集和测试集的分布不一致，我进行了一系列的探测，发现测试集有一些特点应该是人为限定的。测试集的txt大小&gt;=2M（linux下看），时间跨度&gt;=60s，路径点个数&gt;=5。在前半个月的时候，这个限制同样抽出了626个路径作为fake test，无论是否进行后处理，我的fake test得分和线上的得分都保持一致。在最后1周我使用4折交叉验证，作为valid的那一份数据我同样做了上述限制，使得我的cv更加一致。后面我检查了private的分数情况，也符合相同的趋势。</p>\n<h1>失败的尝试</h1>\n<p>1.直接使用png图像特征作为端到端的输入<br>\n2.使用pytorch复现网络结构和lb分数<br>\n3.使用MLP替换并不是时间序列的lstm<br>\n4.使用所有site的数据<br>\n5.尝试序列RNN，但是它效果不好且难以收敛<br>\n。。。。。。</p>\n<h1>我有一个问题</h1>\n<p>我这一个月始终都没有想明白，为什么序列长度为1的lstm模型工作的这么好，mlp为什么无法替代它。。。<br>\n如果您对此发表任何看法。我将非常感激。</p>",
  "messages": [
    {
      "id": "1313271",
      "postDate": "05/18/2021 13:32:07",
      "content": "<p>Hello everyone,</p>\n<p>This competition is very good, thanks to XYZ10 Technology, thanks to the organizer of kaggle, I love kaggle so much. Completing this competition I think is a very good growth for me. I have learned (reviewed) a lot of things, including: tensorfloor, pytorch, multi-threading skills, code specifications, standardized experimental records, and the importance of cv . I haven't made any breakthroughs in algorithms and features. The only thing I haven't seen anyone discuss is the construction method of my cv. I will share it later.</p>\n<h1>step by step</h1>\n<p>Now let’s take a look at my main improvement points in the past month:</p>\n<table>\n<thead>\n<tr>\n<th>Mainline</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Lstm public kernel is used with simple parameter adjustment, also use cost min+snap2grid</td>\n<td>4.72</td>\n</tr>\n<tr>\n<td>A fake test is constructed by simulating the characteristics of test data</td>\n<td></td>\n</tr>\n<tr>\n<td>Interpolate the wifi timestamp of the training set, the test set is still the timestamp of the waypoint</td>\n<td>4.45</td>\n</tr>\n<tr>\n<td>Drop out wifi log with an interval greater than 5 seconds</td>\n<td>4.20</td>\n</tr>\n<tr>\n<td>By find a timestamp very close to the train, fix start and end point</td>\n<td>4.18</td>\n</tr>\n<tr>\n<td>The training set is unchanged, the test set is also the wifi timestamp, and then de-interpolated to get the waypoints</td>\n<td>4.16</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and finally do cost+snap again</td>\n<td>4.0x-390</td>\n</tr>\n<tr>\n<td>4 fold use fake test strategy</td>\n<td>3.82</td>\n</tr>\n<tr>\n<td>More careful fix start and end points</td>\n<td>3.67</td>\n</tr>\n<tr>\n<td>Duild delta modeling of wifi timestamp granularity improves the result of wifi cost min</td>\n<td>3.46</td>\n</tr>\n<tr>\n<td>Duild delta modeling of waypoint timestamp granularity improves the result of wp cost min</td>\n<td>3.36</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and twice cost+snap+fix</td>\n<td>3.23</td>\n</tr>\n<tr>\n<td>Do cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and three times cost+snap+fix</td>\n<td>3.19</td>\n</tr>\n</tbody>\n</table>\n<h1>fake test</h1>\n<p>When I did simple path group Nfold for the first time, there was no improvement. I realized that the distribution of the training set and the test set was inconsistent. I conducted a series of probes and found that the test set has some characteristics that should be artificially limited. <strong>The txt size of the test set&gt;=2M (look in linux), the time span&gt;=60s, and the number of path points&gt;=5</strong>. In the first half of the month, i also selected 626 paths as fake tests using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. In the last week, I used 4-fold cross-validation. I also made the above restrictions to make my cv more consistent. Later, I checked the score of private, which also conformed to the same trend.</p>\n<h1>Failed attempt</h1>\n<ol>\n<li>Directly use png image features as end-to-end input</li>\n<li>Use pytorch to reproduce the network structure and lb score</li>\n<li>Use MLP to replace lstm that is not a time series</li>\n<li>Use all site data</li>\n<li>Try sequential RNN, but it does not work well and it is difficult to converge <br>\n. . . . . .</li>\n</ol>\n<h1>I have a question</h1>\n<p>I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .<br>\n<strong>If you have any comments on this. I will be very grateful.</strong></p>\n<hr>\n<hr>\n<p>大家好，</p>\n<p>这个比赛非常棒，感谢十域科技，感谢kaggle主办方，我太爱kaggle了。完成这个比赛我觉得是对我个人的一次非常好的成长，我学到了（复习了）非常多的东西，包括：tensorfloor、pytorch、多线程技巧、代码规范、规范的实验记录、cv的重要性。我没有在算法和特征上有什么突破，唯一有一点我没有看到有人讨论的是我cv的构建方法。后面我将会分享到它。</p>\n<h1>step by step</h1>\n<p>现在先看一下我这一个月来的主要提升点：</p>\n<table>\n<thead>\n<tr>\n<th>主线</th>\n<th>public</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>使用了 lstm public kernel，进行简单调参，使用cost min+snap2grid</td>\n<td>4.72</td>\n</tr>\n<tr>\n<td>模拟test data 的特点构造了一个 fake test</td>\n<td></td>\n</tr>\n<tr>\n<td>对训练集wifi时间戳进行插值，测试集还是路径点的时间戳</td>\n<td>4.45</td>\n</tr>\n<tr>\n<td>过滤掉间隔大于5秒的wifi</td>\n<td>4.20</td>\n</tr>\n<tr>\n<td>寻找时间非常接近的点，修复开始结束点</td>\n<td>4.18</td>\n</tr>\n<tr>\n<td>训练集不变，测试集也是wifi时间戳，再反插值得到路径点</td>\n<td>4.16</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，最后再嵌套一次cost+snap</td>\n<td>4.0x-390</td>\n</tr>\n<tr>\n<td>引入4折 fake test</td>\n<td>3.82</td>\n</tr>\n<tr>\n<td>更小心的修复开始和结束点</td>\n<td>3.67</td>\n</tr>\n<tr>\n<td>对wifi时间戳粒度的delta建模提高wifi cost min的结果</td>\n<td>3.46</td>\n</tr>\n<tr>\n<td>对wp时间戳粒度误差建模提高cost min</td>\n<td>3.36</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套两次cost+snap+fix</td>\n<td>3.23</td>\n</tr>\n<tr>\n<td>做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套三次cost+snap+fix</td>\n<td>3.19</td>\n</tr>\n</tbody>\n</table>\n<h1>fake test</h1>\n<p>第一次做Nfold的时候没有提升，我意识到训练集和测试集的分布不一致，我进行了一系列的探测，发现测试集有一些特点应该是人为限定的。测试集的txt大小&gt;=2M（linux下看），时间跨度&gt;=60s，路径点个数&gt;=5。在前半个月的时候，这个限制同样抽出了626个路径作为fake test，无论是否进行后处理，我的fake test得分和线上的得分都保持一致。在最后1周我使用4折交叉验证，作为valid的那一份数据我同样做了上述限制，使得我的cv更加一致。后面我检查了private的分数情况，也符合相同的趋势。</p>\n<h1>失败的尝试</h1>\n<p>1.直接使用png图像特征作为端到端的输入<br>\n2.使用pytorch复现网络结构和lb分数<br>\n3.使用MLP替换并不是时间序列的lstm<br>\n4.使用所有site的数据<br>\n5.尝试序列RNN，但是它效果不好且难以收敛<br>\n。。。。。。</p>\n<h1>我有一个问题</h1>\n<p>我这一个月始终都没有想明白，为什么序列长度为1的lstm模型工作的这么好，mlp为什么无法替代它。。。<br>\n如果您对此发表任何看法。我将非常感激。</p>",
      "rawMarkdown": "Hello everyone,\n\n\nThis competition is very good, thanks to XYZ10 Technology, thanks to the organizer of kaggle, I love kaggle so much. Completing this competition I think is a very good growth for me. I have learned (reviewed) a lot of things, including: tensorfloor, pytorch, multi-threading skills, code specifications, standardized experimental records, and the importance of cv . I haven't made any breakthroughs in algorithms and features. The only thing I haven't seen anyone discuss is the construction method of my cv. I will share it later.\n# step by step\nNow let’s take a look at my main improvement points in the past month:\n\n| Mainline | public |\n| --- | --- |\nLstm public kernel is used with simple parameter adjustment, also use cost min+snap2grid | 4.72\nA fake test is constructed by simulating the characteristics of test data|\nInterpolate the wifi timestamp of the training set, the test set is still the timestamp of the waypoint|4.45\nDrop out wifi log with an interval greater than 5 seconds | 4.20\nBy find a timestamp very close to the train, fix start and end point| 4.18\nThe training set is unchanged, the test set is also the wifi timestamp, and then de-interpolated to get the waypoints | 4.16\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and finally do cost+snap again| 4.0x-390\n4 fold use fake test strategy| 3.82\nMore careful fix start and end points | 3.67\nDuild delta modeling of wifi timestamp granularity improves the result of wifi cost min | 3.46\nDuild delta modeling of waypoint timestamp granularity improves the result of wp cost min| 3.36\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and twice cost+snap+fix| 3.23\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and three times cost+snap+fix | 3.19\n\n# fake test\nWhen I did simple path group Nfold for the first time, there was no improvement. I realized that the distribution of the training set and the test set was inconsistent. I conducted a series of probes and found that the test set has some characteristics that should be artificially limited. **The txt size of the test set>=2M (look in linux), the time span>=60s, and the number of path points>=5**. In the first half of the month, i also selected 626 paths as fake tests using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. In the last week, I used 4-fold cross-validation. I also made the above restrictions to make my cv more consistent. Later, I checked the score of private, which also conformed to the same trend.\n\n\n# Failed attempt\n1. Directly use png image features as end-to-end input\n2. Use pytorch to reproduce the network structure and lb score\n3. Use MLP to replace lstm that is not a time series\n4. Use all site data\n5. Try sequential RNN, but it does not work well and it is difficult to converge \n. . . . . .\n\n# I have a question\nI haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .\n**If you have any comments on this. I will be very grateful.**\n\n<hr>\n<hr>\n\n大家好，\n\n这个比赛非常棒，感谢十域科技，感谢kaggle主办方，我太爱kaggle了。完成这个比赛我觉得是对我个人的一次非常好的成长，我学到了（复习了）非常多的东西，包括：tensorfloor、pytorch、多线程技巧、代码规范、规范的实验记录、cv的重要性。我没有在算法和特征上有什么突破，唯一有一点我没有看到有人讨论的是我cv的构建方法。后面我将会分享到它。\n\n# step by step\n现在先看一下我这一个月来的主要提升点：\n\n| 主线| public | \n| --- | --- |\n使用了 lstm public kernel，进行简单调参，使用cost min+snap2grid|  4.72\n模拟test data 的特点构造了一个 fake test|\n对训练集wifi时间戳进行插值，测试集还是路径点的时间戳|4.45\n过滤掉间隔大于5秒的wifi | 4.20\n寻找时间非常接近的点，修复开始结束点  | 4.18\n训练集不变，测试集也是wifi时间戳，再反插值得到路径点|  4.16\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，最后再嵌套一次cost+snap| 4.0x-390\n引入4折 fake test |  3.82\n更小心的修复开始和结束点 |   3.67\n对wifi时间戳粒度的delta建模提高wifi cost min的结果|   3.46\n对wp时间戳粒度误差建模提高cost min|  3.36\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套两次cost+snap+fix|   3.23\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套三次cost+snap+fix |   3.19\n\n# fake test\n第一次做Nfold的时候没有提升，我意识到训练集和测试集的分布不一致，我进行了一系列的探测，发现测试集有一些特点应该是人为限定的。测试集的txt大小>=2M（linux下看），时间跨度>=60s，路径点个数>=5。在前半个月的时候，这个限制同样抽出了626个路径作为fake test，无论是否进行后处理，我的fake test得分和线上的得分都保持一致。在最后1周我使用4折交叉验证，作为valid的那一份数据我同样做了上述限制，使得我的cv更加一致。后面我检查了private的分数情况，也符合相同的趋势。\n\n# 失败的尝试\n1.直接使用png图像特征作为端到端的输入\n2.使用pytorch复现网络结构和lb分数\n3.使用MLP替换并不是时间序列的lstm\n4.使用所有site的数据\n5.尝试序列RNN，但是它效果不好且难以收敛\n。。。。。。\n\n# 我有一个问题\n我这一个月始终都没有想明白，为什么序列长度为1的lstm模型工作的这么好，mlp为什么无法替代它。。。\n如果您对此发表任何看法。我将非常感激。",
      "votes": null
    },
    {
      "id": "1313706",
      "postDate": "05/18/2021 17:08:04",
      "content": "<p>Thanks for sharing. I also learned a lot from this competition.</p>\n<blockquote>\n  <p>I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .</p>\n</blockquote>\n<p>I am new to neural networks and have also been troubled with this question for a long time. My understanding is :<br>\n1)A single-cell LSTM has 4.x times params compared to a Dense layer with the same units.  <br>\n2)LSTM has a different mathematical equation that involves sigmoid and tanh (or relu) activations and element-wise product. And this structure might perform better than a single dense layer with only one activation.<br>\nActually, I find my modified MLP model performs slightly better than LSTM. </p>",
      "rawMarkdown": "Thanks for sharing. I also learned a lot from this competition.\n> I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .\n\nI am new to neural networks and have also been troubled with this question for a long time. My understanding is :\n1)A single-cell LSTM has 4.x times params compared to a Dense layer with the same units.  \n2)LSTM has a different mathematical equation that involves sigmoid and tanh (or relu) activations and element-wise product. And this structure might perform better than a single dense layer with only one activation.\nActually, I find my modified MLP model performs slightly better than LSTM.",
      "votes": null
    },
    {
      "id": "1314105",
      "postDate": "05/19/2021 00:42:33",
      "content": "<blockquote>\n  <p>I find my modified MLP model performs slightly better than LSTM.</p>\n</blockquote>\n<p>great，make sense， Could you please share your network structure？</p>",
      "rawMarkdown": "> I find my modified MLP model performs slightly better than LSTM.\n\ngreat，make sense， Could you please share your network structure？",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1313706,
      "author_name": "hakase1",
      "author_url": "",
      "post_date": "05/18/2021 17:08:04",
      "content": "<p>Thanks for sharing. I also learned a lot from this competition.</p>\n<blockquote>\n  <p>I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .</p>\n</blockquote>\n<p>I am new to neural networks and have also been troubled with this question for a long time. My understanding is :<br>\n1)A single-cell LSTM has 4.x times params compared to a Dense layer with the same units.  <br>\n2)LSTM has a different mathematical equation that involves sigmoid and tanh (or relu) activations and element-wise product. And this structure might perform better than a single dense layer with only one activation.<br>\nActually, I find my modified MLP model performs slightly better than LSTM. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1314105,
          "author_name": "max2020",
          "author_url": "",
          "post_date": "05/19/2021 00:42:33",
          "content": "<blockquote>\n  <p>I find my modified MLP model performs slightly better than LSTM.</p>\n</blockquote>\n<p>great，make sense， Could you please share your network structure？</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1313271": "Hello everyone,\n\n\nThis competition is very good, thanks to XYZ10 Technology, thanks to the organizer of kaggle, I love kaggle so much. Completing this competition I think is a very good growth for me. I have learned (reviewed) a lot of things, including: tensorfloor, pytorch, multi-threading skills, code specifications, standardized experimental records, and the importance of cv . I haven't made any breakthroughs in algorithms and features. The only thing I haven't seen anyone discuss is the construction method of my cv. I will share it later.\n# step by step\nNow let’s take a look at my main improvement points in the past month:\n\n| Mainline | public |\n| --- | --- |\nLstm public kernel is used with simple parameter adjustment, also use cost min+snap2grid | 4.72\nA fake test is constructed by simulating the characteristics of test data|\nInterpolate the wifi timestamp of the training set, the test set is still the timestamp of the waypoint|4.45\nDrop out wifi log with an interval greater than 5 seconds | 4.20\nBy find a timestamp very close to the train, fix start and end point| 4.18\nThe training set is unchanged, the test set is also the wifi timestamp, and then de-interpolated to get the waypoints | 4.16\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and finally do cost+snap again| 4.0x-390\n4 fold use fake test strategy| 3.82\nMore careful fix start and end points | 3.67\nDuild delta modeling of wifi timestamp granularity improves the result of wifi cost min | 3.46\nDuild delta modeling of waypoint timestamp granularity improves the result of wp cost min| 3.36\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and twice cost+snap+fix| 3.23\nDo cost min at the wifi timestamp level, interpolate to get wp, <br>snap, fix the start and end points, and three times cost+snap+fix | 3.19\n\n# fake test\nWhen I did simple path group Nfold for the first time, there was no improvement. I realized that the distribution of the training set and the test set was inconsistent. I conducted a series of probes and found that the test set has some characteristics that should be artificially limited. **The txt size of the test set>=2M (look in linux), the time span>=60s, and the number of path points>=5**. In the first half of the month, i also selected 626 paths as fake tests using above restrictions . No matter whether post-processing is performed or not, my fake test scores are consistent with the lb scores. In the last week, I used 4-fold cross-validation. I also made the above restrictions to make my cv more consistent. Later, I checked the score of private, which also conformed to the same trend.\n\n\n# Failed attempt\n1. Directly use png image features as end-to-end input\n2. Use pytorch to reproduce the network structure and lb score\n3. Use MLP to replace lstm that is not a time series\n4. Use all site data\n5. Try sequential RNN, but it does not work well and it is difficult to converge \n. . . . . .\n\n# I have a question\nI haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .\n**If you have any comments on this. I will be very grateful.**\n\n<hr>\n<hr>\n\n大家好，\n\n这个比赛非常棒，感谢十域科技，感谢kaggle主办方，我太爱kaggle了。完成这个比赛我觉得是对我个人的一次非常好的成长，我学到了（复习了）非常多的东西，包括：tensorfloor、pytorch、多线程技巧、代码规范、规范的实验记录、cv的重要性。我没有在算法和特征上有什么突破，唯一有一点我没有看到有人讨论的是我cv的构建方法。后面我将会分享到它。\n\n# step by step\n现在先看一下我这一个月来的主要提升点：\n\n| 主线| public | \n| --- | --- |\n使用了 lstm public kernel，进行简单调参，使用cost min+snap2grid|  4.72\n模拟test data 的特点构造了一个 fake test|\n对训练集wifi时间戳进行插值，测试集还是路径点的时间戳|4.45\n过滤掉间隔大于5秒的wifi | 4.20\n寻找时间非常接近的点，修复开始结束点  | 4.18\n训练集不变，测试集也是wifi时间戳，再反插值得到路径点|  4.16\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，最后再嵌套一次cost+snap| 4.0x-390\n引入4折 fake test |  3.82\n更小心的修复开始和结束点 |   3.67\n对wifi时间戳粒度的delta建模提高wifi cost min的结果|   3.46\n对wp时间戳粒度误差建模提高cost min|  3.36\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套两次cost+snap+fix|   3.23\n做wifi时间戳级别的cost min，插值得到wp，<br>snap，修复开始结束点，嵌套三次cost+snap+fix |   3.19\n\n# fake test\n第一次做Nfold的时候没有提升，我意识到训练集和测试集的分布不一致，我进行了一系列的探测，发现测试集有一些特点应该是人为限定的。测试集的txt大小>=2M（linux下看），时间跨度>=60s，路径点个数>=5。在前半个月的时候，这个限制同样抽出了626个路径作为fake test，无论是否进行后处理，我的fake test得分和线上的得分都保持一致。在最后1周我使用4折交叉验证，作为valid的那一份数据我同样做了上述限制，使得我的cv更加一致。后面我检查了private的分数情况，也符合相同的趋势。\n\n# 失败的尝试\n1.直接使用png图像特征作为端到端的输入\n2.使用pytorch复现网络结构和lb分数\n3.使用MLP替换并不是时间序列的lstm\n4.使用所有site的数据\n5.尝试序列RNN，但是它效果不好且难以收敛\n。。。。。。\n\n# 我有一个问题\n我这一个月始终都没有想明白，为什么序列长度为1的lstm模型工作的这么好，mlp为什么无法替代它。。。\n如果您对此发表任何看法。我将非常感激。",
    "1313706": "Thanks for sharing. I also learned a lot from this competition.\n> I haven't figured out why the lstm model with a sequence length of 1 works so well, and why can't mlp replace it. . .\n\nI am new to neural networks and have also been troubled with this question for a long time. My understanding is :\n1)A single-cell LSTM has 4.x times params compared to a Dense layer with the same units.  \n2)LSTM has a different mathematical equation that involves sigmoid and tanh (or relu) activations and element-wise product. And this structure might perform better than a single dense layer with only one activation.\nActually, I find my modified MLP model performs slightly better than LSTM.",
    "1314105": "> I find my modified MLP model performs slightly better than LSTM.\n\ngreat，make sense， Could you please share your network structure？"
  },
  "source": "meta"
}