{
  "id": 220835,
  "title": "CV score vs LB score",
  "url": "/competitions/indoor-location-navigation/discussion/220835",
  "author_name": "Jiwei Liu",
  "post_date": "2021-02-19T18:33:38.639000",
  "votes": 28,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I start to see inconsistent CV-LB patterns. This is expected due to the small data size. I've built RNN models recently:</p>\n<p>Model 0<br>\nCV: 8.2 LB:8.3</p>\n<p>Model 1 <br>\nCV: 7.8 LB: 7.6</p>\n<p>Model 2<br>\nCV: 7.4 LB: 8.0</p>\n<p>Model 3<br>\nCV: 7.67 LB: 6.99</p>\n<p>Model 4<br>\nCV: 7.37 LB: 7.49</p>\n<p>I have no proof but I feel public LB and private LB might be from different sites. Please correct me.</p>\n<p>My cross-validation strategy is kind of group kfold w.r.t paths and stratified kfold w.r.t sites. I build one model for all sites, which is different from the best kernel with one model per site. Please share your CV strategy and CV-LB. Thank you. </p>",
  "messages": [
    {
      "id": 1210850,
      "postDate": "2021-02-19T18:33:38.640Z",
      "content": "<p>I start to see inconsistent CV-LB patterns. This is expected due to the small data size. I've built RNN models recently:</p>\n<p>Model 0<br>\nCV: 8.2 LB:8.3</p>\n<p>Model 1 <br>\nCV: 7.8 LB: 7.6</p>\n<p>Model 2<br>\nCV: 7.4 LB: 8.0</p>\n<p>Model 3<br>\nCV: 7.67 LB: 6.99</p>\n<p>Model 4<br>\nCV: 7.37 LB: 7.49</p>\n<p>I have no proof but I feel public LB and private LB might be from different sites. Please correct me.</p>\n<p>My cross-validation strategy is kind of group kfold w.r.t paths and stratified kfold w.r.t sites. I build one model for all sites, which is different from the best kernel with one model per site. Please share your CV strategy and CV-LB. Thank you. </p>",
      "rawMarkdown": "I start to see inconsistent CV-LB patterns. This is expected due to the small data size. I've built RNN models recently:\n\nModel 0\nCV: 8.2 LB:8.3\n\nModel 1 \nCV: 7.8 LB: 7.6\n\nModel 2\nCV: 7.4 LB: 8.0\n\nModel 3\nCV: 7.67 LB: 6.99\n\nModel 4\nCV: 7.37 LB: 7.49\n\nI have no proof but I feel public LB and private LB might be from different sites. Please correct me.\n\nMy cross-validation strategy is kind of group kfold w.r.t paths and stratified kfold w.r.t sites. I build one model for all sites, which is different from the best kernel with one model per site. Please share your CV strategy and CV-LB. Thank you. ",
      "votes": 28
    },
    {
      "id": 1255520,
      "postDate": "2021-03-28T21:41:11.753Z",
      "content": "<p>CV 4.6, LB 3.2 and CV-LB is always correlated, as people say :)</p>",
      "rawMarkdown": "CV 4.6, LB 3.2 and CV-LB is always correlated, as people say :)",
      "votes": 10
    },
    {
      "id": 1210973,
      "postDate": "2021-02-19T20:58:53.050Z",
      "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>12.08</td>\n<td>9.75</td>\n</tr>\n<tr>\n<td>11.64</td>\n<td>9.22</td>\n</tr>\n<tr>\n<td>11.23</td>\n<td>9.27</td>\n</tr>\n<tr>\n<td>11.04</td>\n<td>8.69</td>\n</tr>\n</tbody>\n</table>\n<p>I agree with you. This competition might be shaky.</p>",
      "rawMarkdown": "| CV | LB |\n| --- | --- |\n| 12.08 | 9.75 |\n| 11.64 | 9.22 |\n| 11.23 | 9.27 |\n| 11.04 | 8.69 |\n\nI agree with you. This competition might be shaky.",
      "votes": 7
    },
    {
      "id": 1212445,
      "postDate": "2021-02-21T08:27:12.757Z",
      "content": "<p>Here are what I did to try to minimize CV-LB inconsistencies. </p>\n<ul>\n<li>make sure both train and test data are sorted by time stamp</li>\n<li>look at <code>df[col].describe()</code> to ensure train columns and test columns have the same distribution</li>\n<li>sites in test data have different frequencies from train. So I assigned weights to train sites to mitigate.</li>\n</ul>\n<p>The last point is not necessary if you build one model per site.</p>",
      "rawMarkdown": "Here are what I did to try to minimize CV-LB inconsistencies. \n- make sure both train and test data are sorted by time stamp\n- look at `df[col].describe()` to ensure train columns and test columns have the same distribution\n- sites in test data have different frequencies from train. So I assigned weights to train sites to mitigate.\n\nThe last point is not necessary if you build one model per site.",
      "votes": 5,
      "replies": [
        {
          "id": 1248500,
          "postDate": "2021-03-22T16:30:14.277Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>! I am trying to eliminate inconsistencies CV/LB and I see your post. What do you mean with  the first point? <br>\nSo far my strategy is GroupKfold w.r.t paths.</p>\n<p>Thanks in advance!</p>",
          "rawMarkdown": "Hi @jiweiliu! I am trying to eliminate inconsistencies CV/LB and I see your post. What do you mean with  the first point? \nSo far my strategy is GroupKfold w.r.t paths.\n\nThanks in advance!\n"
        },
        {
          "id": 1254285,
          "postDate": "2021-03-27T13:45:01.823Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1210962,
      "postDate": "2021-02-19T20:45:09.403Z",
      "content": "<p>Hi Jiwei,<br>\nI haven't computed a global CV score yet since I'm using site-specific models. However, it's not difficult to compute this global score and I'll do that.<br>\nIn my models, CV scores vary a lot between the different sites, one reason being that location data points have different densities between sites. So far, improvements on CV and LB scores were pretty consistent.</p>\n<p>Regarding the split between public and private LB, I don't have any hypothesis yet. I think it would be easy to do some probing just by changing floor value by 1 unit for some selected paths, and check the impact on the score :-). <br>\nOne thing to notice is that the number of timestamps by site in the test set is very \"unbalanced\", some sites being over-represented, as illustrated below:</p>\n<pre><code>                          # of timestamps\nsite     \n5a0546857ecc773753327266     299\n5c3c44b80379370013e0fd2b     26\n5d27075f03f801723c2e360f     47\n5d27096c03f801723c31e5e0     654\n5d27097f03f801723c320d97     303\n5d27099f03f801723c32511d     49\n5d2709a003f801723c3251bf     218\n5d2709b303f801723c327472     527\n5d2709bb03f801723c32852c     716\n5d2709c303f801723c3299ee     509\n5d2709d403f801723c32bd39     1223\n5d2709e003f801723c32d896     531\n5da138274db8ce0c98bbd3d2     103\n5da1382d4db8ce0c98bbe92e     311\n5da138314db8ce0c98bbf3a0     171\n5da138364db8ce0c98bc00f1     139\n5da1383b4db8ce0c98bc11ab     380\n5da138754db8ce0c98bca82f     386\n5da138764db8ce0c98bcaa46     573\n5da1389e4db8ce0c98bd0547     174\n5da138b74db8ce0c98bd4774     445\n5da958dd46f8266d0737457b     778\n5dbc1d84c1eb61796cf7c010     923\n5dc8cea7659e181adb076a3f     648\n</code></pre>",
      "rawMarkdown": "Hi Jiwei,\nI haven't computed a global CV score yet since I'm using site-specific models. However, it's not difficult to compute this global score and I'll do that.\nIn my models, CV scores vary a lot between the different sites, one reason being that location data points have different densities between sites. So far, improvements on CV and LB scores were pretty consistent.\n\nRegarding the split between public and private LB, I don't have any hypothesis yet. I think it would be easy to do some probing just by changing floor value by 1 unit for some selected paths, and check the impact on the score :-). \nOne thing to notice is that the number of timestamps by site in the test set is very \"unbalanced\", some sites being over-represented, as illustrated below:\n \n ```\n                          # of timestamps\nsite \t\n5a0546857ecc773753327266 \t299\n5c3c44b80379370013e0fd2b \t26\n5d27075f03f801723c2e360f \t47\n5d27096c03f801723c31e5e0 \t654\n5d27097f03f801723c320d97 \t303\n5d27099f03f801723c32511d \t49\n5d2709a003f801723c3251bf \t218\n5d2709b303f801723c327472 \t527\n5d2709bb03f801723c32852c \t716\n5d2709c303f801723c3299ee \t509\n5d2709d403f801723c32bd39 \t1223\n5d2709e003f801723c32d896 \t531\n5da138274db8ce0c98bbd3d2 \t103\n5da1382d4db8ce0c98bbe92e \t311\n5da138314db8ce0c98bbf3a0 \t171\n5da138364db8ce0c98bc00f1 \t139\n5da1383b4db8ce0c98bc11ab \t380\n5da138754db8ce0c98bca82f \t386\n5da138764db8ce0c98bcaa46 \t573\n5da1389e4db8ce0c98bd0547 \t174\n5da138b74db8ce0c98bd4774 \t445\n5da958dd46f8266d0737457b \t778\n5dbc1d84c1eb61796cf7c010 \t923\n5dc8cea7659e181adb076a3f \t648\n```",
      "votes": 6,
      "replies": [
        {
          "id": 1210970,
          "postDate": "2021-02-19T20:52:20.933Z",
          "content": "<p>Thank you for sharing! Yes, imbalance of sites is critical, especially for the one model approach I'm doing. You got amazing scores. Well done! </p>",
          "rawMarkdown": "Thank you for sharing! Yes, imbalance of sites is critical, especially for the one model approach I'm doing. You got amazing scores. Well done! ",
          "votes": 4
        },
        {
          "id": 1210989,
          "postDate": "2021-02-19T21:18:12.003Z",
          "content": "<p>Thanks. Your recent progression on the LB has been also very impressive.<br>\nMy feeling is that there is still room for a lot of progress. I'm not familiar with the indoor-location domain and I would not make any bet yet about best scores at the end of this competition, but I can notice in my submissions that location predictions are still very noisy (with too many unrealistic locations). For this reason, I'm afraid it is too early to draw reliable statistical patterns.<br>\nFor now, my goal is still to focus on getting a better understanding of the domain and the data to improve my approach. Good luck !</p>",
          "rawMarkdown": "Thanks. Your recent progression on the LB has been also very impressive.\nMy feeling is that there is still room for a lot of progress. I'm not familiar with the indoor-location domain and I would not make any bet yet about best scores at the end of this competition, but I can notice in my submissions that location predictions are still very noisy (with too many unrealistic locations). For this reason, I'm afraid it is too early to draw reliable statistical patterns.\nFor now, my goal is still to focus on getting a better understanding of the domain and the data to improve my approach. Good luck !",
          "votes": 5
        }
      ]
    },
    {
      "id": 1248762,
      "postDate": "2021-03-22T20:42:44.300Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> <br>\nI have just joined and I am new here , it would be great help if you could shed some light on the kind of CV you are using.<br>\nWhat I have thought of right now is for a particular txt file we take the first few timestamps in train and the last few in valid , is it the right way?</p>",
      "rawMarkdown": "Hi @jiweiliu \nI have just joined and I am new here , it would be great help if you could shed some light on the kind of CV you are using.\nWhat I have thought of right now is for a particular txt file we take the first few timestamps in train and the last few in valid , is it the right way?",
      "votes": 3
    },
    {
      "id": 1213621,
      "postDate": "2021-02-22T08:31:21.417Z",
      "content": "<p>for my lgb model( 70k+ training records, not the public baseline):</p>\n<ul>\n<li>CV 11.2 ,  LB 9.1</li>\n<li>CV 10.8 , LB 8.5</li>\n<li>CV 8.41 , LB 7.48</li>\n<li>CV 8.38 , LB 7.43</li>\n</ul>\n<p>LSTM model：</p>\n<ul>\n<li>CV 10.67,  LB 7.38, the gap is worse than lgb</li>\n</ul>",
      "rawMarkdown": "for my lgb model( 70k+ training records, not the public baseline):\n- CV 11.2 ,  LB 9.1\n- CV 10.8 , LB 8.5\n- CV 8.41 , LB 7.48\n- CV 8.38 , LB 7.43\n\nLSTM model：\n- CV 10.67,  LB 7.38, the gap is worse than lgb",
      "votes": 3,
      "replies": [
        {
          "id": 1215399,
          "postDate": "2021-02-23T15:21:44.513Z",
          "content": "<p>Thanks, looks highly correlated and similar to mine</p>",
          "rawMarkdown": "Thanks, looks highly correlated and similar to mine",
          "votes": 1
        }
      ]
    },
    {
      "id": 1217481,
      "postDate": "2021-02-25T05:41:28.307Z",
      "content": "<p>My LSTM LB  = 8.689<br>\nI use one model per site and CV range between [6.xxx; 11.xxx] depending on site</p>\n<p>I may try Transformers based model latter to see how it can helps</p>",
      "rawMarkdown": "My LSTM LB  = 8.689\nI use one model per site and CV range between [6.xxx; 11.xxx] depending on site\n\nI may try Transformers based model latter to see how it can helps",
      "votes": 1
    },
    {
      "id": 1248614,
      "postDate": "2021-03-22T17:50:45Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1248613,
      "postDate": "2021-03-22T17:49:39.423Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1255520,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2021-03-28T21:41:11.753000",
      "content": "<p>CV 4.6, LB 3.2 and CV-LB is always correlated, as people say :)</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1210973,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2021-02-19T20:58:53.050000",
      "content": "<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>12.08</td>\n<td>9.75</td>\n</tr>\n<tr>\n<td>11.64</td>\n<td>9.22</td>\n</tr>\n<tr>\n<td>11.23</td>\n<td>9.27</td>\n</tr>\n<tr>\n<td>11.04</td>\n<td>8.69</td>\n</tr>\n</tbody>\n</table>\n<p>I agree with you. This competition might be shaky.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1212445,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2021-02-21T08:27:12.757000",
      "content": "<p>Here are what I did to try to minimize CV-LB inconsistencies. </p>\n<ul>\n<li>make sure both train and test data are sorted by time stamp</li>\n<li>look at <code>df[col].describe()</code> to ensure train columns and test columns have the same distribution</li>\n<li>sites in test data have different frequencies from train. So I assigned weights to train sites to mitigate.</li>\n</ul>\n<p>The last point is not necessary if you build one model per site.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1248500,
          "author_name": "Fnoa",
          "author_url": "",
          "post_date": "2021-03-22T16:30:14.277000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>! I am trying to eliminate inconsistencies CV/LB and I see your post. What do you mean with  the first point? <br>\nSo far my strategy is GroupKfold w.r.t paths.</p>\n<p>Thanks in advance!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254285,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-27T13:45:01.823000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1210962,
      "author_name": "Hugues",
      "author_url": "",
      "post_date": "2021-02-19T20:45:09.403000",
      "content": "<p>Hi Jiwei,<br>\nI haven't computed a global CV score yet since I'm using site-specific models. However, it's not difficult to compute this global score and I'll do that.<br>\nIn my models, CV scores vary a lot between the different sites, one reason being that location data points have different densities between sites. So far, improvements on CV and LB scores were pretty consistent.</p>\n<p>Regarding the split between public and private LB, I don't have any hypothesis yet. I think it would be easy to do some probing just by changing floor value by 1 unit for some selected paths, and check the impact on the score :-). <br>\nOne thing to notice is that the number of timestamps by site in the test set is very \"unbalanced\", some sites being over-represented, as illustrated below:</p>\n<pre><code>                          # of timestamps\nsite     \n5a0546857ecc773753327266     299\n5c3c44b80379370013e0fd2b     26\n5d27075f03f801723c2e360f     47\n5d27096c03f801723c31e5e0     654\n5d27097f03f801723c320d97     303\n5d27099f03f801723c32511d     49\n5d2709a003f801723c3251bf     218\n5d2709b303f801723c327472     527\n5d2709bb03f801723c32852c     716\n5d2709c303f801723c3299ee     509\n5d2709d403f801723c32bd39     1223\n5d2709e003f801723c32d896     531\n5da138274db8ce0c98bbd3d2     103\n5da1382d4db8ce0c98bbe92e     311\n5da138314db8ce0c98bbf3a0     171\n5da138364db8ce0c98bc00f1     139\n5da1383b4db8ce0c98bc11ab     380\n5da138754db8ce0c98bca82f     386\n5da138764db8ce0c98bcaa46     573\n5da1389e4db8ce0c98bd0547     174\n5da138b74db8ce0c98bd4774     445\n5da958dd46f8266d0737457b     778\n5dbc1d84c1eb61796cf7c010     923\n5dc8cea7659e181adb076a3f     648\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 1210970,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2021-02-19T20:52:20.933000",
          "content": "<p>Thank you for sharing! Yes, imbalance of sites is critical, especially for the one model approach I'm doing. You got amazing scores. Well done! </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1210989,
          "author_name": "Hugues",
          "author_url": "",
          "post_date": "2021-02-19T21:18:12.003000",
          "content": "<p>Thanks. Your recent progression on the LB has been also very impressive.<br>\nMy feeling is that there is still room for a lot of progress. I'm not familiar with the indoor-location domain and I would not make any bet yet about best scores at the end of this competition, but I can notice in my submissions that location predictions are still very noisy (with too many unrealistic locations). For this reason, I'm afraid it is too early to draw reliable statistical patterns.<br>\nFor now, my goal is still to focus on getting a better understanding of the domain and the data to improve my approach. Good luck !</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1248762,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-03-22T20:42:44.300000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> <br>\nI have just joined and I am new here , it would be great help if you could shed some light on the kind of CV you are using.<br>\nWhat I have thought of right now is for a particular txt file we take the first few timestamps in train and the last few in valid , is it the right way?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1213621,
      "author_name": "qyxs",
      "author_url": "",
      "post_date": "2021-02-22T08:31:21.417000",
      "content": "<p>for my lgb model( 70k+ training records, not the public baseline):</p>\n<ul>\n<li>CV 11.2 ,  LB 9.1</li>\n<li>CV 10.8 , LB 8.5</li>\n<li>CV 8.41 , LB 7.48</li>\n<li>CV 8.38 , LB 7.43</li>\n</ul>\n<p>LSTM model：</p>\n<ul>\n<li>CV 10.67,  LB 7.38, the gap is worse than lgb</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 1215399,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2021-02-23T15:21:44.513000",
          "content": "<p>Thanks, looks highly correlated and similar to mine</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1217481,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-02-25T05:41:28.307000",
      "content": "<p>My LSTM LB  = 8.689<br>\nI use one model per site and CV range between [6.xxx; 11.xxx] depending on site</p>\n<p>I may try Transformers based model latter to see how it can helps</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1248614,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-22T17:50:45",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1248613,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-22T17:49:39.423000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210850": "I start to see inconsistent CV-LB patterns. This is expected due to the small data size. I've built RNN models recently:\n\nModel 0\nCV: 8.2 LB:8.3\n\nModel 1 \nCV: 7.8 LB: 7.6\n\nModel 2\nCV: 7.4 LB: 8.0\n\nModel 3\nCV: 7.67 LB: 6.99\n\nModel 4\nCV: 7.37 LB: 7.49\n\nI have no proof but I feel public LB and private LB might be from different sites. Please correct me.\n\nMy cross-validation strategy is kind of group kfold w.r.t paths and stratified kfold w.r.t sites. I build one model for all sites, which is different from the best kernel with one model per site. Please share your CV strategy and CV-LB. Thank you. ",
    "1255520": "CV 4.6, LB 3.2 and CV-LB is always correlated, as people say :)",
    "1210973": "| CV | LB |\n| --- | --- |\n| 12.08 | 9.75 |\n| 11.64 | 9.22 |\n| 11.23 | 9.27 |\n| 11.04 | 8.69 |\n\nI agree with you. This competition might be shaky.",
    "1212445": "Here are what I did to try to minimize CV-LB inconsistencies. \n- make sure both train and test data are sorted by time stamp\n- look at `df[col].describe()` to ensure train columns and test columns have the same distribution\n- sites in test data have different frequencies from train. So I assigned weights to train sites to mitigate.\n\nThe last point is not necessary if you build one model per site.",
    "1210962": "Hi Jiwei,\nI haven't computed a global CV score yet since I'm using site-specific models. However, it's not difficult to compute this global score and I'll do that.\nIn my models, CV scores vary a lot between the different sites, one reason being that location data points have different densities between sites. So far, improvements on CV and LB scores were pretty consistent.\n\nRegarding the split between public and private LB, I don't have any hypothesis yet. I think it would be easy to do some probing just by changing floor value by 1 unit for some selected paths, and check the impact on the score :-). \nOne thing to notice is that the number of timestamps by site in the test set is very \"unbalanced\", some sites being over-represented, as illustrated below:\n \n ```\n                          # of timestamps\nsite \t\n5a0546857ecc773753327266 \t299\n5c3c44b80379370013e0fd2b \t26\n5d27075f03f801723c2e360f \t47\n5d27096c03f801723c31e5e0 \t654\n5d27097f03f801723c320d97 \t303\n5d27099f03f801723c32511d \t49\n5d2709a003f801723c3251bf \t218\n5d2709b303f801723c327472 \t527\n5d2709bb03f801723c32852c \t716\n5d2709c303f801723c3299ee \t509\n5d2709d403f801723c32bd39 \t1223\n5d2709e003f801723c32d896 \t531\n5da138274db8ce0c98bbd3d2 \t103\n5da1382d4db8ce0c98bbe92e \t311\n5da138314db8ce0c98bbf3a0 \t171\n5da138364db8ce0c98bc00f1 \t139\n5da1383b4db8ce0c98bc11ab \t380\n5da138754db8ce0c98bca82f \t386\n5da138764db8ce0c98bcaa46 \t573\n5da1389e4db8ce0c98bd0547 \t174\n5da138b74db8ce0c98bd4774 \t445\n5da958dd46f8266d0737457b \t778\n5dbc1d84c1eb61796cf7c010 \t923\n5dc8cea7659e181adb076a3f \t648\n```",
    "1248762": "Hi @jiweiliu \nI have just joined and I am new here , it would be great help if you could shed some light on the kind of CV you are using.\nWhat I have thought of right now is for a particular txt file we take the first few timestamps in train and the last few in valid , is it the right way?",
    "1213621": "for my lgb model( 70k+ training records, not the public baseline):\n- CV 11.2 ,  LB 9.1\n- CV 10.8 , LB 8.5\n- CV 8.41 , LB 7.48\n- CV 8.38 , LB 7.43\n\nLSTM model：\n- CV 10.67,  LB 7.38, the gap is worse than lgb",
    "1217481": "My LSTM LB  = 8.689\nI use one model per site and CV range between [6.xxx; 11.xxx] depending on site\n\nI may try Transformers based model latter to see how it can helps",
    "1248614": "",
    "1248613": ""
  }
}