{
  "id": 78138,
  "title": "Differences between test and train set?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/78138",
  "author_name": "",
  "post_date": "2019-01-20T09:15:10.134857300Z",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Just wondering if anyone else is seeing a consistent difference in performance of a model on the train set compared to when submitting results.</p>\n\n<p>I have split off 1/8 of the training samples (split on the <code>id_measurement</code> column) as a validation set and I am confident that I have not made any errors in my code that means this set is seen during training. I consistently score around 0.6 --&gt; 0.7 with this validation set but my best score on submission is very much weaker (c. 0.4)</p>\n\n<p>I understand that I have only a small local validation set and there are lots of avenues for me still to explore but it seemed like a large gulf in performance so I thought I'd ask if anyone else had noticed a similar pattern?</p>\n\n<p>Do we have any information from the organisers as to how the data was gathered? Is it simply random chance that a given measurement ends up in the test set instead of the train set?</p>",
  "messages": [
    {
      "id": "458683",
      "postDate": "01/20/2019 09:15:10",
      "content": "<p>Just wondering if anyone else is seeing a consistent difference in performance of a model on the train set compared to when submitting results.</p>\n\n<p>I have split off 1/8 of the training samples (split on the <code>id_measurement</code> column) as a validation set and I am confident that I have not made any errors in my code that means this set is seen during training. I consistently score around 0.6 --&gt; 0.7 with this validation set but my best score on submission is very much weaker (c. 0.4)</p>\n\n<p>I understand that I have only a small local validation set and there are lots of avenues for me still to explore but it seemed like a large gulf in performance so I thought I'd ask if anyone else had noticed a similar pattern?</p>\n\n<p>Do we have any information from the organisers as to how the data was gathered? Is it simply random chance that a given measurement ends up in the test set instead of the train set?</p>",
      "rawMarkdown": "Just wondering if anyone else is seeing a consistent difference in performance of a model on the train set compared to when submitting results.\n\nI have split off 1/8 of the training samples (split on the `id_measurement` column) as a validation set and I am confident that I have not made any errors in my code that means this set is seen during training. I consistently score around 0.6 --&gt; 0.7 with this validation set but my best score on submission is very much weaker (c. 0.4)\n\nI understand that I have only a small local validation set and there are lots of avenues for me still to explore but it seemed like a large gulf in performance so I thought I'd ask if anyone else had noticed a similar pattern?\n\nDo we have any information from the organisers as to how the data was gathered? Is it simply random chance that a given measurement ends up in the test set instead of the train set?",
      "votes": null
    },
    {
      "id": "459105",
      "postDate": "01/21/2019 07:31:47",
      "content": "<p>There is a similar thread:\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984</a></p>",
      "rawMarkdown": "There is a similar thread:\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984",
      "votes": null
    },
    {
      "id": "459416",
      "postDate": "01/21/2019 17:24:25",
      "content": "<p>I'm presently right there with you. </p>\n\n<p>My approach to tackling this problem is to take the raw signal, perform some signal processing, and then extract what I'm perceiving to be useful features into a data frame that can be used to train a Random Forest machine learning model.</p>\n\n<p>I've been running Monte Carlo trials varying the random_state of my <code>train_test_split</code> of my extracted signal features and my model locally is producing Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.</p>\n\n<p>My final step is to use the extracted features from the entire training set to train the model and then produce predictions on the test set and create a file in the submission format. My first submission netted a 0.34 on the leader board.</p>\n\n<p>I understand this is a difficult problem to solve given the heavy imbalance of good and fault labeled data but I thought a Random Forest model would offer some resilience to overfitting to the dominant class.</p>\n\n<p>I too am at a loss on how to improve further with a Random Forest classifier since I already believe I'm reaching the high end locally. From what I'm seeing int he discussion forums, it seems most of the successful teams on the leader board are using k-Fold Cross Validation or Recurrent Neural Nets like LSTM.</p>",
      "rawMarkdown": "I'm presently right there with you. \n\nMy approach to tackling this problem is to take the raw signal, perform some signal processing, and then extract what I'm perceiving to be useful features into a data frame that can be used to train a Random Forest machine learning model.\n\nI've been running Monte Carlo trials varying the random_state of my `train_test_split` of my extracted signal features and my model locally is producing Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.\n\nMy final step is to use the extracted features from the entire training set to train the model and then produce predictions on the test set and create a file in the submission format. My first submission netted a 0.34 on the leader board.\n\nI understand this is a difficult problem to solve given the heavy imbalance of good and fault labeled data but I thought a Random Forest model would offer some resilience to overfitting to the dominant class.\n\nI too am at a loss on how to improve further with a Random Forest classifier since I already believe I'm reaching the high end locally. From what I'm seeing int he discussion forums, it seems most of the successful teams on the leader board are using k-Fold Cross Validation or Recurrent Neural Nets like LSTM.",
      "votes": null
    },
    {
      "id": "459450",
      "postDate": "01/21/2019 18:58:50",
      "content": "<blockquote>\n  <p>Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.</p>\n</blockquote>\n\n<p>This is suspicious! My CV is 0.72 and 0.63 LB (using LSTM). </p>\n\n<p>Before that I tried LGBM with simple features, both CV and LB was low, but the gap was wider.</p>",
      "rawMarkdown": "&gt; Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.\n\nThis is suspicious! My CV is 0.72 and 0.63 LB (using LSTM). \n\nBefore that I tried LGBM with simple features, both CV and LB was low, but the gap was wider.",
      "votes": null
    },
    {
      "id": "459696",
      "postDate": "01/22/2019 07:32:32",
      "content": "<p>Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case. Best Lightgbm model for me has CV = 0.697872, LB = 0.555. \nHope, You are not using id_measurement as a feature also.</p>",
      "rawMarkdown": "Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case. Best Lightgbm model for me has CV = 0.697872, LB = 0.555. \nHope, You are not using id_measurement as a feature also.",
      "votes": null
    },
    {
      "id": "461594",
      "postDate": "01/26/2019 13:12:34",
      "content": "<blockquote>\n  <p>Hope, You are not using <code>id_measurement</code> as a feature also.</p>\n</blockquote>\n\n<p>You had me scared that I had made a mistake there for a minute because that would have made perfect sense to get such high MCC scores with that present. I double checked that I am not inadvertently passing this feature to my model for training and test.</p>\n\n<p>My current approach was to essentially follow the outline laid out in the academic paper published by Tomas Vantuch. I'm doing my own signal processing to de-noise, some basic pulse cancellation, and extract features that I can be used to fit the model. Namely many of the features he found useful in the paper: entropy, peak characteristics, etc are the same ones I'm using. My latest theory is that some of these characteristics between signals correlate just as well as <code>id_measurement</code> and help allow for an unusually high MCC.</p>\n\n<blockquote>\n  <p>Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case.</p>\n</blockquote>\n\n<p>I'm going to spend some time this weekend following up on your suggestion to re-evaluate the distribution of features in each set and re-run models with features that have similar distributions in both sets. In all likelihood I'll probably have to circle back to my signal processing code and make some tweaks in data prep. I attached a pair plot of my current feature distribution (along with \"fault\") if you were curious where they stood at the moment.</p>\n\n<p>Thanks for the feedback and suggestions!</p>\n\n<p><img src=\"https://i.imgur.com/LLVIFge.png\" alt=\"Image\"> \n<a href=\"https://i.imgur.com/LLVIFge.png\">Larger version</a></p>",
      "rawMarkdown": "&gt; Hope, You are not using `id_measurement` as a feature also.\n\nYou had me scared that I had made a mistake there for a minute because that would have made perfect sense to get such high MCC scores with that present. I double checked that I am not inadvertently passing this feature to my model for training and test.\n\nMy current approach was to essentially follow the outline laid out in the academic paper published by Tomas Vantuch. I'm doing my own signal processing to de-noise, some basic pulse cancellation, and extract features that I can be used to fit the model. Namely many of the features he found useful in the paper: entropy, peak characteristics, etc are the same ones I'm using. My latest theory is that some of these characteristics between signals correlate just as well as `id_measurement` and help allow for an unusually high MCC.\n\n&gt; Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case.\n\nI'm going to spend some time this weekend following up on your suggestion to re-evaluate the distribution of features in each set and re-run models with features that have similar distributions in both sets. In all likelihood I'll probably have to circle back to my signal processing code and make some tweaks in data prep. I attached a pair plot of my current feature distribution (along with \"fault\") if you were curious where they stood at the moment.\n\nThanks for the feedback and suggestions!\n\n![Image](https://i.imgur.com/LLVIFge.png) \n[Larger version](https://i.imgur.com/LLVIFge.png)",
      "votes": null
    },
    {
      "id": "461683",
      "postDate": "01/26/2019 18:08:37",
      "content": "<p>Great! Let me know how the boosting is working for you. I wasn’t able to get good score from boosting, hence moved to lstm, but those models are unstable so far, fixing seed can reproduce the results but changing the seed can lead to a lot of variation.</p>",
      "rawMarkdown": "Great! Let me know how the boosting is working for you. I wasn’t able to get good score from boosting, hence moved to lstm, but those models are unstable so far, fixing seed can reproduce the results but changing the seed can lead to a lot of variation.",
      "votes": null
    },
    {
      "id": "462120",
      "postDate": "01/27/2019 18:09:15",
      "content": "<p>A bit of a mixed bag. The \"good\" news is that now my models are scoring between 0.26-0.43 locally which is consistent with my leader board scoring around 0.34. The bad news is that I'm also starting to get the sense that this is the high end of possible performance for Random Forest. I also tried k-Nearest Neighbors and SVM with worse performance.</p>\n\n<p>I'm going to follow your lead and start exploring use of some models I haven't really used before like the LTSM. Looking forward to learning something new!</p>",
      "rawMarkdown": "A bit of a mixed bag. The \"good\" news is that now my models are scoring between 0.26-0.43 locally which is consistent with my leader board scoring around 0.34. The bad news is that I'm also starting to get the sense that this is the high end of possible performance for Random Forest. I also tried k-Nearest Neighbors and SVM with worse performance.\n\nI'm going to follow your lead and start exploring use of some models I haven't really used before like the LTSM. Looking forward to learning something new!",
      "votes": null
    },
    {
      "id": "464523",
      "postDate": "02/01/2019 03:07:26",
      "content": "<p>Hi Jeffrey,</p>\n\n<p>Did you try using those features in a LSTM ? I tried using similar features but same thing I got very different results as well between validation set and LB. </p>",
      "rawMarkdown": "Hi Jeffrey,\n\nDid you try using those features in a LSTM ? I tried using similar features but same thing I got very different results as well between validation set and LB.",
      "votes": null
    },
    {
      "id": "464524",
      "postDate": "02/01/2019 03:09:12",
      "content": "<p>Hi <a href=\"/harshit92\">@harshit92</a>, were you able to fix the seed to make it fully reproductible with Keras ? It seems that the GPU computation is not deterministic. I always have very different results between runs.</p>",
      "rawMarkdown": "Hi @harshit92, were you able to fix the seed to make it fully reproductible with Keras ? It seems that the GPU computation is not deterministic. I always have very different results between runs.",
      "votes": null
    },
    {
      "id": "464529",
      "postDate": "02/01/2019 03:16:13",
      "content": "<p>@Antoine Reveillion, I have only checked on my laptop, i will check on kaggle GPU with same code and will let you know. </p>",
      "rawMarkdown": "Antoine Reveillion, I have only checked on my laptop, i will check on kaggle GPU with same code and will let you know.",
      "votes": null
    },
    {
      "id": "464829",
      "postDate": "02/01/2019 15:35:22",
      "content": "<p>That's probably why then ! Thanks !</p>",
      "rawMarkdown": "That's probably why then ! Thanks !",
      "votes": null
    },
    {
      "id": "464855",
      "postDate": "02/01/2019 16:46:05",
      "content": "<p>@Antoine, I tried it on kaggle platform and getting differences in the results for same seed and everything else. Tried it three times and cv was 0.684,0.699 and 0.7097.</p>",
      "rawMarkdown": "Antoine, I tried it on kaggle platform and getting differences in the results for same seed and everything else. Tried it three times and cv was 0.684,0.699 and 0.7097.",
      "votes": null
    },
    {
      "id": "464962",
      "postDate": "02/01/2019 22:33:36",
      "content": "<p>Yes that exactly what I have, and the LB is moving the same way as well...</p>",
      "rawMarkdown": "Yes that exactly what I have, and the LB is moving the same way as well...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 459105,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "01/21/2019 07:31:47",
      "content": "<p>There is a similar thread:\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 459416,
      "author_name": "jeffreyegan",
      "author_url": "",
      "post_date": "01/21/2019 17:24:25",
      "content": "<p>I'm presently right there with you. </p>\n\n<p>My approach to tackling this problem is to take the raw signal, perform some signal processing, and then extract what I'm perceiving to be useful features into a data frame that can be used to train a Random Forest machine learning model.</p>\n\n<p>I've been running Monte Carlo trials varying the random_state of my <code>train_test_split</code> of my extracted signal features and my model locally is producing Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.</p>\n\n<p>My final step is to use the extracted features from the entire training set to train the model and then produce predictions on the test set and create a file in the submission format. My first submission netted a 0.34 on the leader board.</p>\n\n<p>I understand this is a difficult problem to solve given the heavy imbalance of good and fault labeled data but I thought a Random Forest model would offer some resilience to overfitting to the dominant class.</p>\n\n<p>I too am at a loss on how to improve further with a Random Forest classifier since I already believe I'm reaching the high end locally. From what I'm seeing int he discussion forums, it seems most of the successful teams on the leader board are using k-Fold Cross Validation or Recurrent Neural Nets like LSTM.</p>",
      "votes": null,
      "replies": [
        {
          "id": 459450,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "01/21/2019 18:58:50",
          "content": "<blockquote>\n  <p>Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.</p>\n</blockquote>\n\n<p>This is suspicious! My CV is 0.72 and 0.63 LB (using LSTM). </p>\n\n<p>Before that I tried LGBM with simple features, both CV and LB was low, but the gap was wider.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 459696,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "01/22/2019 07:32:32",
          "content": "<p>Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case. Best Lightgbm model for me has CV = 0.697872, LB = 0.555. \nHope, You are not using id_measurement as a feature also.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461594,
          "author_name": "jeffreyegan",
          "author_url": "",
          "post_date": "01/26/2019 13:12:34",
          "content": "<blockquote>\n  <p>Hope, You are not using <code>id_measurement</code> as a feature also.</p>\n</blockquote>\n\n<p>You had me scared that I had made a mistake there for a minute because that would have made perfect sense to get such high MCC scores with that present. I double checked that I am not inadvertently passing this feature to my model for training and test.</p>\n\n<p>My current approach was to essentially follow the outline laid out in the academic paper published by Tomas Vantuch. I'm doing my own signal processing to de-noise, some basic pulse cancellation, and extract features that I can be used to fit the model. Namely many of the features he found useful in the paper: entropy, peak characteristics, etc are the same ones I'm using. My latest theory is that some of these characteristics between signals correlate just as well as <code>id_measurement</code> and help allow for an unusually high MCC.</p>\n\n<blockquote>\n  <p>Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case.</p>\n</blockquote>\n\n<p>I'm going to spend some time this weekend following up on your suggestion to re-evaluate the distribution of features in each set and re-run models with features that have similar distributions in both sets. In all likelihood I'll probably have to circle back to my signal processing code and make some tweaks in data prep. I attached a pair plot of my current feature distribution (along with \"fault\") if you were curious where they stood at the moment.</p>\n\n<p>Thanks for the feedback and suggestions!</p>\n\n<p><img src=\"https://i.imgur.com/LLVIFge.png\" alt=\"Image\"> \n<a href=\"https://i.imgur.com/LLVIFge.png\">Larger version</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 461683,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "01/26/2019 18:08:37",
          "content": "<p>Great! Let me know how the boosting is working for you. I wasn’t able to get good score from boosting, hence moved to lstm, but those models are unstable so far, fixing seed can reproduce the results but changing the seed can lead to a lot of variation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 462120,
          "author_name": "jeffreyegan",
          "author_url": "",
          "post_date": "01/27/2019 18:09:15",
          "content": "<p>A bit of a mixed bag. The \"good\" news is that now my models are scoring between 0.26-0.43 locally which is consistent with my leader board scoring around 0.34. The bad news is that I'm also starting to get the sense that this is the high end of possible performance for Random Forest. I also tried k-Nearest Neighbors and SVM with worse performance.</p>\n\n<p>I'm going to follow your lead and start exploring use of some models I haven't really used before like the LTSM. Looking forward to learning something new!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464523,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "02/01/2019 03:07:26",
          "content": "<p>Hi Jeffrey,</p>\n\n<p>Did you try using those features in a LSTM ? I tried using similar features but same thing I got very different results as well between validation set and LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464524,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "02/01/2019 03:09:12",
          "content": "<p>Hi <a href=\"/harshit92\">@harshit92</a>, were you able to fix the seed to make it fully reproductible with Keras ? It seems that the GPU computation is not deterministic. I always have very different results between runs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464529,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/01/2019 03:16:13",
          "content": "<p>@Antoine Reveillion, I have only checked on my laptop, i will check on kaggle GPU with same code and will let you know. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464829,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "02/01/2019 15:35:22",
          "content": "<p>That's probably why then ! Thanks !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464855,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "02/01/2019 16:46:05",
          "content": "<p>@Antoine, I tried it on kaggle platform and getting differences in the results for same seed and everything else. Tried it three times and cv was 0.684,0.699 and 0.7097.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 464962,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "02/01/2019 22:33:36",
          "content": "<p>Yes that exactly what I have, and the LB is moving the same way as well...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "458683": "Just wondering if anyone else is seeing a consistent difference in performance of a model on the train set compared to when submitting results.\n\nI have split off 1/8 of the training samples (split on the `id_measurement` column) as a validation set and I am confident that I have not made any errors in my code that means this set is seen during training. I consistently score around 0.6 --&gt; 0.7 with this validation set but my best score on submission is very much weaker (c. 0.4)\n\nI understand that I have only a small local validation set and there are lots of avenues for me still to explore but it seemed like a large gulf in performance so I thought I'd ask if anyone else had noticed a similar pattern?\n\nDo we have any information from the organisers as to how the data was gathered? Is it simply random chance that a given measurement ends up in the test set instead of the train set?",
    "459105": "There is a similar thread:\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/76984",
    "459416": "I'm presently right there with you. \n\nMy approach to tackling this problem is to take the raw signal, perform some signal processing, and then extract what I'm perceiving to be useful features into a data frame that can be used to train a Random Forest machine learning model.\n\nI've been running Monte Carlo trials varying the random_state of my `train_test_split` of my extracted signal features and my model locally is producing Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.\n\nMy final step is to use the extracted features from the entire training set to train the model and then produce predictions on the test set and create a file in the submission format. My first submission netted a 0.34 on the leader board.\n\nI understand this is a difficult problem to solve given the heavy imbalance of good and fault labeled data but I thought a Random Forest model would offer some resilience to overfitting to the dominant class.\n\nI too am at a loss on how to improve further with a Random Forest classifier since I already believe I'm reaching the high end locally. From what I'm seeing int he discussion forums, it seems most of the successful teams on the leader board are using k-Fold Cross Validation or Recurrent Neural Nets like LSTM.",
    "459450": "&gt; Matthews Correlation Coefficient scores consistently within 0.87 - 0.91.\n\nThis is suspicious! My CV is 0.72 and 0.63 LB (using LSTM). \n\nBefore that I tried LGBM with simple features, both CV and LB was low, but the gap was wider.",
    "459696": "Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case. Best Lightgbm model for me has CV = 0.697872, LB = 0.555. \nHope, You are not using id_measurement as a feature also.",
    "461594": "&gt; Hope, You are not using `id_measurement` as a feature also.\n\nYou had me scared that I had made a mistake there for a minute because that would have made perfect sense to get such high MCC scores with that present. I double checked that I am not inadvertently passing this feature to my model for training and test.\n\nMy current approach was to essentially follow the outline laid out in the academic paper published by Tomas Vantuch. I'm doing my own signal processing to de-noise, some basic pulse cancellation, and extract features that I can be used to fit the model. Namely many of the features he found useful in the paper: entropy, peak characteristics, etc are the same ones I'm using. My latest theory is that some of these characteristics between signals correlate just as well as `id_measurement` and help allow for an unusually high MCC.\n\n&gt; Just make sure you use features which have pretty much the same distributions in train and test set. I Initially used Lightgbm and had that problem, though difference wasn't that high in my case.\n\nI'm going to spend some time this weekend following up on your suggestion to re-evaluate the distribution of features in each set and re-run models with features that have similar distributions in both sets. In all likelihood I'll probably have to circle back to my signal processing code and make some tweaks in data prep. I attached a pair plot of my current feature distribution (along with \"fault\") if you were curious where they stood at the moment.\n\nThanks for the feedback and suggestions!\n\n![Image](https://i.imgur.com/LLVIFge.png) \n[Larger version](https://i.imgur.com/LLVIFge.png)",
    "461683": "Great! Let me know how the boosting is working for you. I wasn’t able to get good score from boosting, hence moved to lstm, but those models are unstable so far, fixing seed can reproduce the results but changing the seed can lead to a lot of variation.",
    "462120": "A bit of a mixed bag. The \"good\" news is that now my models are scoring between 0.26-0.43 locally which is consistent with my leader board scoring around 0.34. The bad news is that I'm also starting to get the sense that this is the high end of possible performance for Random Forest. I also tried k-Nearest Neighbors and SVM with worse performance.\n\nI'm going to follow your lead and start exploring use of some models I haven't really used before like the LTSM. Looking forward to learning something new!",
    "464523": "Hi Jeffrey,\n\nDid you try using those features in a LSTM ? I tried using similar features but same thing I got very different results as well between validation set and LB.",
    "464524": "Hi @harshit92, were you able to fix the seed to make it fully reproductible with Keras ? It seems that the GPU computation is not deterministic. I always have very different results between runs.",
    "464529": "Antoine Reveillion, I have only checked on my laptop, i will check on kaggle GPU with same code and will let you know.",
    "464829": "That's probably why then ! Thanks !",
    "464855": "Antoine, I tried it on kaggle platform and getting differences in the results for same seed and everything else. Tried it three times and cv was 0.684,0.699 and 0.7097.",
    "464962": "Yes that exactly what I have, and the LB is moving the same way as well..."
  },
  "source": "meta"
}