{
  "id": 201480,
  "title": "what's your secret sauce? ",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/201480",
  "author_name": "Kamal Das",
  "post_date": "2020-12-05T08:16:53.596000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We have a wonderful notebook shared by <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> <br>\n<a href=\"https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh</a></p>\n<p>Alexander Lyubchenko's notebook is a great one… only 44 of us have managed to cross his score</p>\n<p>I was wondering what your secret sauce is to cross the score?<br>\nComing out of MoA comp, I tried blending/stacking a few models and that has helped me reach #30</p>\n<p>what about others? What is working and not working in this competition?</p>\n<p>Thanks for sharing!!</p>",
  "messages": [
    {
      "id": 1116625,
      "postDate": "2020-12-17T10:44:17.083Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> for mentioning my work in this discussion and its evaluation. Also, pls, I'm sorry for my late answer as I was on vacation.  <br>\nActually, guys, there is no secret in my kernel :) I generated a set of features using tsfresh and applied catboost with even default params, everything is posted in the notebook. <br>\nHonestly, yeah, I tried to tune catboost, checking symmetric and non-symmetric trees, applying the feature selection methods from sklearn. But all checked hypotheses didn't help to overcome the score obtained with the ingv-tsfresh-7730 dataset <a href=\"https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730</a>  and default catboostregressor. <br>\nConcerning NA imputing, I used simple fillna with zeros for time series with more than 50% of NAs.   </p>\n<p>Currently, I have several hypotheses, one of them is to find more or less \"optimal\" window (that gives better ml metrics) to calculate statistics through an optimization experiment. </p>",
      "rawMarkdown": "Thank you @kmldas for mentioning my work in this discussion and its evaluation. Also, pls, I'm sorry for my late answer as I was on vacation.  \nActually, guys, there is no secret in my kernel :) I generated a set of features using tsfresh and applied catboost with even default params, everything is posted in the notebook. \nHonestly, yeah, I tried to tune catboost, checking symmetric and non-symmetric trees, applying the feature selection methods from sklearn. But all checked hypotheses didn't help to overcome the score obtained with the ingv-tsfresh-7730 dataset https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730  and default catboostregressor. \nConcerning NA imputing, I used simple fillna with zeros for time series with more than 50% of NAs.   \n\nCurrently, I have several hypotheses, one of them is to find more or less \"optimal\" window (that gives better ml metrics) to calculate statistics through an optimization experiment. ",
      "votes": 1
    },
    {
      "id": 1104154,
      "postDate": "2020-12-06T17:01:42.640Z",
      "content": "<p>Something that I know would work very well is too basically just hack the test set by running kstests on all train samples versus the test samples so that you can train on the most ideal samples.</p>",
      "rawMarkdown": "Something that I know would work very well is too basically just hack the test set by running kstests on all train samples versus the test samples so that you can train on the most ideal samples.",
      "votes": 1
    },
    {
      "id": 1103632,
      "postDate": "2020-12-06T04:52:41.380Z",
      "content": "<p>it's very helpful and insightful, thanks for sharing this.</p>",
      "rawMarkdown": "it's very helpful and insightful, thanks for sharing this.",
      "votes": 1
    },
    {
      "id": 1102711,
      "postDate": "2020-12-05T08:16:53.597Z",
      "content": "<p>We have a wonderful notebook shared by <a href=\"https://www.kaggle.com/carpediemamigo\" target=\"_blank\">@carpediemamigo</a> <br>\n<a href=\"https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh</a></p>\n<p>Alexander Lyubchenko's notebook is a great one… only 44 of us have managed to cross his score</p>\n<p>I was wondering what your secret sauce is to cross the score?<br>\nComing out of MoA comp, I tried blending/stacking a few models and that has helped me reach #30</p>\n<p>what about others? What is working and not working in this competition?</p>\n<p>Thanks for sharing!!</p>",
      "rawMarkdown": "We have a wonderful notebook shared by @carpediemamigo \nhttps://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\n\nAlexander Lyubchenko's notebook is a great one... only 44 of us have managed to cross his score\n\nI was wondering what your secret sauce is to cross the score?\nComing out of MoA comp, I tried blending/stacking a few models and that has helped me reach #30\n\nwhat about others? What is working and not working in this competition?\n\nThanks for sharing!!",
      "votes": 1
    },
    {
      "id": 1103290,
      "postDate": "2020-12-05T19:45:07.207Z",
      "content": "<p>I am using CWT to extract information from signals.<br>\nto have more information I am trying to use the NA's by some method of impute, I tried impute with KNN, but it did not work well. Then with impute by median, it improved a little. So far what has worked best for me is impute by mean.</p>\n<p>As I generated many characteristics, I had to select, I tried with several methods, the only thing that has given me any result is correlation with target and genetic algorithms.</p>\n<p>I tried various algorithms for prediction, although at the moment I am testing with XGBoost and h2o's autoML, both with similar results.</p>\n<p>In my case, I think the main problem I have is finding and selecting the features that have relevant information.</p>",
      "rawMarkdown": "I am using CWT to extract information from signals.\nto have more information I am trying to use the NA's by some method of impute, I tried impute with KNN, but it did not work well. Then with impute by median, it improved a little. So far what has worked best for me is impute by mean.\n\nAs I generated many characteristics, I had to select, I tried with several methods, the only thing that has given me any result is correlation with target and genetic algorithms.\n\nI tried various algorithms for prediction, although at the moment I am testing with XGBoost and h2o's autoML, both with similar results.\n\nIn my case, I think the main problem I have is finding and selecting the features that have relevant information.",
      "votes": 2,
      "replies": [
        {
          "id": 1103604,
          "postDate": "2020-12-06T04:25:39.613Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1103644,
      "postDate": "2020-12-06T05:00:16.363Z",
      "content": "<p>Thanks its very helpful</p>",
      "rawMarkdown": "Thanks its very helpful"
    }
  ],
  "comments": [
    {
      "id": 1116625,
      "author_name": "Alexander Lyubchenko",
      "author_url": "",
      "post_date": "2020-12-17T10:44:17.083000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/kmldas\" target=\"_blank\">@kmldas</a> for mentioning my work in this discussion and its evaluation. Also, pls, I'm sorry for my late answer as I was on vacation.  <br>\nActually, guys, there is no secret in my kernel :) I generated a set of features using tsfresh and applied catboost with even default params, everything is posted in the notebook. <br>\nHonestly, yeah, I tried to tune catboost, checking symmetric and non-symmetric trees, applying the feature selection methods from sklearn. But all checked hypotheses didn't help to overcome the score obtained with the ingv-tsfresh-7730 dataset <a href=\"https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730\" target=\"_blank\">https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730</a>  and default catboostregressor. <br>\nConcerning NA imputing, I used simple fillna with zeros for time series with more than 50% of NAs.   </p>\n<p>Currently, I have several hypotheses, one of them is to find more or less \"optimal\" window (that gives better ml metrics) to calculate statistics through an optimization experiment. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1104154,
      "author_name": "Pythonian",
      "author_url": "",
      "post_date": "2020-12-06T17:01:42.640000",
      "content": "<p>Something that I know would work very well is too basically just hack the test set by running kstests on all train samples versus the test samples so that you can train on the most ideal samples.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1103632,
      "author_name": "Arpit Bhushan Sharma",
      "author_url": "",
      "post_date": "2020-12-06T04:52:41.380000",
      "content": "<p>it's very helpful and insightful, thanks for sharing this.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1103290,
      "author_name": "Desareca",
      "author_url": "",
      "post_date": "2020-12-05T19:45:07.207000",
      "content": "<p>I am using CWT to extract information from signals.<br>\nto have more information I am trying to use the NA's by some method of impute, I tried impute with KNN, but it did not work well. Then with impute by median, it improved a little. So far what has worked best for me is impute by mean.</p>\n<p>As I generated many characteristics, I had to select, I tried with several methods, the only thing that has given me any result is correlation with target and genetic algorithms.</p>\n<p>I tried various algorithms for prediction, although at the moment I am testing with XGBoost and h2o's autoML, both with similar results.</p>\n<p>In my case, I think the main problem I have is finding and selecting the features that have relevant information.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1103604,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-06T04:25:39.613000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1103644,
      "author_name": "Arpit Bhushan Sharma",
      "author_url": "",
      "post_date": "2020-12-06T05:00:16.363000",
      "content": "<p>Thanks its very helpful</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1116625": "Thank you @kmldas for mentioning my work in this discussion and its evaluation. Also, pls, I'm sorry for my late answer as I was on vacation.  \nActually, guys, there is no secret in my kernel :) I generated a set of features using tsfresh and applied catboost with even default params, everything is posted in the notebook. \nHonestly, yeah, I tried to tune catboost, checking symmetric and non-symmetric trees, applying the feature selection methods from sklearn. But all checked hypotheses didn't help to overcome the score obtained with the ingv-tsfresh-7730 dataset https://www.kaggle.com/carpediemamigo/ingv-tsfresh-7730  and default catboostregressor. \nConcerning NA imputing, I used simple fillna with zeros for time series with more than 50% of NAs.   \n\nCurrently, I have several hypotheses, one of them is to find more or less \"optimal\" window (that gives better ml metrics) to calculate statistics through an optimization experiment. ",
    "1104154": "Something that I know would work very well is too basically just hack the test set by running kstests on all train samples versus the test samples so that you can train on the most ideal samples.",
    "1103632": "it's very helpful and insightful, thanks for sharing this.",
    "1102711": "We have a wonderful notebook shared by @carpediemamigo \nhttps://www.kaggle.com/carpediemamigo/ingv-catboost-baseline-tsfresh\n\nAlexander Lyubchenko's notebook is a great one... only 44 of us have managed to cross his score\n\nI was wondering what your secret sauce is to cross the score?\nComing out of MoA comp, I tried blending/stacking a few models and that has helped me reach #30\n\nwhat about others? What is working and not working in this competition?\n\nThanks for sharing!!",
    "1103290": "I am using CWT to extract information from signals.\nto have more information I am trying to use the NA's by some method of impute, I tried impute with KNN, but it did not work well. Then with impute by median, it improved a little. So far what has worked best for me is impute by mean.\n\nAs I generated many characteristics, I had to select, I tried with several methods, the only thing that has given me any result is correlation with target and genetic algorithms.\n\nI tried various algorithms for prediction, although at the moment I am testing with XGBoost and h2o's autoML, both with similar results.\n\nIn my case, I think the main problem I have is finding and selecting the features that have relevant information.",
    "1103644": "Thanks its very helpful"
  }
}