{
  "id": 415967,
  "title": "Was there a secret?",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/415967",
  "author_name": "Fritz Cremer",
  "post_date": "2023-06-09T00:07:29.601000",
  "votes": 9,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Congrats to all the winners! Up until the end, I wasn't that satisfied with my approach and felt I left out a lot of possibilities to improve my model as I only used the time series data to construct features. Also, I thought a lot of top teams might be overfitting. Turns out the shakeup is not that large, at least not for top positions. So given that there are relatively few scores above 0.4 or 0.35, was there some kind of special ingredient or was it just the tedious task of getting every little bit out of all the methods you could think of?</p>",
  "messages": [
    {
      "id": 2293075,
      "postDate": "2023-06-09T00:07:29.603Z",
      "content": "<p>Congrats to all the winners! Up until the end, I wasn't that satisfied with my approach and felt I left out a lot of possibilities to improve my model as I only used the time series data to construct features. Also, I thought a lot of top teams might be overfitting. Turns out the shakeup is not that large, at least not for top positions. So given that there are relatively few scores above 0.4 or 0.35, was there some kind of special ingredient or was it just the tedious task of getting every little bit out of all the methods you could think of?</p>",
      "rawMarkdown": "Congrats to all the winners! Up until the end, I wasn't that satisfied with my approach and felt I left out a lot of possibilities to improve my model as I only used the time series data to construct features. Also, I thought a lot of top teams might be overfitting. Turns out the shakeup is not that large, at least not for top positions. So given that there are relatively few scores above 0.4 or 0.35, was there some kind of special ingredient or was it just the tedious task of getting every little bit out of all the methods you could think of?",
      "votes": 9
    },
    {
      "id": 2293198,
      "postDate": "2023-06-09T02:52:46.680Z",
      "content": "<p>I found notype data overlapped with train &amp; test(subject-level). Maybe there are some leak information. But I still don't figure out how to use it.<br>\n Waiting for the top solutions.</p>",
      "rawMarkdown": "I found notype data overlapped with train & test(subject-level). Maybe there are some leak information. But I still don't figure out how to use it.\n Waiting for the top solutions.",
      "votes": 3
    },
    {
      "id": 2293130,
      "postDate": "2023-06-09T01:04:58.953Z",
      "content": "<p>I am particularly curious how the famous <strong>time_frac</strong> feature works in the test set. If it's good, how will they apply the algorithm in real-world applications with this feature😂</p>",
      "rawMarkdown": "I am particularly curious how the famous **time_frac** feature works in the test set. If it's good, how will they apply the algorithm in real-world applications with this feature😂",
      "votes": 1,
      "replies": [
        {
          "id": 2293239,
          "postDate": "2023-06-09T03:57:29.680Z",
          "content": "<p>I think one of the main challenge of this competition is defog dataset with long time observation .simple time frac didn’t strong enough to cover it unless constructing some short active time interval .The active interval construct method could be a simulate of tdcsfog dataset , in this way time frac maybe useful</p>",
          "rawMarkdown": "I think one of the main challenge of this competition is defog dataset with long time observation .simple time frac didn’t strong enough to cover it unless constructing some short active time interval .The active interval construct method could be a simulate of tdcsfog dataset , in this way time frac maybe useful",
          "votes": 1
        }
      ]
    },
    {
      "id": 2293108,
      "postDate": "2023-06-09T00:37:49.283Z",
      "content": "<p>Waiting for top solutions too!! what's the magic??</p>",
      "rawMarkdown": "Waiting for top solutions too!! what's the magic??",
      "votes": 1
    },
    {
      "id": 2293106,
      "postDate": "2023-06-09T00:30:52.480Z",
      "content": "<p>I'm waiting for the top solutions.. </p>",
      "rawMarkdown": "I'm waiting for the top solutions.. ",
      "votes": 2
    },
    {
      "id": 2293139,
      "postDate": "2023-06-09T01:12:49.787Z",
      "content": "<p>I guess top solutions must be using some kind of multi-stage deep learning sequence models. With sklearn classification models even after trying all feature engineering with inclusion of fourier transformation, wavelets, step and duration calculations etc., and using possibilities like Multioutput, multiclass, OVR, Calibration, imbalance rectification, stacking them all..etc., the score never crossed 0.315 in public LB, while a basic regression notebook gave a score of 0.311!  The multiclass models could beat regression models, but multioutput classification models could not even beat that 0.311! I joined late and after realizing that deep learning sequence models might do better, I lost interest!</p>",
      "rawMarkdown": "I guess top solutions must be using some kind of multi-stage deep learning sequence models. With sklearn classification models even after trying all feature engineering with inclusion of fourier transformation, wavelets, step and duration calculations etc., and using possibilities like Multioutput, multiclass, OVR, Calibration, imbalance rectification, stacking them all..etc., the score never crossed 0.315 in public LB, while a basic regression notebook gave a score of 0.311!  The multiclass models could beat regression models, but multioutput classification models could not even beat that 0.311! I joined late and after realizing that deep learning sequence models might do better, I lost interest!",
      "replies": [
        {
          "id": 2293325,
          "postDate": "2023-06-09T05:09:41.993Z",
          "content": "<p>This was my observation as well. Moreover, improvements in GroupKFold CV did not correspond to improvements in public LB, even significant ones, which I found confusing. Will have to go over submissions and confirm that CV was indeed king.</p>",
          "rawMarkdown": "This was my observation as well. Moreover, improvements in GroupKFold CV did not correspond to improvements in public LB, even significant ones, which I found confusing. Will have to go over submissions and confirm that CV was indeed king."
        }
      ]
    },
    {
      "id": 2293112,
      "postDate": "2023-06-09T00:53:44.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2293198,
      "author_name": "hyd",
      "author_url": "",
      "post_date": "2023-06-09T02:52:46.680000",
      "content": "<p>I found notype data overlapped with train &amp; test(subject-level). Maybe there are some leak information. But I still don't figure out how to use it.<br>\n Waiting for the top solutions.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2293130,
      "author_name": "gyhs9bkybhy",
      "author_url": "",
      "post_date": "2023-06-09T01:04:58.953000",
      "content": "<p>I am particularly curious how the famous <strong>time_frac</strong> feature works in the test set. If it's good, how will they apply the algorithm in real-world applications with this feature😂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2293239,
          "author_name": "此般浅薄",
          "author_url": "",
          "post_date": "2023-06-09T03:57:29.680000",
          "content": "<p>I think one of the main challenge of this competition is defog dataset with long time observation .simple time frac didn’t strong enough to cover it unless constructing some short active time interval .The active interval construct method could be a simulate of tdcsfog dataset , in this way time frac maybe useful</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2293108,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2023-06-09T00:37:49.283000",
      "content": "<p>Waiting for top solutions too!! what's the magic??</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2293106,
      "author_name": "olivepicker",
      "author_url": "",
      "post_date": "2023-06-09T00:30:52.480000",
      "content": "<p>I'm waiting for the top solutions.. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2293139,
      "author_name": "Murugesan Narayanaswamy",
      "author_url": "",
      "post_date": "2023-06-09T01:12:49.787000",
      "content": "<p>I guess top solutions must be using some kind of multi-stage deep learning sequence models. With sklearn classification models even after trying all feature engineering with inclusion of fourier transformation, wavelets, step and duration calculations etc., and using possibilities like Multioutput, multiclass, OVR, Calibration, imbalance rectification, stacking them all..etc., the score never crossed 0.315 in public LB, while a basic regression notebook gave a score of 0.311!  The multiclass models could beat regression models, but multioutput classification models could not even beat that 0.311! I joined late and after realizing that deep learning sequence models might do better, I lost interest!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2293325,
          "author_name": "Yijie Xu",
          "author_url": "",
          "post_date": "2023-06-09T05:09:41.993000",
          "content": "<p>This was my observation as well. Moreover, improvements in GroupKFold CV did not correspond to improvements in public LB, even significant ones, which I found confusing. Will have to go over submissions and confirm that CV was indeed king.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2293112,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-09T00:53:44.237000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2293075": "Congrats to all the winners! Up until the end, I wasn't that satisfied with my approach and felt I left out a lot of possibilities to improve my model as I only used the time series data to construct features. Also, I thought a lot of top teams might be overfitting. Turns out the shakeup is not that large, at least not for top positions. So given that there are relatively few scores above 0.4 or 0.35, was there some kind of special ingredient or was it just the tedious task of getting every little bit out of all the methods you could think of?",
    "2293198": "I found notype data overlapped with train & test(subject-level). Maybe there are some leak information. But I still don't figure out how to use it.\n Waiting for the top solutions.",
    "2293130": "I am particularly curious how the famous **time_frac** feature works in the test set. If it's good, how will they apply the algorithm in real-world applications with this feature😂",
    "2293108": "Waiting for top solutions too!! what's the magic??",
    "2293106": "I'm waiting for the top solutions.. ",
    "2293139": "I guess top solutions must be using some kind of multi-stage deep learning sequence models. With sklearn classification models even after trying all feature engineering with inclusion of fourier transformation, wavelets, step and duration calculations etc., and using possibilities like Multioutput, multiclass, OVR, Calibration, imbalance rectification, stacking them all..etc., the score never crossed 0.315 in public LB, while a basic regression notebook gave a score of 0.311!  The multiclass models could beat regression models, but multioutput classification models could not even beat that 0.311! I joined late and after realizing that deep learning sequence models might do better, I lost interest!",
    "2293112": ""
  }
}