{
  "id": 416057,
  "title": "2nd place solution",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/waiwai-2nd-place-solution",
  "author_name": "",
  "post_date": "2023-06-20T04:49:52.173Z",
  "votes": 62,
  "comment_count": 20,
  "views": 0,
  "content": "<h2>2nd place solution</h2>\n<p>First of all, thanks to the host for an interesting competition and congratulations to all the winners!</p>\n<h2>Summary</h2>\n<ul>\n<li>We designed separate models for tDCS FOG and DeFog.</li>\n<li>Both tDCS FOG and DeFog were trained using GRUs.</li>\n<li>Each Id was split into sequences of a specified length. During training, we used a shorter length (e.g., 1000 for tdcsfog, 5000 for defog), but for inference, a longer length (e.g., 5000 for tdcsfog, 30000 for defog) was applied. This emerged as the most crucial factor in this competition.</li>\n<li>For the DeFog model, we utilized the ‘notype’ data for training with pseudo-labels, significantly improving the scores; by leveraging the Event column, robust pseudo-labels could be created.</li>\n<li>Although CV and the public score were generally correlated, they became uncorrelated as the public score increased. Additionally, the CV was quite variable and occasionally surprisingly high. Therefore, we employed the public score for model selection, while CV guided the sequence length.</li>\n</ul>\n<h2>tDCS FOG</h2>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>Three provided accelerations</li>\n<li>Feature engineering performed on each acceleration involved:<ul>\n<li>Difference between prior and subsequent accelerations</li>\n<li>Cumulative sum of acceleration</li></ul></li></ul></li>\n<li><p>Standardization</p>\n<ul>\n<li>Used RobustScaler</li>\n<li>Each Id was standardized individually</li></ul></li>\n<li><p>Sequence Creation Method:</p>\n<ul>\n<li><p>Training</p>\n<ul>\n<li>Sequence length: 1000<ul>\n<li>Sequences were created by shifting 500 steps from the starting position.</li></ul></li></ul></li>\n<li><p>Inference</p>\n<ul>\n<li>Sequence length: 3000 or 5000 (3000 if data size was less than 5000, 5000 otherwise)</li>\n<li>Sequences were created by shifting 1500 or 2500 steps from the starting position.<ul>\n<li>During prediction, the sequence section from 750/1250 to 2250/3750 was utilized.</li>\n<li>The initial segment spanned from 0 to 2250/3750, while the final segment used from 750/1250 to the end of the sequence. </li></ul></li></ul></li></ul></li>\n<li><p>Models</p>\n<ul>\n<li><p>For each target, we ensembled four models.</p></li>\n<li><p>The following settings were common to each model</p>\n<ul>\n<li>Model : GRU</li>\n<li>Cross validation method : StratifiedGroupKFold<ul>\n<li>group : Subject</li></ul></li>\n<li>Loss : BCEWithLogitsLoss</li>\n<li>Optimizer : AdamW</li>\n<li>Scheduler : get_linear_schedule_with_warmup<ul>\n<li>(Although not verified in detail) get_linear_schedule_with_warmup seemd to work better in CV than get_cosine_schedule_with_warmup.</li></ul></li>\n<li>Sequence length<ul>\n<li>Train : 1000</li>\n<li>Inference : 3000 / 5000</li>\n<li>Training the model with a longer sequence did not improve CV or public score. However, training with a short sequence and performing inference with a long sequence significantly improved both CV and public score. </li></ul></li></ul></li>\n<li><p>Model1</p>\n<ul>\n<li>This model trained with equal loss for each target.</li>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.493 <br><br>\n(Sequence length 1000 : 0.438)</li></ul></li>\n<li>Ensemble weight : 0.2</li></ul></li>\n<li><p>Model2</p>\n<ul>\n<li>The loss weight of one target was set to 0.6, and the remaining targets were set at 0.4.<ul>\n<li>The following three patterns<ul>\n<li>StartHesitation : 0.6 , Turn &amp; Walking : 0.4</li>\n<li>Turn : 0.6 , StartHesitation &amp; Walking : 0.4</li>\n<li>Walking : 0.6 , StartHesitation &amp; Turn : 0.4</li></ul></li></ul></li>\n<li>The model was saved at the epoch where the target with the weight set to 0.6 had the best score.</li>\n<li>Only the predictions with the loss weight set to 0.6 were used in the test predictions.</li>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.520</li></ul></li>\n<li>Ensemble weight : 0.4</li></ul></li>\n<li><p>Model3 &amp; 4</p>\n<ul>\n<li>The loss weight for two targets was set to 0.8, and the remaining target was set at 0.2.<ul>\n<li>The following three patterns<ul>\n<li>StartHesitation &amp; Turn : 0.8 , Walking : 0.2</li>\n<li>StartHesitation &amp; Walking : 0.8 , Turn : 0.2</li>\n<li>Turn &amp; Walking : 0.8 , StartHesitation : 0.2</li></ul></li></ul></li>\n<li>The model was saved at the epoch where the two targets with the weight set to 0.8 had the best score.</li>\n<li>Only the predictions with a loss weight set to 0.8 were used in the test predictions.</li>\n<li>CV <ul>\n<li>Sequence length 3000 / 5000 : 0.536</li></ul></li>\n<li>Ensemble weight : 0.4</li></ul></li>\n<li><p>ensemble</p>\n<ul>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.537</li></ul></li></ul></li></ul></li>\n</ul>\n<h2>DeFog</h2>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>Three provided accelerations</li>\n<li>Feature engineering performed for each acceleration included<ul>\n<li>Difference between prior and subsequent accelerations</li></ul></li></ul></li>\n<li><p>Standardization</p>\n<ul>\n<li>Used StandardScaler</li>\n<li>Each Id was standardized individually</li></ul></li>\n<li><p>Sequence Creation Method:</p>\n<ul>\n<li><p>Training</p>\n<ul>\n<li>Sequence length: 5000<ul>\n<li>Sequences were created by shifting 2500 steps from the starting position.</li></ul></li></ul></li>\n<li><p>Inference</p>\n<ul>\n<li>Sequence length: 15000 or 30000 (15000 if data size was less than 200000, 30000 otherwise)</li>\n<li>Sequences were created by shifting 7500 or 15000 steps from the starting position.<ul>\n<li>During prediction, the sequence section from 3750/7500 to 11250/22500 was utilized.</li>\n<li>The initial segment spanned from 0 to 11250/22500, while the final segment used from 3750/7500 to the end of the sequence.</li></ul></li></ul></li></ul></li>\n<li><p>Models</p>\n<ul>\n<li><p>We ensembled five models </p></li>\n<li><p>The following settings are common to each model</p>\n<ul>\n<li>Model : GRU</li>\n<li>Cross validation method : StratifiedGroupKFold<ul>\n<li>group : Subject</li></ul></li>\n<li>Optimizer : AdamW</li>\n<li>Loss : BCEWithLogitsLoss</li>\n<li>Scheduler : get_linear_schedule_with_warmup</li>\n<li>Sequence length<ul>\n<li>Train : 5000</li>\n<li>Inference : 15000 / 30000</li></ul></li>\n<li>The loss weights for each target were uniform</li>\n<li>Only instances where both 'Valid' and 'Task' were true were considered for loss calculation.</li></ul></li>\n<li><p>model1</p>\n<ul>\n<li>CV<ul>\n<li>Sequence length 15000 / 30000 : 0.279</li></ul></li>\n<li>Ensemble weight : 0.35</li></ul></li>\n<li><p>model2</p>\n<ul>\n<li>Utilized the first round of pseudo-labeling<ul>\n<li>Applied hard labels, with the label set to 1 only if the data value of the 'Event' was 1, otherwise it was set to 0</li>\n<li>The label was determined based on the highest predictive value among the three target predictions</li>\n<li>Inference results from sequences of length 15000 from model1 were used</li></ul></li>\n<li>The application of pseudo-labeling significantly improved both public and private scores</li>\n<li>CV<ul>\n<li>Sequence length 15000 / 30000 : 0.306</li></ul></li>\n<li>Ensemble weight : 0.25</li></ul></li>\n<li><p>model3</p>\n<ul>\n<li>Utilized the second round of pseudo-labeling</li>\n<li>CV<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.313</li></ul></li>\n<li>Ensemble weight : 0.25</li></ul></li>\n<li><p>model4</p>\n<ul>\n<li>Increased the hidden size of the GRU</li>\n<li>Utilized the first round of pseudo-labeling</li>\n<li>CV<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.3393</li></ul></li>\n<li>Ensemble weight : 0.10</li></ul></li>\n<li><p>model5</p>\n<ul>\n<li>Trained with all data</li>\n<li>Utilized pseudo-labeling</li>\n<li>Ensemble weight : 0.05</li></ul></li>\n<li><p>ensemble(excluding model5)</p>\n<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.33706</li></ul></li></ul></li>\n</ul>\n<h2>tDCS FOG &amp; DeFog</h2>\n<ul>\n<li>CV : 0.548</li>\n<li>Public Score : 0.530</li>\n<li>Private Score : 0.450</li>\n</ul>\n<h2>Inference notebook</h2>\n<p><a href=\"https://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1\" target=\"_blank\">https://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1</a></p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction\" target=\"_blank\">https://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction</a></p>",
  "messages": [
    {
      "id": "2293629",
      "postDate": "06/09/2023 10:59:26",
      "content": "<h2>2nd place solution</h2>\n<p>First of all, thanks to the host for an interesting competition and congratulations to all the winners!</p>\n<h2>Summary</h2>\n<ul>\n<li>We designed separate models for tDCS FOG and DeFog.</li>\n<li>Both tDCS FOG and DeFog were trained using GRUs.</li>\n<li>Each Id was split into sequences of a specified length. During training, we used a shorter length (e.g., 1000 for tdcsfog, 5000 for defog), but for inference, a longer length (e.g., 5000 for tdcsfog, 30000 for defog) was applied. This emerged as the most crucial factor in this competition.</li>\n<li>For the DeFog model, we utilized the ‘notype’ data for training with pseudo-labels, significantly improving the scores; by leveraging the Event column, robust pseudo-labels could be created.</li>\n<li>Although CV and the public score were generally correlated, they became uncorrelated as the public score increased. Additionally, the CV was quite variable and occasionally surprisingly high. Therefore, we employed the public score for model selection, while CV guided the sequence length.</li>\n</ul>\n<h2>tDCS FOG</h2>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>Three provided accelerations</li>\n<li>Feature engineering performed on each acceleration involved:<ul>\n<li>Difference between prior and subsequent accelerations</li>\n<li>Cumulative sum of acceleration</li></ul></li></ul></li>\n<li><p>Standardization</p>\n<ul>\n<li>Used RobustScaler</li>\n<li>Each Id was standardized individually</li></ul></li>\n<li><p>Sequence Creation Method:</p>\n<ul>\n<li><p>Training</p>\n<ul>\n<li>Sequence length: 1000<ul>\n<li>Sequences were created by shifting 500 steps from the starting position.</li></ul></li></ul></li>\n<li><p>Inference</p>\n<ul>\n<li>Sequence length: 3000 or 5000 (3000 if data size was less than 5000, 5000 otherwise)</li>\n<li>Sequences were created by shifting 1500 or 2500 steps from the starting position.<ul>\n<li>During prediction, the sequence section from 750/1250 to 2250/3750 was utilized.</li>\n<li>The initial segment spanned from 0 to 2250/3750, while the final segment used from 750/1250 to the end of the sequence. </li></ul></li></ul></li></ul></li>\n<li><p>Models</p>\n<ul>\n<li><p>For each target, we ensembled four models.</p></li>\n<li><p>The following settings were common to each model</p>\n<ul>\n<li>Model : GRU</li>\n<li>Cross validation method : StratifiedGroupKFold<ul>\n<li>group : Subject</li></ul></li>\n<li>Loss : BCEWithLogitsLoss</li>\n<li>Optimizer : AdamW</li>\n<li>Scheduler : get_linear_schedule_with_warmup<ul>\n<li>(Although not verified in detail) get_linear_schedule_with_warmup seemd to work better in CV than get_cosine_schedule_with_warmup.</li></ul></li>\n<li>Sequence length<ul>\n<li>Train : 1000</li>\n<li>Inference : 3000 / 5000</li>\n<li>Training the model with a longer sequence did not improve CV or public score. However, training with a short sequence and performing inference with a long sequence significantly improved both CV and public score. </li></ul></li></ul></li>\n<li><p>Model1</p>\n<ul>\n<li>This model trained with equal loss for each target.</li>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.493 <br><br>\n(Sequence length 1000 : 0.438)</li></ul></li>\n<li>Ensemble weight : 0.2</li></ul></li>\n<li><p>Model2</p>\n<ul>\n<li>The loss weight of one target was set to 0.6, and the remaining targets were set at 0.4.<ul>\n<li>The following three patterns<ul>\n<li>StartHesitation : 0.6 , Turn &amp; Walking : 0.4</li>\n<li>Turn : 0.6 , StartHesitation &amp; Walking : 0.4</li>\n<li>Walking : 0.6 , StartHesitation &amp; Turn : 0.4</li></ul></li></ul></li>\n<li>The model was saved at the epoch where the target with the weight set to 0.6 had the best score.</li>\n<li>Only the predictions with the loss weight set to 0.6 were used in the test predictions.</li>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.520</li></ul></li>\n<li>Ensemble weight : 0.4</li></ul></li>\n<li><p>Model3 &amp; 4</p>\n<ul>\n<li>The loss weight for two targets was set to 0.8, and the remaining target was set at 0.2.<ul>\n<li>The following three patterns<ul>\n<li>StartHesitation &amp; Turn : 0.8 , Walking : 0.2</li>\n<li>StartHesitation &amp; Walking : 0.8 , Turn : 0.2</li>\n<li>Turn &amp; Walking : 0.8 , StartHesitation : 0.2</li></ul></li></ul></li>\n<li>The model was saved at the epoch where the two targets with the weight set to 0.8 had the best score.</li>\n<li>Only the predictions with a loss weight set to 0.8 were used in the test predictions.</li>\n<li>CV <ul>\n<li>Sequence length 3000 / 5000 : 0.536</li></ul></li>\n<li>Ensemble weight : 0.4</li></ul></li>\n<li><p>ensemble</p>\n<ul>\n<li>CV<ul>\n<li>Sequence length 3000 / 5000 : 0.537</li></ul></li></ul></li></ul></li>\n</ul>\n<h2>DeFog</h2>\n<ul>\n<li><p>Features</p>\n<ul>\n<li>Three provided accelerations</li>\n<li>Feature engineering performed for each acceleration included<ul>\n<li>Difference between prior and subsequent accelerations</li></ul></li></ul></li>\n<li><p>Standardization</p>\n<ul>\n<li>Used StandardScaler</li>\n<li>Each Id was standardized individually</li></ul></li>\n<li><p>Sequence Creation Method:</p>\n<ul>\n<li><p>Training</p>\n<ul>\n<li>Sequence length: 5000<ul>\n<li>Sequences were created by shifting 2500 steps from the starting position.</li></ul></li></ul></li>\n<li><p>Inference</p>\n<ul>\n<li>Sequence length: 15000 or 30000 (15000 if data size was less than 200000, 30000 otherwise)</li>\n<li>Sequences were created by shifting 7500 or 15000 steps from the starting position.<ul>\n<li>During prediction, the sequence section from 3750/7500 to 11250/22500 was utilized.</li>\n<li>The initial segment spanned from 0 to 11250/22500, while the final segment used from 3750/7500 to the end of the sequence.</li></ul></li></ul></li></ul></li>\n<li><p>Models</p>\n<ul>\n<li><p>We ensembled five models </p></li>\n<li><p>The following settings are common to each model</p>\n<ul>\n<li>Model : GRU</li>\n<li>Cross validation method : StratifiedGroupKFold<ul>\n<li>group : Subject</li></ul></li>\n<li>Optimizer : AdamW</li>\n<li>Loss : BCEWithLogitsLoss</li>\n<li>Scheduler : get_linear_schedule_with_warmup</li>\n<li>Sequence length<ul>\n<li>Train : 5000</li>\n<li>Inference : 15000 / 30000</li></ul></li>\n<li>The loss weights for each target were uniform</li>\n<li>Only instances where both 'Valid' and 'Task' were true were considered for loss calculation.</li></ul></li>\n<li><p>model1</p>\n<ul>\n<li>CV<ul>\n<li>Sequence length 15000 / 30000 : 0.279</li></ul></li>\n<li>Ensemble weight : 0.35</li></ul></li>\n<li><p>model2</p>\n<ul>\n<li>Utilized the first round of pseudo-labeling<ul>\n<li>Applied hard labels, with the label set to 1 only if the data value of the 'Event' was 1, otherwise it was set to 0</li>\n<li>The label was determined based on the highest predictive value among the three target predictions</li>\n<li>Inference results from sequences of length 15000 from model1 were used</li></ul></li>\n<li>The application of pseudo-labeling significantly improved both public and private scores</li>\n<li>CV<ul>\n<li>Sequence length 15000 / 30000 : 0.306</li></ul></li>\n<li>Ensemble weight : 0.25</li></ul></li>\n<li><p>model3</p>\n<ul>\n<li>Utilized the second round of pseudo-labeling</li>\n<li>CV<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.313</li></ul></li>\n<li>Ensemble weight : 0.25</li></ul></li>\n<li><p>model4</p>\n<ul>\n<li>Increased the hidden size of the GRU</li>\n<li>Utilized the first round of pseudo-labeling</li>\n<li>CV<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.3393</li></ul></li>\n<li>Ensemble weight : 0.10</li></ul></li>\n<li><p>model5</p>\n<ul>\n<li>Trained with all data</li>\n<li>Utilized pseudo-labeling</li>\n<li>Ensemble weight : 0.05</li></ul></li>\n<li><p>ensemble(excluding model5)</p>\n<ul>\n<li>Sequence length 15000 &amp; 30000 : 0.33706</li></ul></li></ul></li>\n</ul>\n<h2>tDCS FOG &amp; DeFog</h2>\n<ul>\n<li>CV : 0.548</li>\n<li>Public Score : 0.530</li>\n<li>Private Score : 0.450</li>\n</ul>\n<h2>Inference notebook</h2>\n<p><a href=\"https://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1\" target=\"_blank\">https://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1</a></p>\n<h2>Code</h2>\n<p><a href=\"https://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction\" target=\"_blank\">https://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction</a></p>",
      "rawMarkdown": "## 2nd place solution\n\n\nFirst of all, thanks to the host for an interesting competition and congratulations to all the winners!\n\n## Summary\n- We designed separate models for tDCS FOG and DeFog.\n- Both tDCS FOG and DeFog were trained using GRUs.\n- Each Id was split into sequences of a specified length. During training, we used a shorter length (e.g., 1000 for tdcsfog, 5000 for defog), but for inference, a longer length (e.g., 5000 for tdcsfog, 30000 for defog) was applied. This emerged as the most crucial factor in this competition.\n- For the DeFog model, we utilized the ‘notype’ data for training with pseudo-labels, significantly improving the scores; by leveraging the Event column, robust pseudo-labels could be created.\n- Although CV and the public score were generally correlated, they became uncorrelated as the public score increased. Additionally, the CV was quite variable and occasionally surprisingly high. Therefore, we employed the public score for model selection, while CV guided the sequence length.\n\n## tDCS FOG\n- Features\n    - Three provided accelerations\n    - Feature engineering performed on each acceleration involved:\n        - Difference between prior and subsequent accelerations\n        - Cumulative sum of acceleration\n\n- Standardization\n    - Used RobustScaler\n    - Each Id was standardized individually\n\n- Sequence Creation Method:\n    - Training\n        - Sequence length: 1000\n            - Sequences were created by shifting 500 steps from the starting position.\n            \n    - Inference\n        - Sequence length: 3000 or 5000 (3000 if data size was less than 5000, 5000 otherwise)\n        - Sequences were created by shifting 1500 or 2500 steps from the starting position.\n            - During prediction, the sequence section from 750/1250 to 2250/3750 was utilized.\n            - The initial segment spanned from 0 to 2250/3750, while the final segment used from 750/1250 to the end of the sequence. \n\n- Models\n    - For each target, we ensembled four models.\n    - The following settings were common to each model\n        - Model : GRU\n        - Cross validation method : StratifiedGroupKFold\n            - group : Subject\n        - Loss : BCEWithLogitsLoss\n        - Optimizer : AdamW\n        - Scheduler : get_linear_schedule_with_warmup\n             - (Although not verified in detail) get_linear_schedule_with_warmup seemd to work better in CV than get_cosine_schedule_with_warmup.\n        - Sequence length\n            - Train : 1000\n            - Inference : 3000 / 5000\n            - Training the model with a longer sequence did not improve CV or public score. However, training with a short sequence and performing inference with a long sequence significantly improved both CV and public score. \n    - Model1\n        - This model trained with equal loss for each target.\n        - CV\n            - Sequence length 3000 / 5000 : 0.493 </br>\n            (Sequence length 1000 : 0.438)\n        - Ensemble weight : 0.2\n\n    - Model2\n        - The loss weight of one target was set to 0.6, and the remaining targets were set at 0.4.\n            - The following three patterns\n                - StartHesitation : 0.6 , Turn & Walking : 0.4\n                - Turn : 0.6 , StartHesitation & Walking : 0.4\n                - Walking : 0.6 , StartHesitation & Turn : 0.4\n        - The model was saved at the epoch where the target with the weight set to 0.6 had the best score.\n        - Only the predictions with the loss weight set to 0.6 were used in the test predictions.\n        - CV\n            - Sequence length 3000 / 5000 : 0.520\n        - Ensemble weight : 0.4\n\n    - Model3 & 4\n        - The loss weight for two targets was set to 0.8, and the remaining target was set at 0.2.\n            - The following three patterns\n                - StartHesitation & Turn : 0.8 , Walking : 0.2\n                - StartHesitation & Walking : 0.8 , Turn : 0.2\n                - Turn & Walking : 0.8 , StartHesitation : 0.2\n        -  The model was saved at the epoch where the two targets with the weight set to 0.8 had the best score.\n        - Only the predictions with a loss weight set to 0.8 were used in the test predictions.\n        - CV \n            - Sequence length 3000 / 5000 : 0.536\n        - Ensemble weight : 0.4\n\n    - ensemble\n        - CV\n            - Sequence length 3000 / 5000 : 0.537\n\n## DeFog\n- Features\n    - Three provided accelerations\n    - Feature engineering performed for each acceleration included\n        - Difference between prior and subsequent accelerations\n\n- Standardization\n    - Used StandardScaler\n    - Each Id was standardized individually\n\n- Sequence Creation Method:\n    - Training\n        - Sequence length: 5000\n            - Sequences were created by shifting 2500 steps from the starting position.\n            \n    - Inference\n        - Sequence length: 15000 or 30000 (15000 if data size was less than 200000, 30000 otherwise)\n        - Sequences were created by shifting 7500 or 15000 steps from the starting position.\n            - During prediction, the sequence section from 3750/7500 to 11250/22500 was utilized.\n            - The initial segment spanned from 0 to 11250/22500, while the final segment used from 3750/7500 to the end of the sequence.\n\n- Models\n    - We ensembled five models \n    - The following settings are common to each model\n        - Model : GRU\n        - Cross validation method : StratifiedGroupKFold\n            - group : Subject\n        - Optimizer : AdamW\n        - Loss : BCEWithLogitsLoss\n        - Scheduler : get_linear_schedule_with_warmup\n        - Sequence length\n            - Train : 5000\n            - Inference : 15000 / 30000\n        - The loss weights for each target were uniform\n        - Only instances where both 'Valid' and 'Task' were true were considered for loss calculation.\n\n    - model1\n        - CV\n            - Sequence length 15000 / 30000 : 0.279\n        - Ensemble weight : 0.35\n\n    - model2\n        - Utilized the first round of pseudo-labeling\n            - Applied hard labels, with the label set to 1 only if the data value of the 'Event' was 1, otherwise it was set to 0\n            - The label was determined based on the highest predictive value among the three target predictions\n            - Inference results from sequences of length 15000 from model1 were used\n        - The application of pseudo-labeling significantly improved both public and private scores\n        - CV\n            - Sequence length 15000 / 30000 : 0.306\n        - Ensemble weight : 0.25\n\n    - model3\n        - Utilized the second round of pseudo-labeling\n        - CV\n            - Sequence length 15000 & 30000 : 0.313\n        - Ensemble weight : 0.25\n\n    - model4\n        - Increased the hidden size of the GRU\n        - Utilized the first round of pseudo-labeling\n        - CV\n            - Sequence length 15000 & 30000 : 0.3393\n        - Ensemble weight : 0.10\n\n    - model5\n        - Trained with all data\n        - Utilized pseudo-labeling\n        - Ensemble weight : 0.05\n\n    - ensemble(excluding model5)\n        - Sequence length 15000 & 30000 : 0.33706\n\n## tDCS FOG & DeFog\n- CV : 0.548\n- Public Score : 0.530\n- Private Score : 0.450\n\n## Inference notebook\nhttps://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1\n\n## Code\nhttps://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction",
      "votes": null
    },
    {
      "id": "2293856",
      "postDate": "06/09/2023 14:44:59",
      "content": "<p>Nice solution. Just a question, if I am understanding the pseudo-labelling correctly, a trained model's best guess is used to convert the event column into one of the 3 classes. Was this labelling corroborated among differently trained model  (or an ensemble)?</p>",
      "rawMarkdown": "Nice solution. Just a question, if I am understanding the pseudo-labelling correctly, a trained model's best guess is used to convert the event column into one of the 3 classes. Was this labelling corroborated among differently trained model  (or an ensemble)?",
      "votes": null
    },
    {
      "id": "2295072",
      "postDate": "06/10/2023 15:06:35",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> !</p>\n<p>Interesting to see how you used different models for TDCSFOG and DEFOG series and different sequence lenght for training and inference…</p>\n<p>Congratulations and keep up with the great work💪🔥</p>",
      "rawMarkdown": "Greetings, @takoihiraokazu !\n\nInteresting to see how you used different models for TDCSFOG and DEFOG series and different sequence lenght for training and inference...\n\nCongratulations and keep up with the great work💪🔥",
      "votes": null
    },
    {
      "id": "2295451",
      "postDate": "06/11/2023 01:19:57",
      "content": "<p>Congratulations!! <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> <br>\nI have 3 questions about your solution.</p>\n<ol>\n<li>Why don't you use other model like LSTM, Transformer, N-Beats…<br>\nI think GRU is not often used.</li>\n<li>What did you do during CrossValidation? Such as avoiding bias in training data and validation data.</li>\n<li>You mentioned that the sequence length differs between training and inference, but is it possible to input data with different sequence lengths into the model? (I may not be understanding this correctly, sorry, let me know)</li>\n</ol>",
      "rawMarkdown": "Congratulations!! @takoihiraokazu \nI have 3 questions about your solution.\n1. Why don't you use other model like LSTM, Transformer, N-Beats...\n    I think GRU is not often used.\n2. What did you do during CrossValidation? Such as avoiding bias in training data and validation data.\n3. You mentioned that the sequence length differs between training and inference, but is it possible to input data with different sequence lengths into the model? (I may not be understanding this correctly, sorry, let me know)",
      "votes": null
    },
    {
      "id": "2296624",
      "postDate": "06/12/2023 02:38:51",
      "content": "<p>For your 3rd point, the model used GRUs (Gated Recurrent units). This layer recursively evaluate along a specified axis (in this case the the axis of the time-series). So variable inputs are accepted through this model.</p>",
      "rawMarkdown": "For your 3rd point, the model used GRUs (Gated Recurrent units). This layer recursively evaluate along a specified axis (in this case the the axis of the time-series). So variable inputs are accepted through this model.",
      "votes": null
    },
    {
      "id": "2297103",
      "postDate": "06/12/2023 11:10:39",
      "content": "<p>Congratulations waiwai team! And thanks a lot for sharing your great solution!<br>\nI have a couple of questions:</p>\n<ol>\n<li>regarding the target weight patterns you mention in your solution. Do you randomly apply those weights for an entire epoch and save the model for the target with the highest way? If so, this looks to me like a very creative way of training a model per target label within a single loop!</li>\n<li>did you any type of regularization technique?<br>\nThanks again!</li>\n</ol>",
      "rawMarkdown": "Congratulations waiwai team! And thanks a lot for sharing your great solution!\nI have a couple of questions:\n1. regarding the target weight patterns you mention in your solution. Do you randomly apply those weights for an entire epoch and save the model for the target with the highest way? If so, this looks to me like a very creative way of training a model per target label within a single loop!\n2. did you any type of regularization technique?\nThanks again!",
      "votes": null
    },
    {
      "id": "2297187",
      "postDate": "06/12/2023 12:08:39",
      "content": "<p>Thank you for your comment. No, I used one model to create a pseudo-label. As for how to create pseudo-labels, it is the same as in the following discussion. <br><a href=\"url\" target=\"_blank\"> https://www.kaggle.com/competitions/google-quest-challenge/discussion/129840</a></p>",
      "rawMarkdown": "Thank you for your comment. No, I used one model to create a pseudo-label. As for how to create pseudo-labels, it is the same as in the following discussion. </br>[ https://www.kaggle.com/competitions/google-quest-challenge/discussion/129840](url)",
      "votes": null
    },
    {
      "id": "2297202",
      "postDate": "06/12/2023 12:19:59",
      "content": "<p>Thanks!</p>\n<ol>\n<li><p>I chose to use GRU because it yielded the highest CV and public scores among the models I experimented with. I did try LSTM and Transformer, but they didn't perform as well as GRU in terms of scoring. While it's true that GRU is not as widely used in Kaggle competitions compared to LSTM or Transformer, it doesn't mean it's not used at all. If you search through past solutions, you can find approaches that utilize GRU.</p></li>\n<li><p>I didn't make any specific modifications for cross-validation. However, even without elaborate techniques, I observed a reasonable correlation between CV scores and both public and private scores.</p></li>\n<li><p>Yes, GRU can handle input sequences of different lengths.</p></li>\n</ol>",
      "rawMarkdown": "Thanks!\n\n1. I chose to use GRU because it yielded the highest CV and public scores among the models I experimented with. I did try LSTM and Transformer, but they didn't perform as well as GRU in terms of scoring. While it's true that GRU is not as widely used in Kaggle competitions compared to LSTM or Transformer, it doesn't mean it's not used at all. If you search through past solutions, you can find approaches that utilize GRU.\n\n2. I didn't make any specific modifications for cross-validation. However, even without elaborate techniques, I observed a reasonable correlation between CV scores and both public and private scores.\n\n3. Yes, GRU can handle input sequences of different lengths.",
      "votes": null
    },
    {
      "id": "2297217",
      "postDate": "06/12/2023 12:28:07",
      "content": "<p>Thanks!</p>\n<ol>\n<li><p>No, I didn't apply the weights randomly. For example, I first set the weight for the loss of the \"Turn\" target to 0.6 and trained the model using a 5-fold cross-validation. We followed the same approach for the other two targets, \"StartHesitation\" and \"Walking\".</p></li>\n<li><p>I didn't use any specific regularization techniques.</p></li>\n</ol>",
      "rawMarkdown": "Thanks!\n\n1. No, I didn't apply the weights randomly. For example, I first set the weight for the loss of the \"Turn\" target to 0.6 and trained the model using a 5-fold cross-validation. We followed the same approach for the other two targets, \"StartHesitation\" and \"Walking\".\n\n2. I didn't use any specific regularization techniques.",
      "votes": null
    },
    {
      "id": "2297232",
      "postDate": "06/12/2023 12:35:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san &amp; <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> -san!!<br>\nIt was a very good learning experience!</p>",
      "rawMarkdown": "Thanks @takoihiraokazu -san & @coderrkj -san!!\nIt was a very good learning experience!",
      "votes": null
    },
    {
      "id": "2297274",
      "postDate": "06/12/2023 13:26:50",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> for your quick reply. <br>\nI'm sorry I misunderstood you. It's clear now.</p>",
      "rawMarkdown": "Thanks, @takoihiraokazu for your quick reply. \nI'm sorry I misunderstood you. It's clear now.",
      "votes": null
    },
    {
      "id": "2299912",
      "postDate": "06/12/2023 23:56:15",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2300235",
      "postDate": "06/13/2023 05:09:40",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san,<br>\nThank you for sharing a great solution! I can learn a lot from you.<br>\nMay I ask why you tried to change the length of the sequence with learning and inferrence?<br>\nIs it a rule of thumb based on your experience ?  </p>",
      "rawMarkdown": "takoihiraokazu -san,\nThank you for sharing a great solution! I can learn a lot from you.\nMay I ask why you tried to change the length of the sequence with learning and inferrence?\nIs it a rule of thumb based on your experience ?",
      "votes": null
    },
    {
      "id": "2300717",
      "postDate": "06/13/2023 11:38:45",
      "content": "<p>Thanks for the comment.</p>\n<p>Firstly, in this competition, I believed it was important to extend the length of the sequence. This is because it enables a larger amount of information to be reflected in the predictions within the same ID. However, if the length of the sequence is increased in the training data, the score in the evaluation data does not seem to improve significantly, possibly due to overfitting to the training data. Therefore, as an experiment, I tried shortening the sequence length in the training data and lengthening it in the evaluation data, which greatly improved the score. So, I adopted this approach.</p>\n<p>Changing the sequence length during training and inference was something I tried for the first time, but the decision was based on my past experiences. For instance, in image competitions, it can be beneficial to increase the image size during inference compared to during training. Similarly, in matching competitions, it could be helpful to increase the number of candidates during inference.</p>",
      "rawMarkdown": "Thanks for the comment.\n\nFirstly, in this competition, I believed it was important to extend the length of the sequence. This is because it enables a larger amount of information to be reflected in the predictions within the same ID. However, if the length of the sequence is increased in the training data, the score in the evaluation data does not seem to improve significantly, possibly due to overfitting to the training data. Therefore, as an experiment, I tried shortening the sequence length in the training data and lengthening it in the evaluation data, which greatly improved the score. So, I adopted this approach.\n\nChanging the sequence length during training and inference was something I tried for the first time, but the decision was based on my past experiences. For instance, in image competitions, it can be beneficial to increase the image size during inference compared to during training. Similarly, in matching competitions, it could be helpful to increase the number of candidates during inference.",
      "votes": null
    },
    {
      "id": "2302862",
      "postDate": "06/14/2023 22:54:20",
      "content": "<p>Nice work ! <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> </p>",
      "rawMarkdown": "Nice work ! @takoihiraokazu",
      "votes": null
    },
    {
      "id": "2303392",
      "postDate": "06/15/2023 09:17:54",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san,<br>\nThank you for your kind reply. Your solution is based on theory and practice ! I respect your hard work!  Your team is our role model.</p>\n<p>I really appreciate your answer. Thank you so much !</p>",
      "rawMarkdown": "takoihiraokazu -san,\nThank you for your kind reply. Your solution is based on theory and practice ! I respect your hard work!  Your team is our role model.\n\nI really appreciate your answer. Thank you so much !",
      "votes": null
    },
    {
      "id": "2322879",
      "postDate": "06/29/2023 14:43:20",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> <br>\nThank you for sharing CODE in github.<br>\nIn your training code, there are 5 kinds of datasets which are original.</p>\n<ul>\n<li>id_path        = f\"../output/fe/fe022/fe022_id.parquet\"</li>\n<li>numerical_path = f\"../output/fe/fe022/fe022_num_array.npy\"</li>\n<li>target_path    = f\"../output/fe/fe022/fe022_target_array.npy\"</li>\n<li>mask_path      = f\"../output/fe/fe022/fe022_mask_array.npy\"</li>\n<li>pred_use_path  = f\"../output/fe/fe022/fe022_pred_use_array.npy\"</li>\n</ul>\n<p>I would like to know about each of datasets.<br>\nMaybe, <br>\nfe022_num_array.npy -&gt; input data (like AccV, AccAP, AccML)<br>\nfe022_target_array.npy -&gt; target data (like StartHesitation, Turn, Walking)<br>\nis it right?<br>\nHowever, I have no idea about fe022_mask_array.npy &amp; fe022_pred_use_array.npy.<br>\nWhat are those dataset? How do you make them.<br>\nPlease let me know if you are OK.</p>",
      "rawMarkdown": "takoihiraokazu \nThank you for sharing CODE in github.\nIn your training code, there are 5 kinds of datasets which are original.\n- id_path        = f\"../output/fe/fe022/fe022_id.parquet\"\n- numerical_path = f\"../output/fe/fe022/fe022_num_array.npy\"\n- target_path    = f\"../output/fe/fe022/fe022_target_array.npy\"\n- mask_path      = f\"../output/fe/fe022/fe022_mask_array.npy\"\n- pred_use_path  = f\"../output/fe/fe022/fe022_pred_use_array.npy\"\n\nI would like to know about each of datasets.\nMaybe, \nfe022_num_array.npy -> input data (like AccV, AccAP, AccML)\nfe022_target_array.npy -> target data (like StartHesitation, Turn, Walking)\nis it right?\nHowever, I have no idea about fe022_mask_array.npy & fe022_pred_use_array.npy.\nWhat are those dataset? How do you make them.\nPlease let me know if you are OK.",
      "votes": null
    },
    {
      "id": "2323619",
      "postDate": "06/30/2023 04:51:11",
      "content": "<p>Thanks for the comment.<br>\nYou are correct about fe022_num_array.npy and fe022_target_array.npy.</p>\n<p>As for fe022_mask_array.npy, it's created to represent the missing parts when generating sequences for each Id. Sometimes, the last sequence might not be long enough to meet the specified length, and that's where this array comes in. In terms of its usage, it's utilized during the computation of the loss in training. For instance, I don't include parts where mask=0 as demonstrated in the following line:</p>\n<pre><code>loss = (output, y)\n</code></pre>\n<p>fe022_pred_use_array.npy, on the other hand, is used to ensure that I don't utilize the sequence's edge portions during prediction, as this could lead to decreased accuracy. The way I use it is as follows: I specify only the parts I want to use for inference, then measure the CV or use it for the final submission. Here's an example:</p>\n<pre><code>pred_valid_index = val_pred_array == \nStartHesitation = (val_target_array,\n                              val_preds)\nTurn = (val_target_array,\n                               val_preds)\nWalking = (val_target_array,\n                              val_preds\n</code></pre>",
      "rawMarkdown": "Thanks for the comment.\nYou are correct about fe022_num_array.npy and fe022_target_array.npy.\n\nAs for fe022_mask_array.npy, it's created to represent the missing parts when generating sequences for each Id. Sometimes, the last sequence might not be long enough to meet the specified length, and that's where this array comes in. In terms of its usage, it's utilized during the computation of the loss in training. For instance, I don't include parts where mask=0 as demonstrated in the following line:\n```\n\nloss = criterion(output[input_data_mask_array == 1], y[input_data_mask_array == 1])\n\n```\n\nfe022_pred_use_array.npy, on the other hand, is used to ensure that I don't utilize the sequence's edge portions during prediction, as this could lead to decreased accuracy. The way I use it is as follows: I specify only the parts I want to use for inference, then measure the CV or use it for the final submission. Here's an example:\n```\npred_valid_index = val_pred_array == 1\nStartHesitation = average_precision_score(val_target_array[pred_valid_index][:,0],\n                              val_preds[pred_valid_index][:,0])\nTurn = average_precision_score(val_target_array[pred_valid_index][:,1],\n                               val_preds[pred_valid_index][:,1])\nWalking = average_precision_score(val_target_array[pred_valid_index][:,2],\n                              val_preds[pred_valid_index][:,2]\n```",
      "votes": null
    },
    {
      "id": "2327777",
      "postDate": "07/03/2023 06:59:20",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> Congratulations to your team! <br>\nWhat a nice solution, and thank you all for sharing, I have learnt a lot from your sharing!❤️❤️❤️</p>",
      "rawMarkdown": "takoihiraokazu Congratulations to your team! \nWhat a nice solution, and thank you all for sharing, I have learnt a lot from your sharing!❤️❤️❤️",
      "votes": null
    },
    {
      "id": "2328014",
      "postDate": "07/03/2023 09:47:35",
      "content": "<p>Thanks a lot!!</p>",
      "rawMarkdown": "Thanks a lot!!",
      "votes": null
    },
    {
      "id": "3015603",
      "postDate": "10/12/2024 16:49:37",
      "content": "<p>Hi, are you passing data frame between two models (tdcsfog and defog) and if there is dependency in between? I tried if I run the tdcsfog all models first then run defog, the defog results will be close to what you have submitted, but if I skip the tdcsfog and directly run defog, the results was not close to what you have submitted. The prediction probability are very small. </p>",
      "rawMarkdown": "Hi, are you passing data frame between two models (tdcsfog and defog) and if there is dependency in between? I tried if I run the tdcsfog all models first then run defog, the defog results will be close to what you have submitted, but if I skip the tdcsfog and directly run defog, the results was not close to what you have submitted. The prediction probability are very small.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2293856,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/09/2023 14:44:59",
      "content": "<p>Nice solution. Just a question, if I am understanding the pseudo-labelling correctly, a trained model's best guess is used to convert the event column into one of the 3 classes. Was this labelling corroborated among differently trained model  (or an ensemble)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2297187,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/12/2023 12:08:39",
          "content": "<p>Thank you for your comment. No, I used one model to create a pseudo-label. As for how to create pseudo-labels, it is the same as in the following discussion. <br><a href=\"url\" target=\"_blank\"> https://www.kaggle.com/competitions/google-quest-challenge/discussion/129840</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2295072,
      "author_name": "vladiluzjr",
      "author_url": "",
      "post_date": "06/10/2023 15:06:35",
      "content": "<p>Greetings, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> !</p>\n<p>Interesting to see how you used different models for TDCSFOG and DEFOG series and different sequence lenght for training and inference…</p>\n<p>Congratulations and keep up with the great work💪🔥</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2295451,
      "author_name": "hidebu",
      "author_url": "",
      "post_date": "06/11/2023 01:19:57",
      "content": "<p>Congratulations!! <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> <br>\nI have 3 questions about your solution.</p>\n<ol>\n<li>Why don't you use other model like LSTM, Transformer, N-Beats…<br>\nI think GRU is not often used.</li>\n<li>What did you do during CrossValidation? Such as avoiding bias in training data and validation data.</li>\n<li>You mentioned that the sequence length differs between training and inference, but is it possible to input data with different sequence lengths into the model? (I may not be understanding this correctly, sorry, let me know)</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2296624,
          "author_name": "coderrkj",
          "author_url": "",
          "post_date": "06/12/2023 02:38:51",
          "content": "<p>For your 3rd point, the model used GRUs (Gated Recurrent units). This layer recursively evaluate along a specified axis (in this case the the axis of the time-series). So variable inputs are accepted through this model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2297202,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/12/2023 12:19:59",
          "content": "<p>Thanks!</p>\n<ol>\n<li><p>I chose to use GRU because it yielded the highest CV and public scores among the models I experimented with. I did try LSTM and Transformer, but they didn't perform as well as GRU in terms of scoring. While it's true that GRU is not as widely used in Kaggle competitions compared to LSTM or Transformer, it doesn't mean it's not used at all. If you search through past solutions, you can find approaches that utilize GRU.</p></li>\n<li><p>I didn't make any specific modifications for cross-validation. However, even without elaborate techniques, I observed a reasonable correlation between CV scores and both public and private scores.</p></li>\n<li><p>Yes, GRU can handle input sequences of different lengths.</p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2297232,
              "author_name": "hidebu",
              "author_url": "",
              "post_date": "06/12/2023 12:35:54",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san &amp; <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> -san!!<br>\nIt was a very good learning experience!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2297103,
      "author_name": "oguiza",
      "author_url": "",
      "post_date": "06/12/2023 11:10:39",
      "content": "<p>Congratulations waiwai team! And thanks a lot for sharing your great solution!<br>\nI have a couple of questions:</p>\n<ol>\n<li>regarding the target weight patterns you mention in your solution. Do you randomly apply those weights for an entire epoch and save the model for the target with the highest way? If so, this looks to me like a very creative way of training a model per target label within a single loop!</li>\n<li>did you any type of regularization technique?<br>\nThanks again!</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2297217,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/12/2023 12:28:07",
          "content": "<p>Thanks!</p>\n<ol>\n<li><p>No, I didn't apply the weights randomly. For example, I first set the weight for the loss of the \"Turn\" target to 0.6 and trained the model using a 5-fold cross-validation. We followed the same approach for the other two targets, \"StartHesitation\" and \"Walking\".</p></li>\n<li><p>I didn't use any specific regularization techniques.</p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2297274,
              "author_name": "oguiza",
              "author_url": "",
              "post_date": "06/12/2023 13:26:50",
              "content": "<p>Thanks, <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> for your quick reply. <br>\nI'm sorry I misunderstood you. It's clear now.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2299912,
      "author_name": "sergeyfedorchenko",
      "author_url": "",
      "post_date": "06/12/2023 23:56:15",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2300235,
      "author_name": "nayu555",
      "author_url": "",
      "post_date": "06/13/2023 05:09:40",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san,<br>\nThank you for sharing a great solution! I can learn a lot from you.<br>\nMay I ask why you tried to change the length of the sequence with learning and inferrence?<br>\nIs it a rule of thumb based on your experience ?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 2300717,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/13/2023 11:38:45",
          "content": "<p>Thanks for the comment.</p>\n<p>Firstly, in this competition, I believed it was important to extend the length of the sequence. This is because it enables a larger amount of information to be reflected in the predictions within the same ID. However, if the length of the sequence is increased in the training data, the score in the evaluation data does not seem to improve significantly, possibly due to overfitting to the training data. Therefore, as an experiment, I tried shortening the sequence length in the training data and lengthening it in the evaluation data, which greatly improved the score. So, I adopted this approach.</p>\n<p>Changing the sequence length during training and inference was something I tried for the first time, but the decision was based on my past experiences. For instance, in image competitions, it can be beneficial to increase the image size during inference compared to during training. Similarly, in matching competitions, it could be helpful to increase the number of candidates during inference.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2303392,
              "author_name": "nayu555",
              "author_url": "",
              "post_date": "06/15/2023 09:17:54",
              "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> -san,<br>\nThank you for your kind reply. Your solution is based on theory and practice ! I respect your hard work!  Your team is our role model.</p>\n<p>I really appreciate your answer. Thank you so much !</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2302862,
      "author_name": "atefbouzid",
      "author_url": "",
      "post_date": "06/14/2023 22:54:20",
      "content": "<p>Nice work ! <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322879,
      "author_name": "hidebu",
      "author_url": "",
      "post_date": "06/29/2023 14:43:20",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> <br>\nThank you for sharing CODE in github.<br>\nIn your training code, there are 5 kinds of datasets which are original.</p>\n<ul>\n<li>id_path        = f\"../output/fe/fe022/fe022_id.parquet\"</li>\n<li>numerical_path = f\"../output/fe/fe022/fe022_num_array.npy\"</li>\n<li>target_path    = f\"../output/fe/fe022/fe022_target_array.npy\"</li>\n<li>mask_path      = f\"../output/fe/fe022/fe022_mask_array.npy\"</li>\n<li>pred_use_path  = f\"../output/fe/fe022/fe022_pred_use_array.npy\"</li>\n</ul>\n<p>I would like to know about each of datasets.<br>\nMaybe, <br>\nfe022_num_array.npy -&gt; input data (like AccV, AccAP, AccML)<br>\nfe022_target_array.npy -&gt; target data (like StartHesitation, Turn, Walking)<br>\nis it right?<br>\nHowever, I have no idea about fe022_mask_array.npy &amp; fe022_pred_use_array.npy.<br>\nWhat are those dataset? How do you make them.<br>\nPlease let me know if you are OK.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2323619,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/30/2023 04:51:11",
          "content": "<p>Thanks for the comment.<br>\nYou are correct about fe022_num_array.npy and fe022_target_array.npy.</p>\n<p>As for fe022_mask_array.npy, it's created to represent the missing parts when generating sequences for each Id. Sometimes, the last sequence might not be long enough to meet the specified length, and that's where this array comes in. In terms of its usage, it's utilized during the computation of the loss in training. For instance, I don't include parts where mask=0 as demonstrated in the following line:</p>\n<pre><code>loss = (output, y)\n</code></pre>\n<p>fe022_pred_use_array.npy, on the other hand, is used to ensure that I don't utilize the sequence's edge portions during prediction, as this could lead to decreased accuracy. The way I use it is as follows: I specify only the parts I want to use for inference, then measure the CV or use it for the final submission. Here's an example:</p>\n<pre><code>pred_valid_index = val_pred_array == \nStartHesitation = (val_target_array,\n                              val_preds)\nTurn = (val_target_array,\n                               val_preds)\nWalking = (val_target_array,\n                              val_preds\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2328014,
              "author_name": "hidebu",
              "author_url": "",
              "post_date": "07/03/2023 09:47:35",
              "content": "<p>Thanks a lot!!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2327777,
      "author_name": "abrachan",
      "author_url": "",
      "post_date": "07/03/2023 06:59:20",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> Congratulations to your team! <br>\nWhat a nice solution, and thank you all for sharing, I have learnt a lot from your sharing!❤️❤️❤️</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3015603,
      "author_name": "ronaldgchangm",
      "author_url": "",
      "post_date": "10/12/2024 16:49:37",
      "content": "<p>Hi, are you passing data frame between two models (tdcsfog and defog) and if there is dependency in between? I tried if I run the tdcsfog all models first then run defog, the defog results will be close to what you have submitted, but if I skip the tdcsfog and directly run defog, the results was not close to what you have submitted. The prediction probability are very small. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2293629": "## 2nd place solution\n\n\nFirst of all, thanks to the host for an interesting competition and congratulations to all the winners!\n\n## Summary\n- We designed separate models for tDCS FOG and DeFog.\n- Both tDCS FOG and DeFog were trained using GRUs.\n- Each Id was split into sequences of a specified length. During training, we used a shorter length (e.g., 1000 for tdcsfog, 5000 for defog), but for inference, a longer length (e.g., 5000 for tdcsfog, 30000 for defog) was applied. This emerged as the most crucial factor in this competition.\n- For the DeFog model, we utilized the ‘notype’ data for training with pseudo-labels, significantly improving the scores; by leveraging the Event column, robust pseudo-labels could be created.\n- Although CV and the public score were generally correlated, they became uncorrelated as the public score increased. Additionally, the CV was quite variable and occasionally surprisingly high. Therefore, we employed the public score for model selection, while CV guided the sequence length.\n\n## tDCS FOG\n- Features\n    - Three provided accelerations\n    - Feature engineering performed on each acceleration involved:\n        - Difference between prior and subsequent accelerations\n        - Cumulative sum of acceleration\n\n- Standardization\n    - Used RobustScaler\n    - Each Id was standardized individually\n\n- Sequence Creation Method:\n    - Training\n        - Sequence length: 1000\n            - Sequences were created by shifting 500 steps from the starting position.\n            \n    - Inference\n        - Sequence length: 3000 or 5000 (3000 if data size was less than 5000, 5000 otherwise)\n        - Sequences were created by shifting 1500 or 2500 steps from the starting position.\n            - During prediction, the sequence section from 750/1250 to 2250/3750 was utilized.\n            - The initial segment spanned from 0 to 2250/3750, while the final segment used from 750/1250 to the end of the sequence. \n\n- Models\n    - For each target, we ensembled four models.\n    - The following settings were common to each model\n        - Model : GRU\n        - Cross validation method : StratifiedGroupKFold\n            - group : Subject\n        - Loss : BCEWithLogitsLoss\n        - Optimizer : AdamW\n        - Scheduler : get_linear_schedule_with_warmup\n             - (Although not verified in detail) get_linear_schedule_with_warmup seemd to work better in CV than get_cosine_schedule_with_warmup.\n        - Sequence length\n            - Train : 1000\n            - Inference : 3000 / 5000\n            - Training the model with a longer sequence did not improve CV or public score. However, training with a short sequence and performing inference with a long sequence significantly improved both CV and public score. \n    - Model1\n        - This model trained with equal loss for each target.\n        - CV\n            - Sequence length 3000 / 5000 : 0.493 </br>\n            (Sequence length 1000 : 0.438)\n        - Ensemble weight : 0.2\n\n    - Model2\n        - The loss weight of one target was set to 0.6, and the remaining targets were set at 0.4.\n            - The following three patterns\n                - StartHesitation : 0.6 , Turn & Walking : 0.4\n                - Turn : 0.6 , StartHesitation & Walking : 0.4\n                - Walking : 0.6 , StartHesitation & Turn : 0.4\n        - The model was saved at the epoch where the target with the weight set to 0.6 had the best score.\n        - Only the predictions with the loss weight set to 0.6 were used in the test predictions.\n        - CV\n            - Sequence length 3000 / 5000 : 0.520\n        - Ensemble weight : 0.4\n\n    - Model3 & 4\n        - The loss weight for two targets was set to 0.8, and the remaining target was set at 0.2.\n            - The following three patterns\n                - StartHesitation & Turn : 0.8 , Walking : 0.2\n                - StartHesitation & Walking : 0.8 , Turn : 0.2\n                - Turn & Walking : 0.8 , StartHesitation : 0.2\n        -  The model was saved at the epoch where the two targets with the weight set to 0.8 had the best score.\n        - Only the predictions with a loss weight set to 0.8 were used in the test predictions.\n        - CV \n            - Sequence length 3000 / 5000 : 0.536\n        - Ensemble weight : 0.4\n\n    - ensemble\n        - CV\n            - Sequence length 3000 / 5000 : 0.537\n\n## DeFog\n- Features\n    - Three provided accelerations\n    - Feature engineering performed for each acceleration included\n        - Difference between prior and subsequent accelerations\n\n- Standardization\n    - Used StandardScaler\n    - Each Id was standardized individually\n\n- Sequence Creation Method:\n    - Training\n        - Sequence length: 5000\n            - Sequences were created by shifting 2500 steps from the starting position.\n            \n    - Inference\n        - Sequence length: 15000 or 30000 (15000 if data size was less than 200000, 30000 otherwise)\n        - Sequences were created by shifting 7500 or 15000 steps from the starting position.\n            - During prediction, the sequence section from 3750/7500 to 11250/22500 was utilized.\n            - The initial segment spanned from 0 to 11250/22500, while the final segment used from 3750/7500 to the end of the sequence.\n\n- Models\n    - We ensembled five models \n    - The following settings are common to each model\n        - Model : GRU\n        - Cross validation method : StratifiedGroupKFold\n            - group : Subject\n        - Optimizer : AdamW\n        - Loss : BCEWithLogitsLoss\n        - Scheduler : get_linear_schedule_with_warmup\n        - Sequence length\n            - Train : 5000\n            - Inference : 15000 / 30000\n        - The loss weights for each target were uniform\n        - Only instances where both 'Valid' and 'Task' were true were considered for loss calculation.\n\n    - model1\n        - CV\n            - Sequence length 15000 / 30000 : 0.279\n        - Ensemble weight : 0.35\n\n    - model2\n        - Utilized the first round of pseudo-labeling\n            - Applied hard labels, with the label set to 1 only if the data value of the 'Event' was 1, otherwise it was set to 0\n            - The label was determined based on the highest predictive value among the three target predictions\n            - Inference results from sequences of length 15000 from model1 were used\n        - The application of pseudo-labeling significantly improved both public and private scores\n        - CV\n            - Sequence length 15000 / 30000 : 0.306\n        - Ensemble weight : 0.25\n\n    - model3\n        - Utilized the second round of pseudo-labeling\n        - CV\n            - Sequence length 15000 & 30000 : 0.313\n        - Ensemble weight : 0.25\n\n    - model4\n        - Increased the hidden size of the GRU\n        - Utilized the first round of pseudo-labeling\n        - CV\n            - Sequence length 15000 & 30000 : 0.3393\n        - Ensemble weight : 0.10\n\n    - model5\n        - Trained with all data\n        - Utilized pseudo-labeling\n        - Ensemble weight : 0.05\n\n    - ensemble(excluding model5)\n        - Sequence length 15000 & 30000 : 0.33706\n\n## tDCS FOG & DeFog\n- CV : 0.548\n- Public Score : 0.530\n- Private Score : 0.450\n\n## Inference notebook\nhttps://www.kaggle.com/code/takoihiraokazu/cv-ensemble-sub-0607-1\n\n## Code\nhttps://github.com/TakoiHirokazu/Kaggle-Parkinsons-Freezing-of-Gait-Prediction",
    "2293856": "Nice solution. Just a question, if I am understanding the pseudo-labelling correctly, a trained model's best guess is used to convert the event column into one of the 3 classes. Was this labelling corroborated among differently trained model  (or an ensemble)?",
    "2295072": "Greetings, @takoihiraokazu !\n\nInteresting to see how you used different models for TDCSFOG and DEFOG series and different sequence lenght for training and inference...\n\nCongratulations and keep up with the great work💪🔥",
    "2295451": "Congratulations!! @takoihiraokazu \nI have 3 questions about your solution.\n1. Why don't you use other model like LSTM, Transformer, N-Beats...\n    I think GRU is not often used.\n2. What did you do during CrossValidation? Such as avoiding bias in training data and validation data.\n3. You mentioned that the sequence length differs between training and inference, but is it possible to input data with different sequence lengths into the model? (I may not be understanding this correctly, sorry, let me know)",
    "2296624": "For your 3rd point, the model used GRUs (Gated Recurrent units). This layer recursively evaluate along a specified axis (in this case the the axis of the time-series). So variable inputs are accepted through this model.",
    "2297103": "Congratulations waiwai team! And thanks a lot for sharing your great solution!\nI have a couple of questions:\n1. regarding the target weight patterns you mention in your solution. Do you randomly apply those weights for an entire epoch and save the model for the target with the highest way? If so, this looks to me like a very creative way of training a model per target label within a single loop!\n2. did you any type of regularization technique?\nThanks again!",
    "2297187": "Thank you for your comment. No, I used one model to create a pseudo-label. As for how to create pseudo-labels, it is the same as in the following discussion. </br>[ https://www.kaggle.com/competitions/google-quest-challenge/discussion/129840](url)",
    "2297202": "Thanks!\n\n1. I chose to use GRU because it yielded the highest CV and public scores among the models I experimented with. I did try LSTM and Transformer, but they didn't perform as well as GRU in terms of scoring. While it's true that GRU is not as widely used in Kaggle competitions compared to LSTM or Transformer, it doesn't mean it's not used at all. If you search through past solutions, you can find approaches that utilize GRU.\n\n2. I didn't make any specific modifications for cross-validation. However, even without elaborate techniques, I observed a reasonable correlation between CV scores and both public and private scores.\n\n3. Yes, GRU can handle input sequences of different lengths.",
    "2297217": "Thanks!\n\n1. No, I didn't apply the weights randomly. For example, I first set the weight for the loss of the \"Turn\" target to 0.6 and trained the model using a 5-fold cross-validation. We followed the same approach for the other two targets, \"StartHesitation\" and \"Walking\".\n\n2. I didn't use any specific regularization techniques.",
    "2297232": "Thanks @takoihiraokazu -san & @coderrkj -san!!\nIt was a very good learning experience!",
    "2297274": "Thanks, @takoihiraokazu for your quick reply. \nI'm sorry I misunderstood you. It's clear now.",
    "2299912": "Thanks for sharing!",
    "2300235": "takoihiraokazu -san,\nThank you for sharing a great solution! I can learn a lot from you.\nMay I ask why you tried to change the length of the sequence with learning and inferrence?\nIs it a rule of thumb based on your experience ?",
    "2300717": "Thanks for the comment.\n\nFirstly, in this competition, I believed it was important to extend the length of the sequence. This is because it enables a larger amount of information to be reflected in the predictions within the same ID. However, if the length of the sequence is increased in the training data, the score in the evaluation data does not seem to improve significantly, possibly due to overfitting to the training data. Therefore, as an experiment, I tried shortening the sequence length in the training data and lengthening it in the evaluation data, which greatly improved the score. So, I adopted this approach.\n\nChanging the sequence length during training and inference was something I tried for the first time, but the decision was based on my past experiences. For instance, in image competitions, it can be beneficial to increase the image size during inference compared to during training. Similarly, in matching competitions, it could be helpful to increase the number of candidates during inference.",
    "2302862": "Nice work ! @takoihiraokazu",
    "2303392": "takoihiraokazu -san,\nThank you for your kind reply. Your solution is based on theory and practice ! I respect your hard work!  Your team is our role model.\n\nI really appreciate your answer. Thank you so much !",
    "2322879": "takoihiraokazu \nThank you for sharing CODE in github.\nIn your training code, there are 5 kinds of datasets which are original.\n- id_path        = f\"../output/fe/fe022/fe022_id.parquet\"\n- numerical_path = f\"../output/fe/fe022/fe022_num_array.npy\"\n- target_path    = f\"../output/fe/fe022/fe022_target_array.npy\"\n- mask_path      = f\"../output/fe/fe022/fe022_mask_array.npy\"\n- pred_use_path  = f\"../output/fe/fe022/fe022_pred_use_array.npy\"\n\nI would like to know about each of datasets.\nMaybe, \nfe022_num_array.npy -> input data (like AccV, AccAP, AccML)\nfe022_target_array.npy -> target data (like StartHesitation, Turn, Walking)\nis it right?\nHowever, I have no idea about fe022_mask_array.npy & fe022_pred_use_array.npy.\nWhat are those dataset? How do you make them.\nPlease let me know if you are OK.",
    "2323619": "Thanks for the comment.\nYou are correct about fe022_num_array.npy and fe022_target_array.npy.\n\nAs for fe022_mask_array.npy, it's created to represent the missing parts when generating sequences for each Id. Sometimes, the last sequence might not be long enough to meet the specified length, and that's where this array comes in. In terms of its usage, it's utilized during the computation of the loss in training. For instance, I don't include parts where mask=0 as demonstrated in the following line:\n```\n\nloss = criterion(output[input_data_mask_array == 1], y[input_data_mask_array == 1])\n\n```\n\nfe022_pred_use_array.npy, on the other hand, is used to ensure that I don't utilize the sequence's edge portions during prediction, as this could lead to decreased accuracy. The way I use it is as follows: I specify only the parts I want to use for inference, then measure the CV or use it for the final submission. Here's an example:\n```\npred_valid_index = val_pred_array == 1\nStartHesitation = average_precision_score(val_target_array[pred_valid_index][:,0],\n                              val_preds[pred_valid_index][:,0])\nTurn = average_precision_score(val_target_array[pred_valid_index][:,1],\n                               val_preds[pred_valid_index][:,1])\nWalking = average_precision_score(val_target_array[pred_valid_index][:,2],\n                              val_preds[pred_valid_index][:,2]\n```",
    "2327777": "takoihiraokazu Congratulations to your team! \nWhat a nice solution, and thank you all for sharing, I have learnt a lot from your sharing!❤️❤️❤️",
    "2328014": "Thanks a lot!!",
    "3015603": "Hi, are you passing data frame between two models (tdcsfog and defog) and if there is dependency in between? I tried if I run the tdcsfog all models first then run defog, the defog results will be close to what you have submitted, but if I skip the tdcsfog and directly run defog, the results was not close to what you have submitted. The prediction probability are very small."
  },
  "source": "meta"
}