{
  "id": 189318,
  "title": "5th place solution",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/writeups/lukereijnen-5th-place-solution",
  "author_name": "",
  "post_date": "2020-10-07T09:01:47.957566200Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, congratulations to the winners and thanks to kaggle and the hosts for this competition. Special thanks to <a href=\"https://www.kaggle.com/LukeReijnen\" target=\"_blank\">@LukeReijnen</a> for being an awesome teammate. I would also like to thank the community for sharing their ideas on what worked and did not work and for reinforcing the idea to trust your cv. </p>\n<h5>Model:</h5>\n<p>Now to the model. We used a small network with only tabular data. It had the following input features: <code>[WeekInit, WeekTarget, WeekDiff, FVC, Percent, Age, Sex, CurrentlySmokes, Ex-smoker, Never Smoked]</code>. These were followed by two dense hidden layers of each 32 nodes with the swish activation function. The output consisted of a node for the FVC prediction and a node for the direct prediction of sigma. To make the network predict values between 0 and 1, the outputs were multiplied by 5000 and 500 respectively.</p>\n<h5>Loss function:</h5>\n<p>We initially used the competition metric. However as we wanted a higher emphasis on MAE we used the loss function:<br>\n$$\\frac{\\sqrt2 \\Delta}{70} + \\frac{\\sqrt2 \\Delta}{\\sigma_{clipped}} + \\ln(\\sqrt2 \\sigma_{clipped})$$</p>\n<h5>Validation:</h5>\n<p>As validation technique, we used group K-fold. We split the patients into 5 folds of roughly equal size. As train and validation set we used all combinations of data for each patient with <code>WeekTarget &gt; WeekInit</code>. In other words, given a measurement of a patient, we trained on predicting all future measurements.</p>\n<h5>Data augmentation:</h5>\n<p>On the training set, we added Gaussian noise to the input. For FVC we used a rather big standard deviation of 500 mL. However, as we added the same noise to the target FVC that was to be predicted this turned out to work great. To change the percent feature accordingly, we added <code>Percent*FVC_noise/FVC</code> to it on top of the noise on percent. Finally, we normalized the input to lie roughly between 0 and 1.</p>\n<p>As prediction, we used the average of the FVC predictions across folds. For sigma, we used the quadratic mean of the predicted sigmas across folds.</p>\n<p>This was an interesting competition from which I learned a lot. We tried a lot of techniques that did not work out in the end but were fun implementing anyways. It was tempting to look at the public leaderboard score but we had often reminded ourselves to trust our CV. With a jump of ~1500 from the public LB that seems to have paid off. We are very pleased with the results and enjoyed this competition.</p>",
  "messages": [
    {
      "id": "1040618",
      "postDate": "10/07/2020 09:01:47",
      "content": "<p>First of all, congratulations to the winners and thanks to kaggle and the hosts for this competition. Special thanks to <a href=\"https://www.kaggle.com/LukeReijnen\" target=\"_blank\">@LukeReijnen</a> for being an awesome teammate. I would also like to thank the community for sharing their ideas on what worked and did not work and for reinforcing the idea to trust your cv. </p>\n<h5>Model:</h5>\n<p>Now to the model. We used a small network with only tabular data. It had the following input features: <code>[WeekInit, WeekTarget, WeekDiff, FVC, Percent, Age, Sex, CurrentlySmokes, Ex-smoker, Never Smoked]</code>. These were followed by two dense hidden layers of each 32 nodes with the swish activation function. The output consisted of a node for the FVC prediction and a node for the direct prediction of sigma. To make the network predict values between 0 and 1, the outputs were multiplied by 5000 and 500 respectively.</p>\n<h5>Loss function:</h5>\n<p>We initially used the competition metric. However as we wanted a higher emphasis on MAE we used the loss function:<br>\n$$\\frac{\\sqrt2 \\Delta}{70} + \\frac{\\sqrt2 \\Delta}{\\sigma_{clipped}} + \\ln(\\sqrt2 \\sigma_{clipped})$$</p>\n<h5>Validation:</h5>\n<p>As validation technique, we used group K-fold. We split the patients into 5 folds of roughly equal size. As train and validation set we used all combinations of data for each patient with <code>WeekTarget &gt; WeekInit</code>. In other words, given a measurement of a patient, we trained on predicting all future measurements.</p>\n<h5>Data augmentation:</h5>\n<p>On the training set, we added Gaussian noise to the input. For FVC we used a rather big standard deviation of 500 mL. However, as we added the same noise to the target FVC that was to be predicted this turned out to work great. To change the percent feature accordingly, we added <code>Percent*FVC_noise/FVC</code> to it on top of the noise on percent. Finally, we normalized the input to lie roughly between 0 and 1.</p>\n<p>As prediction, we used the average of the FVC predictions across folds. For sigma, we used the quadratic mean of the predicted sigmas across folds.</p>\n<p>This was an interesting competition from which I learned a lot. We tried a lot of techniques that did not work out in the end but were fun implementing anyways. It was tempting to look at the public leaderboard score but we had often reminded ourselves to trust our CV. With a jump of ~1500 from the public LB that seems to have paid off. We are very pleased with the results and enjoyed this competition.</p>",
      "rawMarkdown": "First of all, congratulations to the winners and thanks to kaggle and the hosts for this competition. Special thanks to @LukeReijnen for being an awesome teammate. I would also like to thank the community for sharing their ideas on what worked and did not work and for reinforcing the idea to trust your cv. \n\n##### Model:\nNow to the model. We used a small network with only tabular data. It had the following input features: `[WeekInit, WeekTarget, WeekDiff, FVC, Percent, Age, Sex, CurrentlySmokes, Ex-smoker, Never Smoked]`. These were followed by two dense hidden layers of each 32 nodes with the swish activation function. The output consisted of a node for the FVC prediction and a node for the direct prediction of sigma. To make the network predict values between 0 and 1, the outputs were multiplied by 5000 and 500 respectively.\n\n#####Loss function:\nWe initially used the competition metric. However as we wanted a higher emphasis on MAE we used the loss function:\n$$\\frac{\\sqrt2 \\Delta}{70} + \\frac{\\sqrt2 \\Delta}{\\sigma_{clipped}} + \\ln(\\sqrt2 \\sigma_{clipped})$$\n\n#####Validation:\nAs validation technique, we used group K-fold. We split the patients into 5 folds of roughly equal size. As train and validation set we used all combinations of data for each patient with `WeekTarget > WeekInit`. In other words, given a measurement of a patient, we trained on predicting all future measurements.\n\n#####Data augmentation:\nOn the training set, we added Gaussian noise to the input. For FVC we used a rather big standard deviation of 500 mL. However, as we added the same noise to the target FVC that was to be predicted this turned out to work great. To change the percent feature accordingly, we added `Percent*FVC_noise/FVC` to it on top of the noise on percent. Finally, we normalized the input to lie roughly between 0 and 1.\n\nAs prediction, we used the average of the FVC predictions across folds. For sigma, we used the quadratic mean of the predicted sigmas across folds.\n\nThis was an interesting competition from which I learned a lot. We tried a lot of techniques that did not work out in the end but were fun implementing anyways. It was tempting to look at the public leaderboard score but we had often reminded ourselves to trust our CV. With a jump of ~1500 from the public LB that seems to have paid off. We are very pleased with the results and enjoyed this competition.",
      "votes": null
    },
    {
      "id": "1040631",
      "postDate": "10/07/2020 09:08:08",
      "content": "<p>Congratulations, I try the same augmentation idea using different deviation per Patients and deviation was Patient FVC std, it doesn't work for me.</p>",
      "rawMarkdown": "Congratulations, I try the same augmentation idea using different deviation per Patients and deviation was Patient FVC std, it doesn't work for me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040631,
      "author_name": "amedprof",
      "author_url": "",
      "post_date": "10/07/2020 09:08:08",
      "content": "<p>Congratulations, I try the same augmentation idea using different deviation per Patients and deviation was Patient FVC std, it doesn't work for me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040618": "First of all, congratulations to the winners and thanks to kaggle and the hosts for this competition. Special thanks to @LukeReijnen for being an awesome teammate. I would also like to thank the community for sharing their ideas on what worked and did not work and for reinforcing the idea to trust your cv. \n\n##### Model:\nNow to the model. We used a small network with only tabular data. It had the following input features: `[WeekInit, WeekTarget, WeekDiff, FVC, Percent, Age, Sex, CurrentlySmokes, Ex-smoker, Never Smoked]`. These were followed by two dense hidden layers of each 32 nodes with the swish activation function. The output consisted of a node for the FVC prediction and a node for the direct prediction of sigma. To make the network predict values between 0 and 1, the outputs were multiplied by 5000 and 500 respectively.\n\n#####Loss function:\nWe initially used the competition metric. However as we wanted a higher emphasis on MAE we used the loss function:\n$$\\frac{\\sqrt2 \\Delta}{70} + \\frac{\\sqrt2 \\Delta}{\\sigma_{clipped}} + \\ln(\\sqrt2 \\sigma_{clipped})$$\n\n#####Validation:\nAs validation technique, we used group K-fold. We split the patients into 5 folds of roughly equal size. As train and validation set we used all combinations of data for each patient with `WeekTarget > WeekInit`. In other words, given a measurement of a patient, we trained on predicting all future measurements.\n\n#####Data augmentation:\nOn the training set, we added Gaussian noise to the input. For FVC we used a rather big standard deviation of 500 mL. However, as we added the same noise to the target FVC that was to be predicted this turned out to work great. To change the percent feature accordingly, we added `Percent*FVC_noise/FVC` to it on top of the noise on percent. Finally, we normalized the input to lie roughly between 0 and 1.\n\nAs prediction, we used the average of the FVC predictions across folds. For sigma, we used the quadratic mean of the predicted sigmas across folds.\n\nThis was an interesting competition from which I learned a lot. We tried a lot of techniques that did not work out in the end but were fun implementing anyways. It was tempting to look at the public leaderboard score but we had often reminded ourselves to trust our CV. With a jump of ~1500 from the public LB that seems to have paid off. We are very pleased with the results and enjoyed this competition.",
    "1040631": "Congratulations, I try the same augmentation idea using different deviation per Patients and deviation was Patient FVC std, it doesn't work for me."
  },
  "source": "meta"
}