{
  "id": 189341,
  "title": "Last 3 CV -6.8746 solution which suffered giant shake up to 1525th place",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189341",
  "author_name": "resistance0108",
  "post_date": "2020-10-07T10:00:30.337000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congrats to winners of this shaky competition. For me, it's very disheartening result…<br>\nI made effort to achieve good CV score with 3 fold and it was successful at CV -6.8746.<br>\nI believed that my solution will at least give a silver medal, but my CV betrayed.</p>\n<p>This is my solution:</p>\n<p><strong>Features</strong><br>\nI used tabular features from this notebook<br>\n<a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a><br>\nand image features that my teammate <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a> engineered (following).</p>\n<p>total_lung_kurtosis<br>\n max_slice_lung_skew<br>\n min_slice_lung_skew<br>\n total_lung_skew<br>\n lung_height<br>\n total_lung_volume</p>\n<p>This helped much for improving CV (about + 0.05), but private score became dismal.</p>\n<p><strong>Model 1: Lasso</strong><br>\nlast 3 MAE 182.20  last 3 CV -6.8892</p>\n<p>I made this model on the final day of this comp.<br>\nI thought Lasso greatly reduced overfitting, but…<br>\nThis model is following process, all of prediction process is Lasso:</p>\n<ol>\n<li>Assume we predict fold 3, construct training data with fold 1&amp;2 data + fold 3 first visit data * 6 (6 times sampled). Simultaneously, construct training data with fold 1&amp;2 data + test first visit data * 2 for test prediciton</li>\n<li>Predict first FVC value and replace first FVC value with (prediciton + actual value) / 2</li>\n<li>Predict FVC with adjusted first FVC value + other feats for both training data and validation/ test data, then take absolute value of prediction error of training data</li>\n<li>Predict absolute value of error then Confidence = max( predicted  error * √2, 70)</li>\n</ol>\n<p>Using first visit data of validation data worked greatly for CV.</p>\n<p><strong>Model 2: 4 layers MLP</strong><br>\nlast 3 MAE 185.51  last 3 CV -6.9057</p>\n<p>This is by my teammate <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> .<br>\nThis model predicts FVC and Confidence separately.<br>\nUsing Huber loss for FVC and quantile loss for Confidence. </p>\n<p><strong>Ensemble</strong><br>\nFVC = Lasso * 0.65 + MLP * 0.35<br>\nConfidence = Lasso * 0.5 + MLP * 0.5<br>\nlast 3 MAE 180.50  last 3 CV -6.8746</p>",
  "messages": [
    {
      "id": 1040721,
      "postDate": "2020-10-07T10:00:30.337Z",
      "content": "<p>Congrats to winners of this shaky competition. For me, it's very disheartening result…<br>\nI made effort to achieve good CV score with 3 fold and it was successful at CV -6.8746.<br>\nI believed that my solution will at least give a silver medal, but my CV betrayed.</p>\n<p>This is my solution:</p>\n<p><strong>Features</strong><br>\nI used tabular features from this notebook<br>\n<a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter</a><br>\nand image features that my teammate <a href=\"https://www.kaggle.com/samshipengs\" target=\"_blank\">@samshipengs</a> engineered (following).</p>\n<p>total_lung_kurtosis<br>\n max_slice_lung_skew<br>\n min_slice_lung_skew<br>\n total_lung_skew<br>\n lung_height<br>\n total_lung_volume</p>\n<p>This helped much for improving CV (about + 0.05), but private score became dismal.</p>\n<p><strong>Model 1: Lasso</strong><br>\nlast 3 MAE 182.20  last 3 CV -6.8892</p>\n<p>I made this model on the final day of this comp.<br>\nI thought Lasso greatly reduced overfitting, but…<br>\nThis model is following process, all of prediction process is Lasso:</p>\n<ol>\n<li>Assume we predict fold 3, construct training data with fold 1&amp;2 data + fold 3 first visit data * 6 (6 times sampled). Simultaneously, construct training data with fold 1&amp;2 data + test first visit data * 2 for test prediciton</li>\n<li>Predict first FVC value and replace first FVC value with (prediciton + actual value) / 2</li>\n<li>Predict FVC with adjusted first FVC value + other feats for both training data and validation/ test data, then take absolute value of prediction error of training data</li>\n<li>Predict absolute value of error then Confidence = max( predicted  error * √2, 70)</li>\n</ol>\n<p>Using first visit data of validation data worked greatly for CV.</p>\n<p><strong>Model 2: 4 layers MLP</strong><br>\nlast 3 MAE 185.51  last 3 CV -6.9057</p>\n<p>This is by my teammate <a href=\"https://www.kaggle.com/drtausamaru\" target=\"_blank\">@drtausamaru</a> .<br>\nThis model predicts FVC and Confidence separately.<br>\nUsing Huber loss for FVC and quantile loss for Confidence. </p>\n<p><strong>Ensemble</strong><br>\nFVC = Lasso * 0.65 + MLP * 0.35<br>\nConfidence = Lasso * 0.5 + MLP * 0.5<br>\nlast 3 MAE 180.50  last 3 CV -6.8746</p>",
      "rawMarkdown": "Congrats to winners of this shaky competition. For me, it's very disheartening result...\nI made effort to achieve good CV score with 3 fold and it was successful at CV -6.8746.\nI believed that my solution will at least give a silver medal, but my CV betrayed.\n\nThis is my solution:\n\n**Features**\nI used tabular features from this notebook\nhttps://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\nand image features that my teammate @samshipengs engineered (following).\n\n total_lung_kurtosis\n max_slice_lung_skew\n min_slice_lung_skew\n total_lung_skew\n lung_height\n total_lung_volume\n\nThis helped much for improving CV (about + 0.05), but private score became dismal.\n\n**Model 1: Lasso**\nlast 3 MAE 182.20  last 3 CV -6.8892\n\nI made this model on the final day of this comp.\nI thought Lasso greatly reduced overfitting, but...\nThis model is following process, all of prediction process is Lasso:\n\n1. Assume we predict fold 3, construct training data with fold 1&2 data + fold 3 first visit data * 6 (6 times sampled). Simultaneously, construct training data with fold 1&2 data + test first visit data * 2 for test prediciton\n2. Predict first FVC value and replace first FVC value with (prediciton + actual value) / 2\n3. Predict FVC with adjusted first FVC value + other feats for both training data and validation/ test data, then take absolute value of prediction error of training data\n4. Predict absolute value of error then Confidence = max( predicted  error * √2, 70)\n\nUsing first visit data of validation data worked greatly for CV.\n\n**Model 2: 4 layers MLP**\nlast 3 MAE 185.51  last 3 CV -6.9057\n\nThis is by my teammate @drtausamaru .\nThis model predicts FVC and Confidence separately.\nUsing Huber loss for FVC and quantile loss for Confidence. \n\n**Ensemble**\nFVC = Lasso * 0.65 + MLP * 0.35\nConfidence = Lasso * 0.5 + MLP * 0.5\nlast 3 MAE 180.50  last 3 CV -6.8746",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1040721": "Congrats to winners of this shaky competition. For me, it's very disheartening result...\nI made effort to achieve good CV score with 3 fold and it was successful at CV -6.8746.\nI believed that my solution will at least give a silver medal, but my CV betrayed.\n\nThis is my solution:\n\n**Features**\nI used tabular features from this notebook\nhttps://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\nand image features that my teammate @samshipengs engineered (following).\n\n total_lung_kurtosis\n max_slice_lung_skew\n min_slice_lung_skew\n total_lung_skew\n lung_height\n total_lung_volume\n\nThis helped much for improving CV (about + 0.05), but private score became dismal.\n\n**Model 1: Lasso**\nlast 3 MAE 182.20  last 3 CV -6.8892\n\nI made this model on the final day of this comp.\nI thought Lasso greatly reduced overfitting, but...\nThis model is following process, all of prediction process is Lasso:\n\n1. Assume we predict fold 3, construct training data with fold 1&2 data + fold 3 first visit data * 6 (6 times sampled). Simultaneously, construct training data with fold 1&2 data + test first visit data * 2 for test prediciton\n2. Predict first FVC value and replace first FVC value with (prediciton + actual value) / 2\n3. Predict FVC with adjusted first FVC value + other feats for both training data and validation/ test data, then take absolute value of prediction error of training data\n4. Predict absolute value of error then Confidence = max( predicted  error * √2, 70)\n\nUsing first visit data of validation data worked greatly for CV.\n\n**Model 2: 4 layers MLP**\nlast 3 MAE 185.51  last 3 CV -6.9057\n\nThis is by my teammate @drtausamaru .\nThis model predicts FVC and Confidence separately.\nUsing Huber loss for FVC and quantile loss for Confidence. \n\n**Ensemble**\nFVC = Lasso * 0.65 + MLP * 0.35\nConfidence = Lasso * 0.5 + MLP * 0.5\nlast 3 MAE 180.50  last 3 CV -6.8746"
  }
}