{
  "id": 190016,
  "title": "Some final thoughts on the competition",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/190016",
  "author_name": "Shoval T",
  "post_date": "2020-10-09T17:29:25.170000",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This was my first Kaggle competition, and boy did I choose a hard one to start with. I didn't have nearly as much time as I wanted, but that's not the point of this thread.</p>\n<p>What's been bothering me is this - <strong>was this competition a success?</strong> </p>\n<p>The first model I submitted was a very simple Ridge Regression model, based only on the data in <code>train.csv</code>. The confidence was determined by multiplying the predicted FVC by a constant (although optimized) fraction. I created the model just to see that I could submit successfully. It scored -6.9248 on the private leaderboard. The winning model scored -6.8305. In terms of position on the board itself this is a huge difference, but in real life, in terms of successful prediction, how much better is it really? </p>\n<p>Now, I don't know how others feel about it, but I am nowhere near to feeling intuitive around the Laplace Log Likelihood score they chose for this competition. Is a 0.1 difference really big? Did we even, as community, manage to harness the data in the thousands of images in any significant way? Will someone, some day be able to score a -5.0 on this competition? Perhaps even a -6.5?</p>\n<p>These are just some thoughts that have been running around my head in the last couple of days. <strong>I'd love to hear others' thoughts on this.</strong></p>",
  "messages": [
    {
      "id": 1044351,
      "postDate": "2020-10-09T17:29:25.170Z",
      "content": "<p>This was my first Kaggle competition, and boy did I choose a hard one to start with. I didn't have nearly as much time as I wanted, but that's not the point of this thread.</p>\n<p>What's been bothering me is this - <strong>was this competition a success?</strong> </p>\n<p>The first model I submitted was a very simple Ridge Regression model, based only on the data in <code>train.csv</code>. The confidence was determined by multiplying the predicted FVC by a constant (although optimized) fraction. I created the model just to see that I could submit successfully. It scored -6.9248 on the private leaderboard. The winning model scored -6.8305. In terms of position on the board itself this is a huge difference, but in real life, in terms of successful prediction, how much better is it really? </p>\n<p>Now, I don't know how others feel about it, but I am nowhere near to feeling intuitive around the Laplace Log Likelihood score they chose for this competition. Is a 0.1 difference really big? Did we even, as community, manage to harness the data in the thousands of images in any significant way? Will someone, some day be able to score a -5.0 on this competition? Perhaps even a -6.5?</p>\n<p>These are just some thoughts that have been running around my head in the last couple of days. <strong>I'd love to hear others' thoughts on this.</strong></p>",
      "rawMarkdown": "This was my first Kaggle competition, and boy did I choose a hard one to start with. I didn't have nearly as much time as I wanted, but that's not the point of this thread.\n\nWhat's been bothering me is this - **was this competition a success?** \n\nThe first model I submitted was a very simple Ridge Regression model, based only on the data in `train.csv`. The confidence was determined by multiplying the predicted FVC by a constant (although optimized) fraction. I created the model just to see that I could submit successfully. It scored -6.9248 on the private leaderboard. The winning model scored -6.8305. In terms of position on the board itself this is a huge difference, but in real life, in terms of successful prediction, how much better is it really? \n\nNow, I don't know how others feel about it, but I am nowhere near to feeling intuitive around the Laplace Log Likelihood score they chose for this competition. Is a 0.1 difference really big? Did we even, as community, manage to harness the data in the thousands of images in any significant way? Will someone, some day be able to score a -5.0 on this competition? Perhaps even a -6.5?\n\nThese are just some thoughts that have been running around my head in the last couple of days. **I'd love to hear others' thoughts on this.**",
      "votes": 4
    },
    {
      "id": 1045658,
      "postDate": "2020-10-10T21:31:29.420Z",
      "content": "<p>I created a notebook with some insights regarding the winning model and the LLL score in general: <a href=\"https://www.kaggle.com/shovalt/laplace-log-likelihood-review\" target=\"_blank\">https://www.kaggle.com/shovalt/laplace-log-likelihood-review</a></p>\n<p>What I did was use the training data to evaluate several scenarios. I would love to get your feedback.</p>\n<p>Here is the final graph of the notebook, as well as the summary. There are several more leading to it in the notebook itself:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5539072%2F81176f08b0bfd390aa41bd432a100d75%2FLLL_final_fig.png?generation=1602365355875649&amp;alt=media\" alt=\"\"></p>\n<p>To summarize:</p>\n<ol>\n<li>The CT images had little-to-no effect on the top-scoring prediction models in this contest.</li>\n<li>The use of linear decay as a prediction target was not useful, and <em>even a perfect prediction</em> of that target wouldn't have been very beneficial. </li>\n<li>A direct prediction of FVC could possibly yield much better results, which is probably why the basic tabular-data models succeeded pretty well, especially when combined with a decent estimation of sigma using the quantiles method. </li>\n<li>A perfect prediction in this competition would probably score around LLL=-4.5, which is 1-2 orders of magnitude higher than the range of scores in the top 1000 scoring kernels (between -6.90 and -6.83) - <strong>there is potentially much room to grow here!</strong></li>\n</ol>",
      "rawMarkdown": "I created a notebook with some insights regarding the winning model and the LLL score in general: https://www.kaggle.com/shovalt/laplace-log-likelihood-review\n\nWhat I did was use the training data to evaluate several scenarios. I would love to get your feedback.\n\nHere is the final graph of the notebook, as well as the summary. There are several more leading to it in the notebook itself:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5539072%2F81176f08b0bfd390aa41bd432a100d75%2FLLL_final_fig.png?generation=1602365355875649&alt=media)\n\nTo summarize:\n1. The CT images had little-to-no effect on the top-scoring prediction models in this contest.\n1. The use of linear decay as a prediction target was not useful, and *even a perfect prediction* of that target wouldn't have been very beneficial. \n1. A direct prediction of FVC could possibly yield much better results, which is probably why the basic tabular-data models succeeded pretty well, especially when combined with a decent estimation of sigma using the quantiles method. \n1. A perfect prediction in this competition would probably score around LLL=-4.5, which is 1-2 orders of magnitude higher than the range of scores in the top 1000 scoring kernels (between -6.90 and -6.83) - **there is potentially much room to grow here!**\n",
      "votes": 1
    },
    {
      "id": 1045142,
      "postDate": "2020-10-10T11:29:54.623Z",
      "content": "<p>Personally I think we cannot determine the metrics: there are many ways to decide how \"good\" a model is. For example, in normal image recognition we can use Accuracy, but in Medicine they use F1, or Specificity, since false negative should be minimized (you should not predict a patient is not ill if they are ill since that could cost a life). I think the organizers also try to pick the best one.\nAbout the difference in score, although maybe tiny, we only have one way to know: replace it with real numbers. Lowest possible score does not have limit (yet), so we should try more ways. The competition is ended, but the ideas to build better model is still burning.\nSince there scores are pretty near each other, change in leaderboard is inevitable though.</p>",
      "rawMarkdown": "Personally I think we cannot determine the metrics: there are many ways to decide how \"good\" a model is. For example, in normal image recognition we can use Accuracy, but in Medicine they use F1, or Specificity, since false negative should be minimized (you should not predict a patient is not ill if they are ill since that could cost a life). I think the organizers also try to pick the best one.\nAbout the difference in score, although maybe tiny, we only have one way to know: replace it with real numbers. Lowest possible score does not have limit (yet), so we should try more ways. The competition is ended, but the ideas to build better model is still burning.\nSince there scores are pretty near each other, change in leaderboard is inevitable though.",
      "votes": 1,
      "replies": [
        {
          "id": 1045656,
          "postDate": "2020-10-10T21:24:30.427Z",
          "content": "<p>I tend to agree that we will see some changes in the leaderboard soon. Perhaps now that all of the secrecy is over people will be able to truly cooperate and share their insights.</p>",
          "rawMarkdown": "I tend to agree that we will see some changes in the leaderboard soon. Perhaps now that all of the secrecy is over people will be able to truly cooperate and share their insights."
        }
      ]
    },
    {
      "id": 1044701,
      "postDate": "2020-10-10T04:06:00.163Z",
      "content": "<p>I totally agree.</p>\n<p>My score was -6.9, but all I used was… linear regression. That's it. The confidence interval was fixed and the prediction was the same for every patient. The model I created was completely unsuitable for medical use, but still did better than half of all other models. EDIT: apparently constant-value predictions can do even BETTER than that! Proves my point further.</p>\n<p>This being my first contest, I was merely trying to get my feet wet and, like you, submit a functional model. What I didn't expect (and admittedly disappointed me a bit) was for it to be merely 0.07 points worse than the best model. What that means is unclear because like you said, Laplace Log likelihood loss is not intuitive - it would've been nice if we had a human expert's baseline to compare it to.</p>\n<p>I applaud the winners and their hard work and determination. But the lack of one standout model makes me me wonder if the creators and participants of this competition might feel a bit disappointed. Frankly, it seems unlikely that any of these solutions will have a real impact on Pulm. Fibrosis regression, whether to inspire new ways of predicting the disease's progression or in actual field use. That's partly due to the raw performance I talked about and the models' black-box nature.</p>\n<p>This competition demonstrates perfectly that <strong>the cutthroat nature of Kaggle competitions - to earn the best possible score by whatever means - often overshadows the real goal of making a positive impact on peoples' lives by building an transparent model that can be trusted and used by real people</strong>. This is clear from the fact that, like you alluded to, one of the most valuable pieces of information available - the chest CT scan - was discarded or under-utilized in many of the best models because it would perform worse than a tabular-only approach. I don't think a real doctor would ever dismiss a patient's CT scan when evaluating them. You could argue that if bringing in the CT scans don't help the model, then you don't have any other choice, or that there weren't many CT scans available. But I ask: <strong>are we here to earn extra thousandths of a point on a leaderboard using a model that will never see the light of day, or to build a truly insightful model that will expand our current knowledge, even if by a hair?</strong></p>",
      "rawMarkdown": "I totally agree.\n\nMy score was -6.9, but all I used was... linear regression. That's it. The confidence interval was fixed and the prediction was the same for every patient. The model I created was completely unsuitable for medical use, but still did better than half of all other models. EDIT: apparently constant-value predictions can do even BETTER than that! Proves my point further.\n\nThis being my first contest, I was merely trying to get my feet wet and, like you, submit a functional model. What I didn't expect (and admittedly disappointed me a bit) was for it to be merely 0.07 points worse than the best model. What that means is unclear because like you said, Laplace Log likelihood loss is not intuitive - it would've been nice if we had a human expert's baseline to compare it to.\n\nI applaud the winners and their hard work and determination. But the lack of one standout model makes me me wonder if the creators and participants of this competition might feel a bit disappointed. Frankly, it seems unlikely that any of these solutions will have a real impact on Pulm. Fibrosis regression, whether to inspire new ways of predicting the disease's progression or in actual field use. That's partly due to the raw performance I talked about and the models' black-box nature.\n\nThis competition demonstrates perfectly that **the cutthroat nature of Kaggle competitions - to earn the best possible score by whatever means - often overshadows the real goal of making a positive impact on peoples' lives by building an transparent model that can be trusted and used by real people**. This is clear from the fact that, like you alluded to, one of the most valuable pieces of information available - the chest CT scan - was discarded or under-utilized in many of the best models because it would perform worse than a tabular-only approach. I don't think a real doctor would ever dismiss a patient's CT scan when evaluating them. You could argue that if bringing in the CT scans don't help the model, then you don't have any other choice, or that there weren't many CT scans available. But I ask: **are we here to earn extra thousandths of a point on a leaderboard using a model that will never see the light of day, or to build a truly insightful model that will expand our current knowledge, even if by a hair?**\n",
      "votes": 2,
      "replies": [
        {
          "id": 1045655,
          "postDate": "2020-10-10T21:22:55.243Z",
          "content": "<p>Very well written, I feel the same way.</p>",
          "rawMarkdown": "Very well written, I feel the same way."
        }
      ]
    },
    {
      "id": 1044807,
      "postDate": "2020-10-10T06:19:13.537Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1045658,
      "author_name": "Shoval T",
      "author_url": "",
      "post_date": "2020-10-10T21:31:29.420000",
      "content": "<p>I created a notebook with some insights regarding the winning model and the LLL score in general: <a href=\"https://www.kaggle.com/shovalt/laplace-log-likelihood-review\" target=\"_blank\">https://www.kaggle.com/shovalt/laplace-log-likelihood-review</a></p>\n<p>What I did was use the training data to evaluate several scenarios. I would love to get your feedback.</p>\n<p>Here is the final graph of the notebook, as well as the summary. There are several more leading to it in the notebook itself:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5539072%2F81176f08b0bfd390aa41bd432a100d75%2FLLL_final_fig.png?generation=1602365355875649&amp;alt=media\" alt=\"\"></p>\n<p>To summarize:</p>\n<ol>\n<li>The CT images had little-to-no effect on the top-scoring prediction models in this contest.</li>\n<li>The use of linear decay as a prediction target was not useful, and <em>even a perfect prediction</em> of that target wouldn't have been very beneficial. </li>\n<li>A direct prediction of FVC could possibly yield much better results, which is probably why the basic tabular-data models succeeded pretty well, especially when combined with a decent estimation of sigma using the quantiles method. </li>\n<li>A perfect prediction in this competition would probably score around LLL=-4.5, which is 1-2 orders of magnitude higher than the range of scores in the top 1000 scoring kernels (between -6.90 and -6.83) - <strong>there is potentially much room to grow here!</strong></li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1045142,
      "author_name": "Long Luu",
      "author_url": "",
      "post_date": "2020-10-10T11:29:54.623000",
      "content": "<p>Personally I think we cannot determine the metrics: there are many ways to decide how \"good\" a model is. For example, in normal image recognition we can use Accuracy, but in Medicine they use F1, or Specificity, since false negative should be minimized (you should not predict a patient is not ill if they are ill since that could cost a life). I think the organizers also try to pick the best one.\nAbout the difference in score, although maybe tiny, we only have one way to know: replace it with real numbers. Lowest possible score does not have limit (yet), so we should try more ways. The competition is ended, but the ideas to build better model is still burning.\nSince there scores are pretty near each other, change in leaderboard is inevitable though.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1045656,
          "author_name": "Shoval T",
          "author_url": "",
          "post_date": "2020-10-10T21:24:30.427000",
          "content": "<p>I tend to agree that we will see some changes in the leaderboard soon. Perhaps now that all of the secrecy is over people will be able to truly cooperate and share their insights.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044701,
      "author_name": "Henry Zhang",
      "author_url": "",
      "post_date": "2020-10-10T04:06:00.163000",
      "content": "<p>I totally agree.</p>\n<p>My score was -6.9, but all I used was… linear regression. That's it. The confidence interval was fixed and the prediction was the same for every patient. The model I created was completely unsuitable for medical use, but still did better than half of all other models. EDIT: apparently constant-value predictions can do even BETTER than that! Proves my point further.</p>\n<p>This being my first contest, I was merely trying to get my feet wet and, like you, submit a functional model. What I didn't expect (and admittedly disappointed me a bit) was for it to be merely 0.07 points worse than the best model. What that means is unclear because like you said, Laplace Log likelihood loss is not intuitive - it would've been nice if we had a human expert's baseline to compare it to.</p>\n<p>I applaud the winners and their hard work and determination. But the lack of one standout model makes me me wonder if the creators and participants of this competition might feel a bit disappointed. Frankly, it seems unlikely that any of these solutions will have a real impact on Pulm. Fibrosis regression, whether to inspire new ways of predicting the disease's progression or in actual field use. That's partly due to the raw performance I talked about and the models' black-box nature.</p>\n<p>This competition demonstrates perfectly that <strong>the cutthroat nature of Kaggle competitions - to earn the best possible score by whatever means - often overshadows the real goal of making a positive impact on peoples' lives by building an transparent model that can be trusted and used by real people</strong>. This is clear from the fact that, like you alluded to, one of the most valuable pieces of information available - the chest CT scan - was discarded or under-utilized in many of the best models because it would perform worse than a tabular-only approach. I don't think a real doctor would ever dismiss a patient's CT scan when evaluating them. You could argue that if bringing in the CT scans don't help the model, then you don't have any other choice, or that there weren't many CT scans available. But I ask: <strong>are we here to earn extra thousandths of a point on a leaderboard using a model that will never see the light of day, or to build a truly insightful model that will expand our current knowledge, even if by a hair?</strong></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1045655,
          "author_name": "Shoval T",
          "author_url": "",
          "post_date": "2020-10-10T21:22:55.243000",
          "content": "<p>Very well written, I feel the same way.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044807,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-10T06:19:13.537000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1044351": "This was my first Kaggle competition, and boy did I choose a hard one to start with. I didn't have nearly as much time as I wanted, but that's not the point of this thread.\n\nWhat's been bothering me is this - **was this competition a success?** \n\nThe first model I submitted was a very simple Ridge Regression model, based only on the data in `train.csv`. The confidence was determined by multiplying the predicted FVC by a constant (although optimized) fraction. I created the model just to see that I could submit successfully. It scored -6.9248 on the private leaderboard. The winning model scored -6.8305. In terms of position on the board itself this is a huge difference, but in real life, in terms of successful prediction, how much better is it really? \n\nNow, I don't know how others feel about it, but I am nowhere near to feeling intuitive around the Laplace Log Likelihood score they chose for this competition. Is a 0.1 difference really big? Did we even, as community, manage to harness the data in the thousands of images in any significant way? Will someone, some day be able to score a -5.0 on this competition? Perhaps even a -6.5?\n\nThese are just some thoughts that have been running around my head in the last couple of days. **I'd love to hear others' thoughts on this.**",
    "1045658": "I created a notebook with some insights regarding the winning model and the LLL score in general: https://www.kaggle.com/shovalt/laplace-log-likelihood-review\n\nWhat I did was use the training data to evaluate several scenarios. I would love to get your feedback.\n\nHere is the final graph of the notebook, as well as the summary. There are several more leading to it in the notebook itself:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5539072%2F81176f08b0bfd390aa41bd432a100d75%2FLLL_final_fig.png?generation=1602365355875649&alt=media)\n\nTo summarize:\n1. The CT images had little-to-no effect on the top-scoring prediction models in this contest.\n1. The use of linear decay as a prediction target was not useful, and *even a perfect prediction* of that target wouldn't have been very beneficial. \n1. A direct prediction of FVC could possibly yield much better results, which is probably why the basic tabular-data models succeeded pretty well, especially when combined with a decent estimation of sigma using the quantiles method. \n1. A perfect prediction in this competition would probably score around LLL=-4.5, which is 1-2 orders of magnitude higher than the range of scores in the top 1000 scoring kernels (between -6.90 and -6.83) - **there is potentially much room to grow here!**\n",
    "1045142": "Personally I think we cannot determine the metrics: there are many ways to decide how \"good\" a model is. For example, in normal image recognition we can use Accuracy, but in Medicine they use F1, or Specificity, since false negative should be minimized (you should not predict a patient is not ill if they are ill since that could cost a life). I think the organizers also try to pick the best one.\nAbout the difference in score, although maybe tiny, we only have one way to know: replace it with real numbers. Lowest possible score does not have limit (yet), so we should try more ways. The competition is ended, but the ideas to build better model is still burning.\nSince there scores are pretty near each other, change in leaderboard is inevitable though.",
    "1044701": "I totally agree.\n\nMy score was -6.9, but all I used was... linear regression. That's it. The confidence interval was fixed and the prediction was the same for every patient. The model I created was completely unsuitable for medical use, but still did better than half of all other models. EDIT: apparently constant-value predictions can do even BETTER than that! Proves my point further.\n\nThis being my first contest, I was merely trying to get my feet wet and, like you, submit a functional model. What I didn't expect (and admittedly disappointed me a bit) was for it to be merely 0.07 points worse than the best model. What that means is unclear because like you said, Laplace Log likelihood loss is not intuitive - it would've been nice if we had a human expert's baseline to compare it to.\n\nI applaud the winners and their hard work and determination. But the lack of one standout model makes me me wonder if the creators and participants of this competition might feel a bit disappointed. Frankly, it seems unlikely that any of these solutions will have a real impact on Pulm. Fibrosis regression, whether to inspire new ways of predicting the disease's progression or in actual field use. That's partly due to the raw performance I talked about and the models' black-box nature.\n\nThis competition demonstrates perfectly that **the cutthroat nature of Kaggle competitions - to earn the best possible score by whatever means - often overshadows the real goal of making a positive impact on peoples' lives by building an transparent model that can be trusted and used by real people**. This is clear from the fact that, like you alluded to, one of the most valuable pieces of information available - the chest CT scan - was discarded or under-utilized in many of the best models because it would perform worse than a tabular-only approach. I don't think a real doctor would ever dismiss a patient's CT scan when evaluating them. You could argue that if bringing in the CT scans don't help the model, then you don't have any other choice, or that there weren't many CT scans available. But I ask: **are we here to earn extra thousandths of a point on a leaderboard using a model that will never see the light of day, or to build a truly insightful model that will expand our current knowledge, even if by a hair?**\n",
    "1044807": ""
  }
}