{
  "id": 244450,
  "title": "Why did my model got [CV:0.98, LB:0.96] but didn't detect anything?",
  "url": "/competitions/seti-breakthrough-listen/discussion/244450",
  "author_name": "Vasanth",
  "post_date": "2021-06-06T20:42:19.266000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I just made my first submission in this competition using the predictions of a single fold model that got a CV of 0.9798 and LB score of 0.96, since I already got to know about the CV vs LB score difference from the discussions, it wasn't a surprise to me. But, what surprised me was the fact that, my model didn't even detect a single positive case at a threshold of 0.5 in the test set, with the maximum probability of a positive class being less than 0.2 for all 35k+ samples, which is not the case with CV that had a recall of 0.91 at a threshold of 0.5.  I wonder if it's caused by the difference in data distributions or the data pre-processing provided it's same for both the train and test sets. Does anyone have any thoughts on this?</p>",
  "messages": [
    {
      "id": 1338939,
      "postDate": "2021-06-06T21:50:17.897Z",
      "content": "<p>I guess that the threshold should be set at the target mean, not 0.5.</p>",
      "rawMarkdown": "I guess that the threshold should be set at the target mean, not 0.5.",
      "votes": 1
    },
    {
      "id": 1338894,
      "postDate": "2021-06-06T20:42:19.267Z",
      "content": "<p>I just made my first submission in this competition using the predictions of a single fold model that got a CV of 0.9798 and LB score of 0.96, since I already got to know about the CV vs LB score difference from the discussions, it wasn't a surprise to me. But, what surprised me was the fact that, my model didn't even detect a single positive case at a threshold of 0.5 in the test set, with the maximum probability of a positive class being less than 0.2 for all 35k+ samples, which is not the case with CV that had a recall of 0.91 at a threshold of 0.5.  I wonder if it's caused by the difference in data distributions or the data pre-processing provided it's same for both the train and test sets. Does anyone have any thoughts on this?</p>",
      "rawMarkdown": "I just made my first submission in this competition using the predictions of a single fold model that got a CV of 0.9798 and LB score of 0.96, since I already got to know about the CV vs LB score difference from the discussions, it wasn't a surprise to me. But, what surprised me was the fact that, my model didn't even detect a single positive case at a threshold of 0.5 in the test set, with the maximum probability of a positive class being less than 0.2 for all 35k+ samples, which is not the case with CV that had a recall of 0.91 at a threshold of 0.5.  I wonder if it's caused by the difference in data distributions or the data pre-processing provided it's same for both the train and test sets. Does anyone have any thoughts on this?",
      "votes": 1
    },
    {
      "id": 1338912,
      "postDate": "2021-06-06T21:07:39.513Z",
      "content": "<p>As the metric is rocauc, only the ranking matters. There is no need for hard labels.<br>\nThreshold 0.5 in local validation (for whatever purpose) is maybe a bit high for a dataset that is very skewed, i am not that surprised.</p>",
      "rawMarkdown": "As the metric is rocauc, only the ranking matters. There is no need for hard labels.\nThreshold 0.5 in local validation (for whatever purpose) is maybe a bit high for a dataset that is very skewed, i am not that surprised.",
      "votes": 2,
      "replies": [
        {
          "id": 1338923,
          "postDate": "2021-06-06T21:25:51.993Z",
          "content": "<p>It still is a good question, what could be causing this much of lower predictions for him.</p>",
          "rawMarkdown": "It still is a good question, what could be causing this much of lower predictions for him.",
          "votes": 1
        },
        {
          "id": 1338928,
          "postDate": "2021-06-06T21:35:34.890Z",
          "content": "<p>Indeed, the difference between test and train is quite interesting. I am not participating here (yet), so I don't know what could cause that. </p>",
          "rawMarkdown": "Indeed, the difference between test and train is quite interesting. I am not participating here (yet), so I don't know what could cause that. ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1338939,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-06-06T21:50:17.897000",
      "content": "<p>I guess that the threshold should be set at the target mean, not 0.5.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1338912,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2021-06-06T21:07:39.513000",
      "content": "<p>As the metric is rocauc, only the ranking matters. There is no need for hard labels.<br>\nThreshold 0.5 in local validation (for whatever purpose) is maybe a bit high for a dataset that is very skewed, i am not that surprised.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1338923,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-06T21:25:51.993000",
          "content": "<p>It still is a good question, what could be causing this much of lower predictions for him.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1338928,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2021-06-06T21:35:34.890000",
          "content": "<p>Indeed, the difference between test and train is quite interesting. I am not participating here (yet), so I don't know what could cause that. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1338939": "I guess that the threshold should be set at the target mean, not 0.5.",
    "1338894": "I just made my first submission in this competition using the predictions of a single fold model that got a CV of 0.9798 and LB score of 0.96, since I already got to know about the CV vs LB score difference from the discussions, it wasn't a surprise to me. But, what surprised me was the fact that, my model didn't even detect a single positive case at a threshold of 0.5 in the test set, with the maximum probability of a positive class being less than 0.2 for all 35k+ samples, which is not the case with CV that had a recall of 0.91 at a threshold of 0.5.  I wonder if it's caused by the difference in data distributions or the data pre-processing provided it's same for both the train and test sets. Does anyone have any thoughts on this?",
    "1338912": "As the metric is rocauc, only the ranking matters. There is no need for hard labels.\nThreshold 0.5 in local validation (for whatever purpose) is maybe a bit high for a dataset that is very skewed, i am not that surprised."
  }
}