{
  "id": 227785,
  "title": "Questions about the eval metrics on self-made val dataset",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/227785",
  "author_name": "Nin7a1",
  "post_date": "2021-03-22T08:46:44.258000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I follow the approach of Darek Kłeczek on <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550</a>. And I split the cropped cells into training and val dataset with a ratio of 4:1 just like the approach does. First, when I train a resnet50 on training dataset and eval on val dataset with <code>average_precision_score</code>, I get average_precision_score=0.5+, and the result on leader board is 0.32. When I train the resnet50 with more epochs with a decay learning rate scheduler, I get a higher average_precision_score=0.8+ on val dataset. Howerver, the result on leader board becomes worse. (0.29). I dont know whether the metrics like average_precision_score is unreliable or the splitting method causes some leakage. Do you have any idea? Thanks a lot. </p>",
  "messages": [
    {
      "id": 1248149,
      "postDate": "2021-03-22T11:40:31.867Z",
      "content": "<p>I still cannot find reliable val metric that can correlate with LB. Will be glad to know if someone found it.</p>",
      "rawMarkdown": "I still cannot find reliable val metric that can correlate with LB. Will be glad to know if someone found it.",
      "votes": 5
    },
    {
      "id": 1247994,
      "postDate": "2021-03-22T08:46:44.260Z",
      "content": "<p>I follow the approach of Darek Kłeczek on <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550</a>. And I split the cropped cells into training and val dataset with a ratio of 4:1 just like the approach does. First, when I train a resnet50 on training dataset and eval on val dataset with <code>average_precision_score</code>, I get average_precision_score=0.5+, and the result on leader board is 0.32. When I train the resnet50 with more epochs with a decay learning rate scheduler, I get a higher average_precision_score=0.8+ on val dataset. Howerver, the result on leader board becomes worse. (0.29). I dont know whether the metrics like average_precision_score is unreliable or the splitting method causes some leakage. Do you have any idea? Thanks a lot. </p>",
      "rawMarkdown": "I follow the approach of Darek Kłeczek on https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550. And I split the cropped cells into training and val dataset with a ratio of 4:1 just like the approach does. First, when I train a resnet50 on training dataset and eval on val dataset with `average_precision_score`, I get average_precision_score=0.5+, and the result on leader board is 0.32. When I train the resnet50 with more epochs with a decay learning rate scheduler, I get a higher average_precision_score=0.8+ on val dataset. Howerver, the result on leader board becomes worse. (0.29). I dont know whether the metrics like average_precision_score is unreliable or the splitting method causes some leakage. Do you have any idea? Thanks a lot. ",
      "votes": 5
    },
    {
      "id": 1258562,
      "postDate": "2021-03-31T17:56:37.977Z",
      "content": "<p>Isn't this awesome, my number of trained epochs correlate negatively with the Public LB -&gt; easy metrics :)</p>\n<p><img src=\"https://i.ibb.co/wsZx3N5/Unbenannt.jpg\" alt=\"\"></p>",
      "rawMarkdown": "Isn't this awesome, my number of trained epochs correlate negatively with the Public LB -> easy metrics :)\n\n![](https://i.ibb.co/wsZx3N5/Unbenannt.jpg)",
      "votes": 3,
      "replies": [
        {
          "id": 1258590,
          "postDate": "2021-03-31T18:29:10.803Z",
          "content": "<p>Where is 0.428 on the chart? ;) </p>",
          "rawMarkdown": "Where is 0.428 on the chart? ;) "
        },
        {
          "id": 1258595,
          "postDate": "2021-03-31T18:31:07.593Z",
          "content": "<p>Thats ensembling :)</p>",
          "rawMarkdown": "Thats ensembling :)",
          "votes": 2
        },
        {
          "id": 1258621,
          "postDate": "2021-03-31T18:53:52.647Z",
          "content": "<p>Ensemble works best for me as well</p>",
          "rawMarkdown": "Ensemble works best for me as well"
        },
        {
          "id": 1258980,
          "postDate": "2021-04-01T04:21:40.230Z",
          "content": "<p>My experiments also show negative correlation…..</p>",
          "rawMarkdown": "My experiments also show negative correlation....."
        },
        {
          "id": 1259160,
          "postDate": "2021-04-01T07:24:51.793Z",
          "content": "<p>how do you explain the negative correlation?</p>",
          "rawMarkdown": "how do you explain the negative correlation?"
        }
      ]
    },
    {
      "id": 1257947,
      "postDate": "2021-03-31T08:21:17.380Z",
      "content": "<p>Its like try and error. I guess because the training set is weakly labeled. </p>",
      "rawMarkdown": "Its like try and error. I guess because the training set is weakly labeled. ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1248149,
      "author_name": "Vladislav Ostankovich",
      "author_url": "",
      "post_date": "2021-03-22T11:40:31.867000",
      "content": "<p>I still cannot find reliable val metric that can correlate with LB. Will be glad to know if someone found it.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1258562,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-03-31T17:56:37.977000",
      "content": "<p>Isn't this awesome, my number of trained epochs correlate negatively with the Public LB -&gt; easy metrics :)</p>\n<p><img src=\"https://i.ibb.co/wsZx3N5/Unbenannt.jpg\" alt=\"\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1258590,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-31T18:29:10.803000",
          "content": "<p>Where is 0.428 on the chart? ;) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258595,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-03-31T18:31:07.593000",
          "content": "<p>Thats ensembling :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1258621,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-03-31T18:53:52.647000",
          "content": "<p>Ensemble works best for me as well</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1258980,
          "author_name": "Nin7a1",
          "author_url": "",
          "post_date": "2021-04-01T04:21:40.230000",
          "content": "<p>My experiments also show negative correlation…..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1259160,
          "author_name": "LucaMTB",
          "author_url": "",
          "post_date": "2021-04-01T07:24:51.793000",
          "content": "<p>how do you explain the negative correlation?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1257947,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2021-03-31T08:21:17.380000",
      "content": "<p>Its like try and error. I guess because the training set is weakly labeled. </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1248149": "I still cannot find reliable val metric that can correlate with LB. Will be glad to know if someone found it.",
    "1247994": "I follow the approach of Darek Kłeczek on https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/221550. And I split the cropped cells into training and val dataset with a ratio of 4:1 just like the approach does. First, when I train a resnet50 on training dataset and eval on val dataset with `average_precision_score`, I get average_precision_score=0.5+, and the result on leader board is 0.32. When I train the resnet50 with more epochs with a decay learning rate scheduler, I get a higher average_precision_score=0.8+ on val dataset. Howerver, the result on leader board becomes worse. (0.29). I dont know whether the metrics like average_precision_score is unreliable or the splitting method causes some leakage. Do you have any idea? Thanks a lot. ",
    "1258562": "Isn't this awesome, my number of trained epochs correlate negatively with the Public LB -> easy metrics :)\n\n![](https://i.ibb.co/wsZx3N5/Unbenannt.jpg)",
    "1257947": "Its like try and error. I guess because the training set is weakly labeled. "
  }
}