{
  "id": 198679,
  "title": "More digits in the LB",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/198679",
  "author_name": "",
  "post_date": "2020-11-22T12:20:50.760481200Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I see so many people have the same score. If the host can got more digits, we better know how more improved. 4 digits is quite good.</p>",
  "messages": [
    {
      "id": "1087160",
      "postDate": "11/22/2020 12:20:50",
      "content": "<p>I see so many people have the same score. If the host can got more digits, we better know how more improved. 4 digits is quite good.</p>",
      "rawMarkdown": "I see so many people have the same score. If the host can got more digits, we better know how more improved. 4 digits is quite good.",
      "votes": null
    },
    {
      "id": "1087414",
      "postDate": "11/22/2020 17:29:00",
      "content": "<p>Short answer: I don't agree. Trust your CV. </p>\n<p>Long answer: I was holding 1st place for most of the time in the last cassava competition. I didn't choose my models with best CV because they had very bad public LB. I learned the lesson and finished 2nd. Winning solution was based on CV.</p>",
      "rawMarkdown": "Short answer: I don't agree. Trust your CV. \n\nLong answer: I was holding 1st place for most of the time in the last cassava competition. I didn't choose my models with best CV because they had very bad public LB. I learned the lesson and finished 2nd. Winning solution was based on CV.",
      "votes": null
    },
    {
      "id": "1089566",
      "postDate": "11/24/2020 15:55:58",
      "content": "<p>Hello I'm  new here. And I want to know how to look at my own CV</p>",
      "rawMarkdown": "Hello I'm  new here. And I want to know how to look at my own CV",
      "votes": null
    },
    {
      "id": "1124440",
      "postDate": "12/23/2020 23:58:26",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/whutddmm\" target=\"_blank\">@whutddmm</a>, Incase you're still stuck, you could look at bulit-in functions from scikit-learn <a href=\"https://scikit-learn.org/stable/modules/cross_validation.html\" target=\"_blank\">cv options</a>. Basically, you split your dataset into n parts and train your model on n-1 parts while testing on the nth part. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fd24a0857eb6e262e0d6db36d59e5bc24%2Fgrid_search_cross_validation.png?generation=1608767792018337&amp;alt=media\" alt=\"k-fold cv\"></p>\n<p>By doing this for different combinations of train and test sets, you get an idea of your model performance. You can take the mean of the metrics from different folds as reference cv score</p>",
      "rawMarkdown": "Hi @whutddmm, Incase you're still stuck, you could look at bulit-in functions from scikit-learn [cv options](https://scikit-learn.org/stable/modules/cross_validation.html). Basically, you split your dataset into n parts and train your model on n-1 parts while testing on the nth part. \n\n ![k-fold cv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fd24a0857eb6e262e0d6db36d59e5bc24%2Fgrid_search_cross_validation.png?generation=1608767792018337&alt=media)\n\n\nBy doing this for different combinations of train and test sets, you get an idea of your model performance. You can take the mean of the metrics from different folds as reference cv score",
      "votes": null
    },
    {
      "id": "1124899",
      "postDate": "12/24/2020 09:05:07",
      "content": "<p>At some point Kaggle will probably show us another digit - but I 100% echo what Miroslav Valan stated.  </p>\n<p>Also - if your working to achieve a 0.0001 size improvement you will almost always end up with an over-fitting model.  </p>\n<p>But you probably will not take that advice (heck I probably will not take my own advice).  </p>\n<p>Kaggle lists your results with all the digits they possess.     If your the very first 0.901 than you score likely 0.90199.   If your the very last than your are likely 0.90101.   So you can tell your results compared to all the others with the same displayed value.</p>\n<p>As you make new submissions -<br>\nYou can display your submission results by selecting the Public Score option -  that will reorder your submissions from best to worse - if you have 50 submissions at 0.901 than the very first one is the better of the bunch.</p>\n<p>So with an extra mouse click or two you can infer the digits!</p>",
      "rawMarkdown": "At some point Kaggle will probably show us another digit - but I 100% echo what Miroslav Valan stated.  \n\nAlso - if your working to achieve a 0.0001 size improvement you will almost always end up with an over-fitting model.  \n\nBut you probably will not take that advice (heck I probably will not take my own advice).  \n\nKaggle lists your results with all the digits they possess.     If your the very first 0.901 than you score likely 0.90199.   If your the very last than your are likely 0.90101.   So you can tell your results compared to all the others with the same displayed value.\n\nAs you make new submissions -\nYou can display your submission results by selecting the Public Score option -  that will reorder your submissions from best to worse - if you have 50 submissions at 0.901 than the very first one is the better of the bunch.\n\nSo with an extra mouse click or two you can infer the digits!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1087414,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "11/22/2020 17:29:00",
      "content": "<p>Short answer: I don't agree. Trust your CV. </p>\n<p>Long answer: I was holding 1st place for most of the time in the last cassava competition. I didn't choose my models with best CV because they had very bad public LB. I learned the lesson and finished 2nd. Winning solution was based on CV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1089566,
          "author_name": "whutddmm",
          "author_url": "",
          "post_date": "11/24/2020 15:55:58",
          "content": "<p>Hello I'm  new here. And I want to know how to look at my own CV</p>",
          "votes": null,
          "replies": [
            {
              "id": 1124440,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "12/23/2020 23:58:26",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/whutddmm\" target=\"_blank\">@whutddmm</a>, Incase you're still stuck, you could look at bulit-in functions from scikit-learn <a href=\"https://scikit-learn.org/stable/modules/cross_validation.html\" target=\"_blank\">cv options</a>. Basically, you split your dataset into n parts and train your model on n-1 parts while testing on the nth part. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fd24a0857eb6e262e0d6db36d59e5bc24%2Fgrid_search_cross_validation.png?generation=1608767792018337&amp;alt=media\" alt=\"k-fold cv\"></p>\n<p>By doing this for different combinations of train and test sets, you get an idea of your model performance. You can take the mean of the metrics from different folds as reference cv score</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1124899,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/24/2020 09:05:07",
      "content": "<p>At some point Kaggle will probably show us another digit - but I 100% echo what Miroslav Valan stated.  </p>\n<p>Also - if your working to achieve a 0.0001 size improvement you will almost always end up with an over-fitting model.  </p>\n<p>But you probably will not take that advice (heck I probably will not take my own advice).  </p>\n<p>Kaggle lists your results with all the digits they possess.     If your the very first 0.901 than you score likely 0.90199.   If your the very last than your are likely 0.90101.   So you can tell your results compared to all the others with the same displayed value.</p>\n<p>As you make new submissions -<br>\nYou can display your submission results by selecting the Public Score option -  that will reorder your submissions from best to worse - if you have 50 submissions at 0.901 than the very first one is the better of the bunch.</p>\n<p>So with an extra mouse click or two you can infer the digits!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1087160": "I see so many people have the same score. If the host can got more digits, we better know how more improved. 4 digits is quite good.",
    "1087414": "Short answer: I don't agree. Trust your CV. \n\nLong answer: I was holding 1st place for most of the time in the last cassava competition. I didn't choose my models with best CV because they had very bad public LB. I learned the lesson and finished 2nd. Winning solution was based on CV.",
    "1089566": "Hello I'm  new here. And I want to know how to look at my own CV",
    "1124440": "Hi @whutddmm, Incase you're still stuck, you could look at bulit-in functions from scikit-learn [cv options](https://scikit-learn.org/stable/modules/cross_validation.html). Basically, you split your dataset into n parts and train your model on n-1 parts while testing on the nth part. \n\n ![k-fold cv](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fd24a0857eb6e262e0d6db36d59e5bc24%2Fgrid_search_cross_validation.png?generation=1608767792018337&alt=media)\n\n\nBy doing this for different combinations of train and test sets, you get an idea of your model performance. You can take the mean of the metrics from different folds as reference cv score",
    "1124899": "At some point Kaggle will probably show us another digit - but I 100% echo what Miroslav Valan stated.  \n\nAlso - if your working to achieve a 0.0001 size improvement you will almost always end up with an over-fitting model.  \n\nBut you probably will not take that advice (heck I probably will not take my own advice).  \n\nKaggle lists your results with all the digits they possess.     If your the very first 0.901 than you score likely 0.90199.   If your the very last than your are likely 0.90101.   So you can tell your results compared to all the others with the same displayed value.\n\nAs you make new submissions -\nYou can display your submission results by selecting the Public Score option -  that will reorder your submissions from best to worse - if you have 50 submissions at 0.901 than the very first one is the better of the bunch.\n\nSo with an extra mouse click or two you can infer the digits!"
  },
  "source": "meta"
}