{
  "id": 107202,
  "title": "How to use K-fold in this competitions, and K-fold will really boost the LB Score?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/107202",
  "author_name": "",
  "post_date": "2019-09-03T01:43:57.433434400Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi, guys. I want to use K-fold ,but I don't know how to use it in keras ,  please give me a kernel to refer if someone have (or on github). The last three days to try? please.</p>",
  "messages": [
    {
      "id": "616309",
      "postDate": "09/03/2019 01:43:57",
      "content": "<p>Hi, guys. I want to use K-fold ,but I don't know how to use it in keras ,  please give me a kernel to refer if someone have (or on github). The last three days to try? please.</p>",
      "rawMarkdown": "Hi, guys. I want to use K-fold ,but I don't know how to use it in keras ,  please give me a kernel to refer if someone have (or on github). The last three days to try? please.",
      "votes": null
    },
    {
      "id": "616508",
      "postDate": "09/03/2019 07:35:08",
      "content": "<p>I use kernel from other competition, to make cross validation in keras.\n<a href=\"https://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation\">https://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation</a></p>\n\n<p>I cant answer on question about boost LB score: in my case local kappa was ~0.81, but submit average of cross validation model predictions give 0. :( Models overfit and  I don`t know what i doing wrong.</p>",
      "rawMarkdown": "I use kernel from other competition, to make cross validation in keras.\nhttps://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation\n\nI cant answer on question about boost LB score: in my case local kappa was ~0.81, but submit average of cross validation model predictions give 0. :( Models overfit and  I don`t know what i doing wrong.",
      "votes": null
    },
    {
      "id": "617325",
      "postDate": "09/04/2019 02:27:18",
      "content": "<p>Thank you a lot, I will try ,you can use ensemble,it's really useful.</p>",
      "rawMarkdown": "Thank you a lot, I will try ,you can use ensemble,it's really useful.",
      "votes": null
    },
    {
      "id": "617423",
      "postDate": "09/04/2019 05:43:43",
      "content": "<p>You can use K folds in two options : \n1. Train K different models (usually  5 or 10, but due to the time limit use 5 ) and select the model which has the best CV <br>\n2. You can take all the models and ensemble them  </p>\n\n<p>Here is a simple example to create 5 folds : </p>\n\n<p>```\nfrom sklearn.model_selection import ShuffleSplit\nimage_ids = pd.read_csv(\"../input/train.csv\")\ncv = ShuffleSplit(n_splits=5, test_size=0.2, random_state=2026)\ni = 0\nfor train_index, test_index in cv.split(image_ids):\n    vars()[\"image_ids_train\"+str(i)] = image_ids.iloc[train_index]\n    vars()[\"image_ids_valid\"+str(i)] = image_ids.iloc[test_index]\n    vars()[\"image_ids_train\"+str(i)].reset_index(inplace=True,drop = True)\n    vars()[\"image_ids_valid\"+str(i)].reset_index(inplace=True,drop=True)\n    i +=1</p>\n\n<p>```</p>",
      "rawMarkdown": "You can use K folds in two options : \n1. Train K different models (usually  5 or 10, but due to the time limit use 5 ) and select the model which has the best CV  \n2. You can take all the models and ensemble them  \n\nHere is a simple example to create 5 folds : \n\n```\nfrom sklearn.model_selection import ShuffleSplit\nimage_ids = pd.read_csv(\"../input/train.csv\")\ncv = ShuffleSplit(n_splits=5, test_size=0.2, random_state=2026)\ni = 0\nfor train_index, test_index in cv.split(image_ids):\n    vars()[\"image_ids_train\"+str(i)] = image_ids.iloc[train_index]\n    vars()[\"image_ids_valid\"+str(i)] = image_ids.iloc[test_index]\n    vars()[\"image_ids_train\"+str(i)].reset_index(inplace=True,drop = True)\n    vars()[\"image_ids_valid\"+str(i)].reset_index(inplace=True,drop=True)\n    i +=1\n\n```",
      "votes": null
    },
    {
      "id": "617469",
      "postDate": "09/04/2019 06:55:25",
      "content": "<p>Nice code,Thank you so much@OmerS</p>",
      "rawMarkdown": "Nice code,Thank you so much@OmerS",
      "votes": null
    },
    {
      "id": "617629",
      "postDate": "09/04/2019 10:25:15",
      "content": "<p>I tried this method and it give me 0 :( I think models were overfit, but I don`t understand why.</p>",
      "rawMarkdown": "I tried this method and it give me 0 :( I think models were overfit, but I don`t understand why.",
      "votes": null
    },
    {
      "id": "617708",
      "postDate": "09/04/2019 12:02:21",
      "content": "<p>Maybe you messed up the indices somewhere? For instance, accidentally used the same subset for training and validation, so that the validation score didn't show overfitting.</p>",
      "rawMarkdown": "Maybe you messed up the indices somewhere? For instance, accidentally used the same subset for training and validation, so that the validation score didn't show overfitting.",
      "votes": null
    },
    {
      "id": "617731",
      "postDate": "09/04/2019 12:24:02",
      "content": "<p>No, i doesn't :(</p>",
      "rawMarkdown": "No, i doesn't :(",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 616508,
      "author_name": "yanahontarenko",
      "author_url": "",
      "post_date": "09/03/2019 07:35:08",
      "content": "<p>I use kernel from other competition, to make cross validation in keras.\n<a href=\"https://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation\">https://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation</a></p>\n\n<p>I cant answer on question about boost LB score: in my case local kappa was ~0.81, but submit average of cross validation model predictions give 0. :( Models overfit and  I don`t know what i doing wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 617325,
          "author_name": "chopinforest1986",
          "author_url": "",
          "post_date": "09/04/2019 02:27:18",
          "content": "<p>Thank you a lot, I will try ,you can use ensemble,it's really useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617629,
          "author_name": "yanahontarenko",
          "author_url": "",
          "post_date": "09/04/2019 10:25:15",
          "content": "<p>I tried this method and it give me 0 :( I think models were overfit, but I don`t understand why.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617708,
          "author_name": "blackitten13",
          "author_url": "",
          "post_date": "09/04/2019 12:02:21",
          "content": "<p>Maybe you messed up the indices somewhere? For instance, accidentally used the same subset for training and validation, so that the validation score didn't show overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617731,
          "author_name": "yanahontarenko",
          "author_url": "",
          "post_date": "09/04/2019 12:24:02",
          "content": "<p>No, i doesn't :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 617423,
      "author_name": "omershect",
      "author_url": "",
      "post_date": "09/04/2019 05:43:43",
      "content": "<p>You can use K folds in two options : \n1. Train K different models (usually  5 or 10, but due to the time limit use 5 ) and select the model which has the best CV <br>\n2. You can take all the models and ensemble them  </p>\n\n<p>Here is a simple example to create 5 folds : </p>\n\n<p>```\nfrom sklearn.model_selection import ShuffleSplit\nimage_ids = pd.read_csv(\"../input/train.csv\")\ncv = ShuffleSplit(n_splits=5, test_size=0.2, random_state=2026)\ni = 0\nfor train_index, test_index in cv.split(image_ids):\n    vars()[\"image_ids_train\"+str(i)] = image_ids.iloc[train_index]\n    vars()[\"image_ids_valid\"+str(i)] = image_ids.iloc[test_index]\n    vars()[\"image_ids_train\"+str(i)].reset_index(inplace=True,drop = True)\n    vars()[\"image_ids_valid\"+str(i)].reset_index(inplace=True,drop=True)\n    i +=1</p>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 617469,
          "author_name": "chopinforest1986",
          "author_url": "",
          "post_date": "09/04/2019 06:55:25",
          "content": "<p>Nice code,Thank you so much@OmerS</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "616309": "Hi, guys. I want to use K-fold ,but I don't know how to use it in keras ,  please give me a kernel to refer if someone have (or on github). The last three days to try? please.",
    "616508": "I use kernel from other competition, to make cross validation in keras.\nhttps://www.kaggle.com/stefanie04736/simple-keras-model-with-k-fold-cross-validation\n\nI cant answer on question about boost LB score: in my case local kappa was ~0.81, but submit average of cross validation model predictions give 0. :( Models overfit and  I don`t know what i doing wrong.",
    "617325": "Thank you a lot, I will try ,you can use ensemble,it's really useful.",
    "617423": "You can use K folds in two options : \n1. Train K different models (usually  5 or 10, but due to the time limit use 5 ) and select the model which has the best CV  \n2. You can take all the models and ensemble them  \n\nHere is a simple example to create 5 folds : \n\n```\nfrom sklearn.model_selection import ShuffleSplit\nimage_ids = pd.read_csv(\"../input/train.csv\")\ncv = ShuffleSplit(n_splits=5, test_size=0.2, random_state=2026)\ni = 0\nfor train_index, test_index in cv.split(image_ids):\n    vars()[\"image_ids_train\"+str(i)] = image_ids.iloc[train_index]\n    vars()[\"image_ids_valid\"+str(i)] = image_ids.iloc[test_index]\n    vars()[\"image_ids_train\"+str(i)].reset_index(inplace=True,drop = True)\n    vars()[\"image_ids_valid\"+str(i)].reset_index(inplace=True,drop=True)\n    i +=1\n\n```",
    "617469": "Nice code,Thank you so much@OmerS",
    "617629": "I tried this method and it give me 0 :( I think models were overfit, but I don`t understand why.",
    "617708": "Maybe you messed up the indices somewhere? For instance, accidentally used the same subset for training and validation, so that the validation score didn't show overfitting.",
    "617731": "No, i doesn't :("
  },
  "source": "meta"
}