{
  "id": 92858,
  "title": "[SOLVED] Cross validation failing",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/92858",
  "author_name": "",
  "post_date": "2019-05-21T05:12:02.969444300Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>[EDIT] - I wrote my inference code from scratch and now it is fixed. I still don't know where the bug was but at least I have something working! :P</p>\n\n<p>Hello all,</p>\n\n<p>I am trying to perform 5-fold cross-validation but it is failing miserably.</p>\n\n<p>I tried a single fold case, can got 0.645 on the public LB. I took the same code and trained 5 folds in one kernel. My OOF lwlrap was 0.81. In another kernel, I loaded the models, and performed inferences with all 5 models and averaged them. I got 0.327 on the Public LB. Here is some samples of my code:</p>\n\n<p>Training:</p>\n\n<p>```\nfrom sklearn.model_selection import KFold</p>\n\n<p>n_fold = 5\nseed = 42\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=seed)</p>\n\n<p>for fold_, (trn_idx, val_idx) in enumerate(folds.split(trn_curated_df['fname'],trn_curated_df['labels'])):\n    tfms = get_transforms(do_flip=True, max_rotate=0, max_lighting=0.1, max_zoom=0, max_warp=0.)\n    src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n    data = (src.transform(tfms, size=128)\n            .databunch(bs=128).normalize(imagenet_stats)\n    )\n    learn = cnn_learner(data, borrowed_model, pretrained=False, metrics=[lwlrap]).mixup(stack_y=False)\n    learn.unfreeze()\n    learn.fit_one_cycle(200, 2e-2,callbacks=[SaveModelCallback(learn, every='improvement', monitor='lwlrap', name='best')])\n    learn.export('FAT2019_Simple'+str(fold_)+'.pkl')\n```</p>\n\n<p>Inference:</p>\n\n<p>```</p>\n\n<p>predictions = torch.from_numpy(np.zeros((1120,80))).float()\nprint(predictions.shape)\nnum_folds = 5\nfor i in range(num_folds):\n    learn = load_learner(path = '.', file='../input/simple-model-fat2019/FAT2019_Simple'+str(i)+'.pkl',test = test)\n    preds, _ = learn.TTA(ds_type=DatasetType.Test,aug_tfms=tfms)\n    print(preds.shape)\n    predictions = predictions + preds\npredictions = predictions/num_folds\n```\nWhat could be wrong here?</p>",
  "messages": [
    {
      "id": "534342",
      "postDate": "05/21/2019 05:12:02",
      "content": "<p>[EDIT] - I wrote my inference code from scratch and now it is fixed. I still don't know where the bug was but at least I have something working! :P</p>\n\n<p>Hello all,</p>\n\n<p>I am trying to perform 5-fold cross-validation but it is failing miserably.</p>\n\n<p>I tried a single fold case, can got 0.645 on the public LB. I took the same code and trained 5 folds in one kernel. My OOF lwlrap was 0.81. In another kernel, I loaded the models, and performed inferences with all 5 models and averaged them. I got 0.327 on the Public LB. Here is some samples of my code:</p>\n\n<p>Training:</p>\n\n<p>```\nfrom sklearn.model_selection import KFold</p>\n\n<p>n_fold = 5\nseed = 42\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=seed)</p>\n\n<p>for fold_, (trn_idx, val_idx) in enumerate(folds.split(trn_curated_df['fname'],trn_curated_df['labels'])):\n    tfms = get_transforms(do_flip=True, max_rotate=0, max_lighting=0.1, max_zoom=0, max_warp=0.)\n    src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n    data = (src.transform(tfms, size=128)\n            .databunch(bs=128).normalize(imagenet_stats)\n    )\n    learn = cnn_learner(data, borrowed_model, pretrained=False, metrics=[lwlrap]).mixup(stack_y=False)\n    learn.unfreeze()\n    learn.fit_one_cycle(200, 2e-2,callbacks=[SaveModelCallback(learn, every='improvement', monitor='lwlrap', name='best')])\n    learn.export('FAT2019_Simple'+str(fold_)+'.pkl')\n```</p>\n\n<p>Inference:</p>\n\n<p>```</p>\n\n<p>predictions = torch.from_numpy(np.zeros((1120,80))).float()\nprint(predictions.shape)\nnum_folds = 5\nfor i in range(num_folds):\n    learn = load_learner(path = '.', file='../input/simple-model-fat2019/FAT2019_Simple'+str(i)+'.pkl',test = test)\n    preds, _ = learn.TTA(ds_type=DatasetType.Test,aug_tfms=tfms)\n    print(preds.shape)\n    predictions = predictions + preds\npredictions = predictions/num_folds\n```\nWhat could be wrong here?</p>",
      "rawMarkdown": "[EDIT] - I wrote my inference code from scratch and now it is fixed. I still don't know where the bug was but at least I have something working! :P\n\nHello all,\n\nI am trying to perform 5-fold cross-validation but it is failing miserably.\n\nI tried a single fold case, can got 0.645 on the public LB. I took the same code and trained 5 folds in one kernel. My OOF lwlrap was 0.81. In another kernel, I loaded the models, and performed inferences with all 5 models and averaged them. I got 0.327 on the Public LB. Here is some samples of my code:\n\nTraining:\n\n```\nfrom sklearn.model_selection import KFold\n\nn_fold = 5\nseed = 42\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=seed)\n\nfor fold_, (trn_idx, val_idx) in enumerate(folds.split(trn_curated_df['fname'],trn_curated_df['labels'])):\n    tfms = get_transforms(do_flip=True, max_rotate=0, max_lighting=0.1, max_zoom=0, max_warp=0.)\n    src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n    data = (src.transform(tfms, size=128)\n            .databunch(bs=128).normalize(imagenet_stats)\n    )\n    learn = cnn_learner(data, borrowed_model, pretrained=False, metrics=[lwlrap]).mixup(stack_y=False)\n    learn.unfreeze()\n    learn.fit_one_cycle(200, 2e-2,callbacks=[SaveModelCallback(learn, every='improvement', monitor='lwlrap', name='best')])\n    learn.export('FAT2019_Simple'+str(fold_)+'.pkl')\n```\n\nInference:\n\n```\n\npredictions = torch.from_numpy(np.zeros((1120,80))).float()\nprint(predictions.shape)\nnum_folds = 5\nfor i in range(num_folds):\n    learn = load_learner(path = '.', file='../input/simple-model-fat2019/FAT2019_Simple'+str(i)+'.pkl',test = test)\n    preds, _ = learn.TTA(ds_type=DatasetType.Test,aug_tfms=tfms)\n    print(preds.shape)\n    predictions = predictions + preds\npredictions = predictions/num_folds\n```\nWhat could be wrong here?",
      "votes": null
    },
    {
      "id": "534536",
      "postDate": "05/21/2019 12:24:31",
      "content": "<p>Maybe different package version?</p>",
      "rawMarkdown": "Maybe different package version?",
      "votes": null
    },
    {
      "id": "534672",
      "postDate": "05/21/2019 16:47:21",
      "content": "<p>Not like I'm pro in cross-validation, but the following code looks very odd:</p>\n\n<p>&gt;     src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )</p>\n\n<p>KFold already gave you <strong>trn idx</strong> and <strong>val idx</strong> so why do you need another 20% split?</p>",
      "rawMarkdown": "Not like I'm pro in cross-validation, but the following code looks very odd:\n\n&gt;     src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n\nKFold already gave you **trn idx** and **val idx** so why do you need another 20% split?",
      "votes": null
    },
    {
      "id": "534749",
      "postDate": "05/21/2019 18:58:36",
      "content": "<p>Indeed I realized this error and am going to resubmit later at 0:00 GMT (ran out of submissions for the day). But why would that lead to such a low score. I feel that this code would only lead to bunch of averaged models, but if each model is doing about 0.645, the ensemble shouldn't do so terrible. </p>\n\n<p>If there is a flaw in my thinking, please let me know.</p>",
      "rawMarkdown": "Indeed I realized this error and am going to resubmit later at 0:00 GMT (ran out of submissions for the day). But why would that lead to such a low score. I feel that this code would only lead to bunch of averaged models, but if each model is doing about 0.645, the ensemble shouldn't do so terrible. \n\nIf there is a flaw in my thinking, please let me know.",
      "votes": null
    },
    {
      "id": "534923",
      "postDate": "05/22/2019 04:10:13",
      "content": "<p>This did not fix the issue.</p>",
      "rawMarkdown": "This did not fix the issue.",
      "votes": null
    },
    {
      "id": "538804",
      "postDate": "05/29/2019 06:19:54",
      "content": "<p>Hi, <code>do_flip=True</code> could change the sound, I think that might be the reason</p>",
      "rawMarkdown": "Hi, `do_flip=True` could change the sound, I think that might be the reason",
      "votes": null
    },
    {
      "id": "538833",
      "postDate": "05/29/2019 06:56:42",
      "content": "<p>I forgot to note but this was solved. I just kinda started from scratch and it worked. I will edit the post to reflect that.</p>",
      "rawMarkdown": "I forgot to note but this was solved. I just kinda started from scratch and it worked. I will edit the post to reflect that.",
      "votes": null
    },
    {
      "id": "538847",
      "postDate": "05/29/2019 07:20:18",
      "content": "<p>Could you share the new inference code if that's ok with you?</p>",
      "rawMarkdown": "Could you share the new inference code if that's ok with you?",
      "votes": null
    },
    {
      "id": "539290",
      "postDate": "05/29/2019 20:44:21",
      "content": "<p>I don't think there was anything wrong with the code I posted above I think it was something else in my inference code.</p>\n\n<p>Only problem was mentioned below, that I split randomly when I should pass in the indices from KFold:\n<code>.split_by__idx(val_idx)</code></p>\n\n<p>So feel free to use the above code for KFold with fastai</p>",
      "rawMarkdown": "I don't think there was anything wrong with the code I posted above I think it was something else in my inference code.\n\nOnly problem was mentioned below, that I split randomly when I should pass in the indices from KFold:\n`.split_by__idx(val_idx)`\n\nSo feel free to use the above code for KFold with fastai",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 534536,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "05/21/2019 12:24:31",
      "content": "<p>Maybe different package version?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 534672,
      "author_name": "vandalko",
      "author_url": "",
      "post_date": "05/21/2019 16:47:21",
      "content": "<p>Not like I'm pro in cross-validation, but the following code looks very odd:</p>\n\n<p>&gt;     src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )</p>\n\n<p>KFold already gave you <strong>trn idx</strong> and <strong>val idx</strong> so why do you need another 20% split?</p>",
      "votes": null,
      "replies": [
        {
          "id": 534749,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "05/21/2019 18:58:36",
          "content": "<p>Indeed I realized this error and am going to resubmit later at 0:00 GMT (ran out of submissions for the day). But why would that lead to such a low score. I feel that this code would only lead to bunch of averaged models, but if each model is doing about 0.645, the ensemble shouldn't do so terrible. </p>\n\n<p>If there is a flaw in my thinking, please let me know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 534923,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "05/22/2019 04:10:13",
          "content": "<p>This did not fix the issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 538804,
      "author_name": "tanmaypandey",
      "author_url": "",
      "post_date": "05/29/2019 06:19:54",
      "content": "<p>Hi, <code>do_flip=True</code> could change the sound, I think that might be the reason</p>",
      "votes": null,
      "replies": [
        {
          "id": 538833,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "05/29/2019 06:56:42",
          "content": "<p>I forgot to note but this was solved. I just kinda started from scratch and it worked. I will edit the post to reflect that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538847,
          "author_name": "tanmaypandey",
          "author_url": "",
          "post_date": "05/29/2019 07:20:18",
          "content": "<p>Could you share the new inference code if that's ok with you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539290,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "05/29/2019 20:44:21",
          "content": "<p>I don't think there was anything wrong with the code I posted above I think it was something else in my inference code.</p>\n\n<p>Only problem was mentioned below, that I split randomly when I should pass in the indices from KFold:\n<code>.split_by__idx(val_idx)</code></p>\n\n<p>So feel free to use the above code for KFold with fastai</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "534342": "[EDIT] - I wrote my inference code from scratch and now it is fixed. I still don't know where the bug was but at least I have something working! :P\n\nHello all,\n\nI am trying to perform 5-fold cross-validation but it is failing miserably.\n\nI tried a single fold case, can got 0.645 on the public LB. I took the same code and trained 5 folds in one kernel. My OOF lwlrap was 0.81. In another kernel, I loaded the models, and performed inferences with all 5 models and averaged them. I got 0.327 on the Public LB. Here is some samples of my code:\n\nTraining:\n\n```\nfrom sklearn.model_selection import KFold\n\nn_fold = 5\nseed = 42\nfolds = KFold(n_splits=n_fold, shuffle=True, random_state=seed)\n\nfor fold_, (trn_idx, val_idx) in enumerate(folds.split(trn_curated_df['fname'],trn_curated_df['labels'])):\n    tfms = get_transforms(do_flip=True, max_rotate=0, max_lighting=0.1, max_zoom=0, max_warp=0.)\n    src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n    data = (src.transform(tfms, size=128)\n            .databunch(bs=128).normalize(imagenet_stats)\n    )\n    learn = cnn_learner(data, borrowed_model, pretrained=False, metrics=[lwlrap]).mixup(stack_y=False)\n    learn.unfreeze()\n    learn.fit_one_cycle(200, 2e-2,callbacks=[SaveModelCallback(learn, every='improvement', monitor='lwlrap', name='best')])\n    learn.export('FAT2019_Simple'+str(fold_)+'.pkl')\n```\n\nInference:\n\n```\n\npredictions = torch.from_numpy(np.zeros((1120,80))).float()\nprint(predictions.shape)\nnum_folds = 5\nfor i in range(num_folds):\n    learn = load_learner(path = '.', file='../input/simple-model-fat2019/FAT2019_Simple'+str(i)+'.pkl',test = test)\n    preds, _ = learn.TTA(ds_type=DatasetType.Test,aug_tfms=tfms)\n    print(preds.shape)\n    predictions = predictions + preds\npredictions = predictions/num_folds\n```\nWhat could be wrong here?",
    "534536": "Maybe different package version?",
    "534672": "Not like I'm pro in cross-validation, but the following code looks very odd:\n\n&gt;     src = (ImageList.from_csv(WORK, Path('..')/CSV_TRN_CURATED, folder='trn_curated')\n           .split_by_rand_pct(0.2)\n           .label_from_df(label_delim=',')\n    )\n\nKFold already gave you **trn idx** and **val idx** so why do you need another 20% split?",
    "534749": "Indeed I realized this error and am going to resubmit later at 0:00 GMT (ran out of submissions for the day). But why would that lead to such a low score. I feel that this code would only lead to bunch of averaged models, but if each model is doing about 0.645, the ensemble shouldn't do so terrible. \n\nIf there is a flaw in my thinking, please let me know.",
    "534923": "This did not fix the issue.",
    "538804": "Hi, `do_flip=True` could change the sound, I think that might be the reason",
    "538833": "I forgot to note but this was solved. I just kinda started from scratch and it worked. I will edit the post to reflect that.",
    "538847": "Could you share the new inference code if that's ok with you?",
    "539290": "I don't think there was anything wrong with the code I posted above I think it was something else in my inference code.\n\nOnly problem was mentioned below, that I split randomly when I should pass in the indices from KFold:\n`.split_by__idx(val_idx)`\n\nSo feel free to use the above code for KFold with fastai"
  },
  "source": "meta"
}