{
  "id": 75923,
  "title": "flow_from_... messing up order in submission?",
  "url": "/competitions/histopathologic-cancer-detection/discussion/75923",
  "author_name": "",
  "post_date": "2018-12-27T16:29:36.129009Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>final_test_generator = testdatagen.flow_from_directory(directory=\"../test\", \n                                                 target_size=(96,96), \n                                                 color_mode='rgb', \n                                                 classes=None, \n                                                 class_mode=None,\n                                                 shuffle=False)</p>\n\n<p>result = model[1].predict_generator(final_test_generator, \n                                    verbose=1,\n                                    steps=test_generator.n)</p>\n\n<p>Im using the above to build out my image generator for my prediction/submission. Ive split my test data 3 ways, train, validation and a final validation (just to check on unseen data). My AUC is reasonably high. Last night's commit AUC was in the .92 range on my third set of data not touched by the model, but my submissions are .5 and below consistently. Even trying other's models I score poorly, likely my data prep and predicting. Could my generator be mixing up my prediction values and the id?</p>\n\n<p>Any insights?</p>",
  "messages": [
    {
      "id": "446168",
      "postDate": "12/27/2018 16:29:36",
      "content": "<p>final_test_generator = testdatagen.flow_from_directory(directory=\"../test\", \n                                                 target_size=(96,96), \n                                                 color_mode='rgb', \n                                                 classes=None, \n                                                 class_mode=None,\n                                                 shuffle=False)</p>\n\n<p>result = model[1].predict_generator(final_test_generator, \n                                    verbose=1,\n                                    steps=test_generator.n)</p>\n\n<p>Im using the above to build out my image generator for my prediction/submission. Ive split my test data 3 ways, train, validation and a final validation (just to check on unseen data). My AUC is reasonably high. Last night's commit AUC was in the .92 range on my third set of data not touched by the model, but my submissions are .5 and below consistently. Even trying other's models I score poorly, likely my data prep and predicting. Could my generator be mixing up my prediction values and the id?</p>\n\n<p>Any insights?</p>",
      "rawMarkdown": "final_test_generator = testdatagen.flow_from_directory(directory=\"../test\", \n                                                 target_size=(96,96), \n                                                 color_mode='rgb', \n                                                 classes=None, \n                                                 class_mode=None,\n                                                 shuffle=False)\n\nresult = model[1].predict_generator(final_test_generator, \n                                    verbose=1,\n                                    steps=test_generator.n)\n\nIm using the above to build out my image generator for my prediction/submission. Ive split my test data 3 ways, train, validation and a final validation (just to check on unseen data). My AUC is reasonably high. Last night's commit AUC was in the .92 range on my third set of data not touched by the model, but my submissions are .5 and below consistently. Even trying other's models I score poorly, likely my data prep and predicting. Could my generator be mixing up my prediction values and the id?\n\nAny insights?",
      "votes": null
    },
    {
      "id": "446437",
      "postDate": "12/28/2018 04:31:02",
      "content": "<p>I had some how similar issues. Make sure you are submitting the probabilities of the output Not 0 and 1. Second AUC of .92 is not that high i guess. I am getting 0.99xx and still have lower lb score.</p>",
      "rawMarkdown": "I had some how similar issues. Make sure you are submitting the probabilities of the output Not 0 and 1. Second AUC of .92 is not that high i guess. I am getting 0.99xx and still have lower lb score.",
      "votes": null
    },
    {
      "id": "446968",
      "postDate": "12/28/2018 23:38:42",
      "content": "<p>Looks like your lb is .97. I’d be happy to get that score. </p>\n\n<p>Yeah, I’ve been submitting probabilities for a few dozen commits now. While .92 might not be high it’s way more than what I’m getting scored on my submissions. It doesn’t seem like AUC from validation really translates to AUC with the final data. </p>",
      "rawMarkdown": "Looks like your lb is .97. I’d be happy to get that score. \n\nYeah, I’ve been submitting probabilities for a few dozen commits now. While .92 might not be high it’s way more than what I’m getting scored on my submissions. It doesn’t seem like AUC from validation really translates to AUC with the final data.",
      "votes": null
    },
    {
      "id": "447076",
      "postDate": "12/29/2018 05:16:40",
      "content": "<p>What is your accuracy number? did you try to use other metrics like accuracy or loss metrics and see how your model performs. Try also to see your confusion matrix to see how many images are misclassified. The best explanation for AUC metrics i have found so far is here. it might help youyouhttps://www.dataschool.io/roc-curves-and-auc-explained/</p>\n\n<p>Yeah. I tried a couple of ensembling method yesterday and boosted my score a little bit higher..</p>",
      "rawMarkdown": "What is your accuracy number? did you try to use other metrics like accuracy or loss metrics and see how your model performs. Try also to see your confusion matrix to see how many images are misclassified. The best explanation for AUC metrics i have found so far is here. it might help youyouhttps://www.dataschool.io/roc-curves-and-auc-explained/\n\nYeah. I tried a couple of ensembling method yesterday and boosted my score a little bit higher..",
      "votes": null
    },
    {
      "id": "447518",
      "postDate": "12/30/2018 01:56:01",
      "content": "<p>Accuracy was so-so in the 80s. I ensembled 5 of the same model together earlier today with label being the average of their outputs and jumped my score up to .95 from my previous highest of .50!!!! I had to manually upload the file because my commit failed to drop the columns from my final submission DataFrame &gt;:( The data was there though so Im running again after working on my mistakes. Waiting for a commit to finish now. Fingers crossed. </p>\n\n<p>Thanks for all your help!</p>",
      "rawMarkdown": "Accuracy was so-so in the 80s. I ensembled 5 of the same model together earlier today with label being the average of their outputs and jumped my score up to .95 from my previous highest of .50!!!! I had to manually upload the file because my commit failed to drop the columns from my final submission DataFrame &gt;:( The data was there though so Im running again after working on my mistakes. Waiting for a commit to finish now. Fingers crossed. \n\nThanks for all your help!",
      "votes": null
    },
    {
      "id": "448634",
      "postDate": "01/01/2019 18:11:34",
      "content": "<p>I had the same issue where I got good validation accuracy according to <code>model.fit()</code> callback but at the end when I wanted to evaluate the model myself, I had the score of a random prediction. \nI came to the conclusion that keras flow_from_directory doesn't produce an ordered flow. Thus, you can't simply replace the last column of the submission file with the output of <code>model.predict_generator</code> you need to correctly match the id column with the correct prediction. \nThus, I advise you to use concatenate <code>final_test_generator.filenames</code> with the output of <code>model.predict_generator</code>. </p>",
      "rawMarkdown": "I had the same issue where I got good validation accuracy according to `model.fit()` callback but at the end when I wanted to evaluate the model myself, I had the score of a random prediction. \nI came to the conclusion that keras flow_from_directory doesn't produce an ordered flow. Thus, you can't simply replace the last column of the submission file with the output of `model.predict_generator` you need to correctly match the id column with the correct prediction. \nThus, I advise you to use concatenate `final_test_generator.filenames` with the output of `model.predict_generator`.",
      "votes": null
    },
    {
      "id": "451612",
      "postDate": "01/07/2019 11:06:46",
      "content": "<p>I am stuck with ~0.5 (wrong order?) output with <code>.flow_from_dataframe</code> for a while. May I ask why would ensemble theory be of help to resolve this problem?</p>",
      "rawMarkdown": "I am stuck with ~0.5 (wrong order?) output with ```.flow_from_dataframe``` for a while. May I ask why would ensemble theory be of help to resolve this problem?",
      "votes": null
    },
    {
      "id": "451614",
      "postDate": "01/07/2019 11:16:51",
      "content": "<p>This is of great help! Just realised the order of the <code>.flow_from_dataframe</code> is sorted. Really appreciate the advice!</p>",
      "rawMarkdown": "This is of great help! Just realised the order of the ```.flow_from_dataframe``` is sorted. Really appreciate the advice!",
      "votes": null
    },
    {
      "id": "451858",
      "postDate": "01/07/2019 19:27:02",
      "content": "<p>If you’ve ruled out the order issue from the flow_from_whatever generators then look into ensembles. In simple terms they are averaging outputs from several training sessions. It’s like getting multiple opinions on something. Running your model several times likely doesn’t produce the exact same AUC every run. There is some randomness involved especially if you’re splitting your train and validate data correctly. Using different architectures will pick up on different features. The probability that the average of these runs is near the correct answer is good. </p>\n\n<p>I’m dealing with time allocation with ensembling now. Do I get better AUC with an ensemble where time is split between models, each having to train fewer epochs so the kernel doesn’t get killed beyond 32400 seconds or whatever, versus one model allowed to use the entire time to itself? </p>",
      "rawMarkdown": "If you’ve ruled out the order issue from the flow_from_whatever generators then look into ensembles. In simple terms they are averaging outputs from several training sessions. It’s like getting multiple opinions on something. Running your model several times likely doesn’t produce the exact same AUC every run. There is some randomness involved especially if you’re splitting your train and validate data correctly. Using different architectures will pick up on different features. The probability that the average of these runs is near the correct answer is good. \n\nI’m dealing with time allocation with ensembling now. Do I get better AUC with an ensemble where time is split between models, each having to train fewer epochs so the kernel doesn’t get killed beyond 32400 seconds or whatever, versus one model allowed to use the entire time to itself?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 446437,
      "author_name": "abdishakuur",
      "author_url": "",
      "post_date": "12/28/2018 04:31:02",
      "content": "<p>I had some how similar issues. Make sure you are submitting the probabilities of the output Not 0 and 1. Second AUC of .92 is not that high i guess. I am getting 0.99xx and still have lower lb score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 446968,
          "author_name": "reidtc",
          "author_url": "",
          "post_date": "12/28/2018 23:38:42",
          "content": "<p>Looks like your lb is .97. I’d be happy to get that score. </p>\n\n<p>Yeah, I’ve been submitting probabilities for a few dozen commits now. While .92 might not be high it’s way more than what I’m getting scored on my submissions. It doesn’t seem like AUC from validation really translates to AUC with the final data. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447076,
          "author_name": "abdishakuur",
          "author_url": "",
          "post_date": "12/29/2018 05:16:40",
          "content": "<p>What is your accuracy number? did you try to use other metrics like accuracy or loss metrics and see how your model performs. Try also to see your confusion matrix to see how many images are misclassified. The best explanation for AUC metrics i have found so far is here. it might help youyouhttps://www.dataschool.io/roc-curves-and-auc-explained/</p>\n\n<p>Yeah. I tried a couple of ensembling method yesterday and boosted my score a little bit higher..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447518,
          "author_name": "reidtc",
          "author_url": "",
          "post_date": "12/30/2018 01:56:01",
          "content": "<p>Accuracy was so-so in the 80s. I ensembled 5 of the same model together earlier today with label being the average of their outputs and jumped my score up to .95 from my previous highest of .50!!!! I had to manually upload the file because my commit failed to drop the columns from my final submission DataFrame &gt;:( The data was there though so Im running again after working on my mistakes. Waiting for a commit to finish now. Fingers crossed. </p>\n\n<p>Thanks for all your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 451612,
          "author_name": "realethanzou",
          "author_url": "",
          "post_date": "01/07/2019 11:06:46",
          "content": "<p>I am stuck with ~0.5 (wrong order?) output with <code>.flow_from_dataframe</code> for a while. May I ask why would ensemble theory be of help to resolve this problem?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 451858,
          "author_name": "reidtc",
          "author_url": "",
          "post_date": "01/07/2019 19:27:02",
          "content": "<p>If you’ve ruled out the order issue from the flow_from_whatever generators then look into ensembles. In simple terms they are averaging outputs from several training sessions. It’s like getting multiple opinions on something. Running your model several times likely doesn’t produce the exact same AUC every run. There is some randomness involved especially if you’re splitting your train and validate data correctly. Using different architectures will pick up on different features. The probability that the average of these runs is near the correct answer is good. </p>\n\n<p>I’m dealing with time allocation with ensembling now. Do I get better AUC with an ensemble where time is split between models, each having to train fewer epochs so the kernel doesn’t get killed beyond 32400 seconds or whatever, versus one model allowed to use the entire time to itself? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448634,
      "author_name": "sdelecourt",
      "author_url": "",
      "post_date": "01/01/2019 18:11:34",
      "content": "<p>I had the same issue where I got good validation accuracy according to <code>model.fit()</code> callback but at the end when I wanted to evaluate the model myself, I had the score of a random prediction. \nI came to the conclusion that keras flow_from_directory doesn't produce an ordered flow. Thus, you can't simply replace the last column of the submission file with the output of <code>model.predict_generator</code> you need to correctly match the id column with the correct prediction. \nThus, I advise you to use concatenate <code>final_test_generator.filenames</code> with the output of <code>model.predict_generator</code>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 451614,
          "author_name": "realethanzou",
          "author_url": "",
          "post_date": "01/07/2019 11:16:51",
          "content": "<p>This is of great help! Just realised the order of the <code>.flow_from_dataframe</code> is sorted. Really appreciate the advice!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "446168": "final_test_generator = testdatagen.flow_from_directory(directory=\"../test\", \n                                                 target_size=(96,96), \n                                                 color_mode='rgb', \n                                                 classes=None, \n                                                 class_mode=None,\n                                                 shuffle=False)\n\nresult = model[1].predict_generator(final_test_generator, \n                                    verbose=1,\n                                    steps=test_generator.n)\n\nIm using the above to build out my image generator for my prediction/submission. Ive split my test data 3 ways, train, validation and a final validation (just to check on unseen data). My AUC is reasonably high. Last night's commit AUC was in the .92 range on my third set of data not touched by the model, but my submissions are .5 and below consistently. Even trying other's models I score poorly, likely my data prep and predicting. Could my generator be mixing up my prediction values and the id?\n\nAny insights?",
    "446437": "I had some how similar issues. Make sure you are submitting the probabilities of the output Not 0 and 1. Second AUC of .92 is not that high i guess. I am getting 0.99xx and still have lower lb score.",
    "446968": "Looks like your lb is .97. I’d be happy to get that score. \n\nYeah, I’ve been submitting probabilities for a few dozen commits now. While .92 might not be high it’s way more than what I’m getting scored on my submissions. It doesn’t seem like AUC from validation really translates to AUC with the final data.",
    "447076": "What is your accuracy number? did you try to use other metrics like accuracy or loss metrics and see how your model performs. Try also to see your confusion matrix to see how many images are misclassified. The best explanation for AUC metrics i have found so far is here. it might help youyouhttps://www.dataschool.io/roc-curves-and-auc-explained/\n\nYeah. I tried a couple of ensembling method yesterday and boosted my score a little bit higher..",
    "447518": "Accuracy was so-so in the 80s. I ensembled 5 of the same model together earlier today with label being the average of their outputs and jumped my score up to .95 from my previous highest of .50!!!! I had to manually upload the file because my commit failed to drop the columns from my final submission DataFrame &gt;:( The data was there though so Im running again after working on my mistakes. Waiting for a commit to finish now. Fingers crossed. \n\nThanks for all your help!",
    "448634": "I had the same issue where I got good validation accuracy according to `model.fit()` callback but at the end when I wanted to evaluate the model myself, I had the score of a random prediction. \nI came to the conclusion that keras flow_from_directory doesn't produce an ordered flow. Thus, you can't simply replace the last column of the submission file with the output of `model.predict_generator` you need to correctly match the id column with the correct prediction. \nThus, I advise you to use concatenate `final_test_generator.filenames` with the output of `model.predict_generator`.",
    "451612": "I am stuck with ~0.5 (wrong order?) output with ```.flow_from_dataframe``` for a while. May I ask why would ensemble theory be of help to resolve this problem?",
    "451614": "This is of great help! Just realised the order of the ```.flow_from_dataframe``` is sorted. Really appreciate the advice!",
    "451858": "If you’ve ruled out the order issue from the flow_from_whatever generators then look into ensembles. In simple terms they are averaging outputs from several training sessions. It’s like getting multiple opinions on something. Running your model several times likely doesn’t produce the exact same AUC every run. There is some randomness involved especially if you’re splitting your train and validate data correctly. Using different architectures will pick up on different features. The probability that the average of these runs is near the correct answer is good. \n\nI’m dealing with time allocation with ensembling now. Do I get better AUC with an ensemble where time is split between models, each having to train fewer epochs so the kernel doesn’t get killed beyond 32400 seconds or whatever, versus one model allowed to use the entire time to itself?"
  },
  "source": "meta"
}