{
  "id": 175181,
  "title": "ROC of OOF Predicitions (Model Effnet)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175181",
  "author_name": "",
  "post_date": "2020-08-17T12:17:16.031506800Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Basically decided to play around a bit with the post processing data, and got my model's out of fold prediction on one fold of training data, and found that the ROC is around 0.5?! This got me confused but atleast I got to know that my model is making random predictions on OOF data, however when I ensemble it with metadata even in 9:1 ratio this problem is solved, and the test prediction does well on LB. What is happening with OOF preds?</p>",
  "messages": [
    {
      "id": "973636",
      "postDate": "08/17/2020 12:17:16",
      "content": "<p>Basically decided to play around a bit with the post processing data, and got my model's out of fold prediction on one fold of training data, and found that the ROC is around 0.5?! This got me confused but atleast I got to know that my model is making random predictions on OOF data, however when I ensemble it with metadata even in 9:1 ratio this problem is solved, and the test prediction does well on LB. What is happening with OOF preds?</p>",
      "rawMarkdown": "Basically decided to play around a bit with the post processing data, and got my model's out of fold prediction on one fold of training data, and found that the ROC is around 0.5?! This got me confused but atleast I got to know that my model is making random predictions on OOF data, however when I ensemble it with metadata even in 9:1 ratio this problem is solved, and the test prediction does well on LB. What is happening with OOF preds?",
      "votes": null
    },
    {
      "id": "974336",
      "postDate": "08/17/2020 22:43:45",
      "content": "<p>Is your training data shuffled?  Are you just taking predictions and throwing them into an array/list? Its a bit tricky doing OOF ROC on just one fold.  You would need to properly place your indexes for each prediction into the data structure, and then you would need to run that against the true targets for those exact indexes.  Remember order matters.  When you do all the folds, the indexes all eventually populate the data structure, and that will match correctly to the list of targets.  But like I said, doing just one fold you have to put more thought into it.</p>",
      "rawMarkdown": "Is your training data shuffled?  Are you just taking predictions and throwing them into an array/list? Its a bit tricky doing OOF ROC on just one fold.  You would need to properly place your indexes for each prediction into the data structure, and then you would need to run that against the true targets for those exact indexes.  Remember order matters.  When you do all the folds, the indexes all eventually populate the data structure, and that will match correctly to the list of targets.  But like I said, doing just one fold you have to put more thought into it.",
      "votes": null
    },
    {
      "id": "974734",
      "postDate": "08/18/2020 03:13:50",
      "content": "<p>Its shuffled but the order is maintained. </p>",
      "rawMarkdown": "Its shuffled but the order is maintained.",
      "votes": null
    },
    {
      "id": "974746",
      "postDate": "08/18/2020 03:22:15",
      "content": "<p>with ROC you are comparing two lists (preds and targets). You are saying you have somehow put your shuffled train indexes back into order? How would you do that?  If you created an array of size fold, your indexes will span much larger than that correct?  But lets say you did manage to put them into index order, back into a list.  Now how did you get your ground truths/targets into the same order?</p>",
      "rawMarkdown": "with ROC you are comparing two lists (preds and targets). You are saying you have somehow put your shuffled train indexes back into order? How would you do that?  If you created an array of size fold, your indexes will span much larger than that correct?  But lets say you did manage to put them into index order, back into a list.  Now how did you get your ground truths/targets into the same order?",
      "votes": null
    },
    {
      "id": "974752",
      "postDate": "08/18/2020 03:27:33",
      "content": "<p>Sorry I shouldve been more clear, Im not using the train index order but image to prediction/target order is correct. If you use StratifiedKFold from sklearn.model_selection it gives you the train index to make a fold for.</p>",
      "rawMarkdown": "Sorry I shouldve been more clear, Im not using the train index order but image to prediction/target order is correct. If you use StratifiedKFold from sklearn.model_selection it gives you the train index to make a fold for.",
      "votes": null
    },
    {
      "id": "974776",
      "postDate": "08/18/2020 03:46:04",
      "content": "<p>gotcha, so you are saving the train preds in a list in the same order as you save the associated targets.  Are you also calculating accuracy off the same data, and is that reasonably high?  I ask because thats sort of a check that your preds are aligned reasonably to your targets.</p>",
      "rawMarkdown": "gotcha, so you are saving the train preds in a list in the same order as you save the associated targets.  Are you also calculating accuracy off the same data, and is that reasonably high?  I ask because thats sort of a check that your preds are aligned reasonably to your targets.",
      "votes": null
    },
    {
      "id": "974880",
      "postDate": "08/18/2020 04:39:29",
      "content": "<p>Yeah accuracy for the same oof prediction wrt labels were 97-98, though roc_auc is between 0.48-0.5. I don't understand the explanation for this phenomenon. But exploration of the oof predictions revealed that the values are relatively lower (for eg, array.max was around 0.42), whereas for subs array.max was 0.94. My model seems to be working in making submissions but failing on oof preds.</p>",
      "rawMarkdown": "Yeah accuracy for the same oof prediction wrt labels were 97-98, though roc_auc is between 0.48-0.5. I don't understand the explanation for this phenomenon. But exploration of the oof predictions revealed that the values are relatively lower (for eg, array.max was around 0.42), whereas for subs array.max was 0.94. My model seems to be working in making submissions but failing on oof preds.",
      "votes": null
    },
    {
      "id": "974911",
      "postDate": "08/18/2020 04:52:55",
      "content": "<p>sounds like you are not using sigmoid perhaps on one but are on the other.  when you take a prediction do you run it through something like round(sigmoid(pred)) then compare it to target?  Typically you need to do something like that for Accuracy……and even then its a stretch since .5 is not really a \"true\" threshold for our targets.  With ROC just run your prediction thru sigmoid (unless it was already done at the end of your model dont do it twice), then compare to the like list of targets.</p>",
      "rawMarkdown": "sounds like you are not using sigmoid perhaps on one but are on the other.  when you take a prediction do you run it through something like round(sigmoid(pred)) then compare it to target?  Typically you need to do something like that for Accuracy......and even then its a stretch since .5 is not really a \"true\" threshold for our targets.  With ROC just run your prediction thru sigmoid (unless it was already done at the end of your model dont do it twice), then compare to the like list of targets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974336,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/17/2020 22:43:45",
      "content": "<p>Is your training data shuffled?  Are you just taking predictions and throwing them into an array/list? Its a bit tricky doing OOF ROC on just one fold.  You would need to properly place your indexes for each prediction into the data structure, and then you would need to run that against the true targets for those exact indexes.  Remember order matters.  When you do all the folds, the indexes all eventually populate the data structure, and that will match correctly to the list of targets.  But like I said, doing just one fold you have to put more thought into it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974734,
      "author_name": "dijorajsenroy",
      "author_url": "",
      "post_date": "08/18/2020 03:13:50",
      "content": "<p>Its shuffled but the order is maintained. </p>",
      "votes": null,
      "replies": [
        {
          "id": 974746,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/18/2020 03:22:15",
          "content": "<p>with ROC you are comparing two lists (preds and targets). You are saying you have somehow put your shuffled train indexes back into order? How would you do that?  If you created an array of size fold, your indexes will span much larger than that correct?  But lets say you did manage to put them into index order, back into a list.  Now how did you get your ground truths/targets into the same order?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974752,
      "author_name": "dijorajsenroy",
      "author_url": "",
      "post_date": "08/18/2020 03:27:33",
      "content": "<p>Sorry I shouldve been more clear, Im not using the train index order but image to prediction/target order is correct. If you use StratifiedKFold from sklearn.model_selection it gives you the train index to make a fold for.</p>",
      "votes": null,
      "replies": [
        {
          "id": 974776,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/18/2020 03:46:04",
          "content": "<p>gotcha, so you are saving the train preds in a list in the same order as you save the associated targets.  Are you also calculating accuracy off the same data, and is that reasonably high?  I ask because thats sort of a check that your preds are aligned reasonably to your targets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 974880,
      "author_name": "dijorajsenroy",
      "author_url": "",
      "post_date": "08/18/2020 04:39:29",
      "content": "<p>Yeah accuracy for the same oof prediction wrt labels were 97-98, though roc_auc is between 0.48-0.5. I don't understand the explanation for this phenomenon. But exploration of the oof predictions revealed that the values are relatively lower (for eg, array.max was around 0.42), whereas for subs array.max was 0.94. My model seems to be working in making submissions but failing on oof preds.</p>",
      "votes": null,
      "replies": [
        {
          "id": 974911,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/18/2020 04:52:55",
          "content": "<p>sounds like you are not using sigmoid perhaps on one but are on the other.  when you take a prediction do you run it through something like round(sigmoid(pred)) then compare it to target?  Typically you need to do something like that for Accuracy……and even then its a stretch since .5 is not really a \"true\" threshold for our targets.  With ROC just run your prediction thru sigmoid (unless it was already done at the end of your model dont do it twice), then compare to the like list of targets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "973636": "Basically decided to play around a bit with the post processing data, and got my model's out of fold prediction on one fold of training data, and found that the ROC is around 0.5?! This got me confused but atleast I got to know that my model is making random predictions on OOF data, however when I ensemble it with metadata even in 9:1 ratio this problem is solved, and the test prediction does well on LB. What is happening with OOF preds?",
    "974336": "Is your training data shuffled?  Are you just taking predictions and throwing them into an array/list? Its a bit tricky doing OOF ROC on just one fold.  You would need to properly place your indexes for each prediction into the data structure, and then you would need to run that against the true targets for those exact indexes.  Remember order matters.  When you do all the folds, the indexes all eventually populate the data structure, and that will match correctly to the list of targets.  But like I said, doing just one fold you have to put more thought into it.",
    "974734": "Its shuffled but the order is maintained.",
    "974746": "with ROC you are comparing two lists (preds and targets). You are saying you have somehow put your shuffled train indexes back into order? How would you do that?  If you created an array of size fold, your indexes will span much larger than that correct?  But lets say you did manage to put them into index order, back into a list.  Now how did you get your ground truths/targets into the same order?",
    "974752": "Sorry I shouldve been more clear, Im not using the train index order but image to prediction/target order is correct. If you use StratifiedKFold from sklearn.model_selection it gives you the train index to make a fold for.",
    "974776": "gotcha, so you are saving the train preds in a list in the same order as you save the associated targets.  Are you also calculating accuracy off the same data, and is that reasonably high?  I ask because thats sort of a check that your preds are aligned reasonably to your targets.",
    "974880": "Yeah accuracy for the same oof prediction wrt labels were 97-98, though roc_auc is between 0.48-0.5. I don't understand the explanation for this phenomenon. But exploration of the oof predictions revealed that the values are relatively lower (for eg, array.max was around 0.42), whereas for subs array.max was 0.94. My model seems to be working in making submissions but failing on oof preds.",
    "974911": "sounds like you are not using sigmoid perhaps on one but are on the other.  when you take a prediction do you run it through something like round(sigmoid(pred)) then compare it to target?  Typically you need to do something like that for Accuracy......and even then its a stretch since .5 is not really a \"true\" threshold for our targets.  With ROC just run your prediction thru sigmoid (unless it was already done at the end of your model dont do it twice), then compare to the like list of targets."
  },
  "source": "meta"
}