{
  "id": 143626,
  "title": "AUC score is quite low. Why?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/143626",
  "author_name": "",
  "post_date": "2020-04-15T17:04:05.946289700Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My train data:-</p>\n\n<p>Epoch 1/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.3664 - auc: 0.9152 - val_loss: 0.2803 - val_auc: 0.9553\nEpoch 2/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.2637 - auc: 0.9570 - val_loss: 0.2405 - val_auc: 0.9653\nEpoch 3/3\n5750/5750 [==============================] - 692s 120ms/step - loss: 0.2405 - auc: 0.9641 - val_loss: 0.2223 - val_auc: 0.9693</p>\n\n<p>But on Leaderboard I got only 0.82 score.\nCan any one tell me why this happen? </p>",
  "messages": [
    {
      "id": "808847",
      "postDate": "04/15/2020 17:04:05",
      "content": "<p>My train data:-</p>\n\n<p>Epoch 1/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.3664 - auc: 0.9152 - val_loss: 0.2803 - val_auc: 0.9553\nEpoch 2/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.2637 - auc: 0.9570 - val_loss: 0.2405 - val_auc: 0.9653\nEpoch 3/3\n5750/5750 [==============================] - 692s 120ms/step - loss: 0.2405 - auc: 0.9641 - val_loss: 0.2223 - val_auc: 0.9693</p>\n\n<p>But on Leaderboard I got only 0.82 score.\nCan any one tell me why this happen? </p>",
      "rawMarkdown": "My train data:-\n\nEpoch 1/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.3664 - auc: 0.9152 - val_loss: 0.2803 - val_auc: 0.9553\nEpoch 2/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.2637 - auc: 0.9570 - val_loss: 0.2405 - val_auc: 0.9653\nEpoch 3/3\n5750/5750 [==============================] - 692s 120ms/step - loss: 0.2405 - auc: 0.9641 - val_loss: 0.2223 - val_auc: 0.9693\n\n\nBut on Leaderboard I got only 0.82 score.\nCan any one tell me why this happen?",
      "votes": null
    },
    {
      "id": "808950",
      "postDate": "04/15/2020 18:32:34",
      "content": "<p><a href=\"/lucca9211\">@lucca9211</a> What is your cv strategy? Maybe just split train set to 5 folds? Do you use validation set which is already prepared by host?</p>",
      "rawMarkdown": "lucca9211 What is your cv strategy? Maybe just split train set to 5 folds? Do you use validation set which is already prepared by host?",
      "votes": null
    },
    {
      "id": "808962",
      "postDate": "04/15/2020 18:40:48",
      "content": "<p>I have meeged the data including validation set but not test set.\nThen train test split.</p>",
      "rawMarkdown": "I have meeged the data including validation set but not test set.\nThen train test split.",
      "votes": null
    },
    {
      "id": "808963",
      "postDate": "04/15/2020 18:41:54",
      "content": "<p>I know my error\nValidation set should be used original</p>",
      "rawMarkdown": "I know my error\nValidation set should be used original",
      "votes": null
    },
    {
      "id": "808998",
      "postDate": "04/15/2020 19:01:09",
      "content": "<p><a href=\"/lucca9211\">@lucca9211</a> Even if you include val set, which is multi-lingual text, to train set, most of your train set is English test, thus your high AUC comes from English text. That doesn't assure high AUC on test set, which is multi-lingual text. You should keep val set for validation, at least a fold if you want to use them as a training set.</p>",
      "rawMarkdown": "lucca9211 Even if you include val set, which is multi-lingual text, to train set, most of your train set is English test, thus your high AUC comes from English text. That doesn't assure high AUC on test set, which is multi-lingual text. You should keep val set for validation, at least a fold if you want to use them as a training set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 808950,
      "author_name": "bamps53",
      "author_url": "",
      "post_date": "04/15/2020 18:32:34",
      "content": "<p><a href=\"/lucca9211\">@lucca9211</a> What is your cv strategy? Maybe just split train set to 5 folds? Do you use validation set which is already prepared by host?</p>",
      "votes": null,
      "replies": [
        {
          "id": 808962,
          "author_name": "lucca9211",
          "author_url": "",
          "post_date": "04/15/2020 18:40:48",
          "content": "<p>I have meeged the data including validation set but not test set.\nThen train test split.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 808963,
          "author_name": "lucca9211",
          "author_url": "",
          "post_date": "04/15/2020 18:41:54",
          "content": "<p>I know my error\nValidation set should be used original</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 808998,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "04/15/2020 19:01:09",
          "content": "<p><a href=\"/lucca9211\">@lucca9211</a> Even if you include val set, which is multi-lingual text, to train set, most of your train set is English test, thus your high AUC comes from English text. That doesn't assure high AUC on test set, which is multi-lingual text. You should keep val set for validation, at least a fold if you want to use them as a training set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "808847": "My train data:-\n\nEpoch 1/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.3664 - auc: 0.9152 - val_loss: 0.2803 - val_auc: 0.9553\nEpoch 2/3\n5750/5750 [==============================] - 690s 120ms/step - loss: 0.2637 - auc: 0.9570 - val_loss: 0.2405 - val_auc: 0.9653\nEpoch 3/3\n5750/5750 [==============================] - 692s 120ms/step - loss: 0.2405 - auc: 0.9641 - val_loss: 0.2223 - val_auc: 0.9693\n\n\nBut on Leaderboard I got only 0.82 score.\nCan any one tell me why this happen?",
    "808950": "lucca9211 What is your cv strategy? Maybe just split train set to 5 folds? Do you use validation set which is already prepared by host?",
    "808962": "I have meeged the data including validation set but not test set.\nThen train test split.",
    "808963": "I know my error\nValidation set should be used original",
    "808998": "lucca9211 Even if you include val set, which is multi-lingual text, to train set, most of your train set is English test, thus your high AUC comes from English text. That doesn't assure high AUC on test set, which is multi-lingual text. You should keep val set for validation, at least a fold if you want to use them as a training set."
  },
  "source": "meta"
}