{
  "id": 49320,
  "title": "Optimization or Magic features",
  "url": "/competitions/sp-society-camera-model-identification/writeups/ods-ai-nokia3310-optimization-or-magic-features",
  "author_name": "",
  "post_date": "2018-02-09T11:30:51.725203700Z",
  "votes": 10,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I would like to thank to Kaggle to organize this challenge, a big thank to all people especially Andres Torrubia, and of course many thanks to my awesome team mates. In this thread I would like to share one idea/improvement that were not mentioned by top teams.</p>\n\n<p>The idea is to apply a post-processing step to balance number of images for each camera type. The input is the actual probabilities, the number of images for each group and the output is the final prediction. The optimal objective is the sum of all (probability * assignment), where assignment can be zero or one.</p>\n\n<p>What is the number of images for each camera? We discovered that the proportion of different type of images are almost equal. Even more, the number of unalt images and manip images are the same as well. And we also discovered that the public test and private test set are split by file timestamp.</p>\n\n<p>The optimization code is quite simple and it runs within a second, you can have a look at one of our first version at <a href=\"https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984\">https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984</a></p>",
  "messages": [
    {
      "id": "280105",
      "postDate": "02/09/2018 11:30:51",
      "content": "<p>First of all, I would like to thank to Kaggle to organize this challenge, a big thank to all people especially Andres Torrubia, and of course many thanks to my awesome team mates. In this thread I would like to share one idea/improvement that were not mentioned by top teams.</p>\n\n<p>The idea is to apply a post-processing step to balance number of images for each camera type. The input is the actual probabilities, the number of images for each group and the output is the final prediction. The optimal objective is the sum of all (probability * assignment), where assignment can be zero or one.</p>\n\n<p>What is the number of images for each camera? We discovered that the proportion of different type of images are almost equal. Even more, the number of unalt images and manip images are the same as well. And we also discovered that the public test and private test set are split by file timestamp.</p>\n\n<p>The optimization code is quite simple and it runs within a second, you can have a look at one of our first version at <a href=\"https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984\">https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984</a></p>",
      "rawMarkdown": "First of all, I would like to thank to Kaggle to organize this challenge, a big thank to all people especially Andres Torrubia, and of course many thanks to my awesome team mates. In this thread I would like to share one idea/improvement that were not mentioned by top teams.\n\nThe idea is to apply a post-processing step to balance number of images for each camera type. The input is the actual probabilities, the number of images for each group and the output is the final prediction. The optimal objective is the sum of all (probability * assignment), where assignment can be zero or one.\n\nWhat is the number of images for each camera? We discovered that the proportion of different type of images are almost equal. Even more, the number of unalt images and manip images are the same as well. And we also discovered that the public test and private test set are split by file timestamp.\n\nThe optimization code is quite simple and it runs within a second, you can have a look at one of our first version at https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984",
      "votes": null
    },
    {
      "id": "280342",
      "postDate": "02/09/2018 18:36:23",
      "content": "<p>I also did post-processing to equalize number of images for each camera, but at the end I have determined I overfitted the public set and my regular inference worked better on the private test set.</p>\n\n<p>How did you find about about file timestamp? How much did you have to probe the public LB?</p>",
      "rawMarkdown": "I also did post-processing to equalize number of images for each camera, but at the end I have determined I overfitted the public set and my regular inference worked better on the private test set.\n\nHow did you find about about file timestamp? How much did you have to probe the public LB?",
      "votes": null
    },
    {
      "id": "280362",
      "postDate": "02/09/2018 20:05:15",
      "content": "<p>@Andres. I probed the LB quite a lot during early phase of the competition to discover the distribution of images as I already knew that it was possible to optimize the predictions. Thanks to your code we had a pretty good result and we decided to probe the LB in the other way to get submissions with zero accuracy.</p>\n\n<p>On the subject of the LB split, I sorted the test files by timestamp and made an accumulate count for both manip and unalt images. The plot already shows it clearly.  <img src=\"https://image.ibb.co/b0G73x/unalt_manip.png\" alt=\"Number of manip/unalt images by timestamp \"> </p>\n\n<p>We even verified that by giving the best predictions to the first 36% images and the wrong predictions to the rest. And the public score for this new submission did not change.</p>",
      "rawMarkdown": "Andres. I probed the LB quite a lot during early phase of the competition to discover the distribution of images as I already knew that it was possible to optimize the predictions. Thanks to your code we had a pretty good result and we decided to probe the LB in the other way to get submissions with zero accuracy.\n\nOn the subject of the LB split, I sorted the test files by timestamp and made an accumulate count for both manip and unalt images. The plot already shows it clearly.  ![Number of manip/unalt images by timestamp ][1] \n\nWe even verified that by giving the best predictions to the first 36% images and the wrong predictions to the rest. And the public score for this new submission did not change.\n\n\n  [1]: https://image.ibb.co/b0G73x/unalt_manip.png",
      "votes": null
    },
    {
      "id": "280410",
      "postDate": "02/09/2018 22:59:55",
      "content": "<p>Here's my attempt at understanding your technique, please correct me:</p>\n\n<p>The gist: Use an estimated test set class distribution to evaluate your predictions, and use that evaluation to improve your predictions. For example, if the distribution is 10% per class, and you're predicting 15% of the images as class_1, you know at least 5% of the predictions are wrong. Use this knowledge to improve your predictions.</p>\n\n<p>If I'm correct so far, I don't know how you're using the knowledge. Is that where MCF comes in?</p>\n\n<p>One way you might be able to use the knowledge is to assign the second most confident label to the most uncertainly labeled images in class_1, until you've moved the 5% to other classes (1/3 of class_1 labels). The thought is: \"while we know that 5% of these labels are wrong, which 5% is it? Let's assume the least confident predictions are the incorrect ones.\" One could repeat this for each class until hitting a threshold.</p>",
      "rawMarkdown": "Here's my attempt at understanding your technique, please correct me:\n\nThe gist: Use an estimated test set class distribution to evaluate your predictions, and use that evaluation to improve your predictions. For example, if the distribution is 10% per class, and you're predicting 15% of the images as class_1, you know at least 5% of the predictions are wrong. Use this knowledge to improve your predictions.\n\nIf I'm correct so far, I don't know how you're using the knowledge. Is that where MCF comes in?\n\nOne way you might be able to use the knowledge is to assign the second most confident label to the most uncertainly labeled images in class_1, until you've moved the 5% to other classes (1/3 of class_1 labels). The thought is: \"while we know that 5% of these labels are wrong, which 5% is it? Let's assume the least confident predictions are the incorrect ones.\" One could repeat this for each class until hitting a threshold.",
      "votes": null
    },
    {
      "id": "280418",
      "postDate": "02/09/2018 23:32:14",
      "content": "<p>Yes, that's most basic \"greedy\" algorithm you may use. There are more complicated schemes possible. </p>",
      "rawMarkdown": "Yes, that's most basic \"greedy\" algorithm you may use. There are more complicated schemes possible.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 280342,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "02/09/2018 18:36:23",
      "content": "<p>I also did post-processing to equalize number of images for each camera, but at the end I have determined I overfitted the public set and my regular inference worked better on the private test set.</p>\n\n<p>How did you find about about file timestamp? How much did you have to probe the public LB?</p>",
      "votes": null,
      "replies": [
        {
          "id": 280362,
          "author_name": "tarobxl",
          "author_url": "",
          "post_date": "02/09/2018 20:05:15",
          "content": "<p>@Andres. I probed the LB quite a lot during early phase of the competition to discover the distribution of images as I already knew that it was possible to optimize the predictions. Thanks to your code we had a pretty good result and we decided to probe the LB in the other way to get submissions with zero accuracy.</p>\n\n<p>On the subject of the LB split, I sorted the test files by timestamp and made an accumulate count for both manip and unalt images. The plot already shows it clearly.  <img src=\"https://image.ibb.co/b0G73x/unalt_manip.png\" alt=\"Number of manip/unalt images by timestamp \"> </p>\n\n<p>We even verified that by giving the best predictions to the first 36% images and the wrong predictions to the rest. And the public score for this new submission did not change.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 280410,
      "author_name": "kleinsmith",
      "author_url": "",
      "post_date": "02/09/2018 22:59:55",
      "content": "<p>Here's my attempt at understanding your technique, please correct me:</p>\n\n<p>The gist: Use an estimated test set class distribution to evaluate your predictions, and use that evaluation to improve your predictions. For example, if the distribution is 10% per class, and you're predicting 15% of the images as class_1, you know at least 5% of the predictions are wrong. Use this knowledge to improve your predictions.</p>\n\n<p>If I'm correct so far, I don't know how you're using the knowledge. Is that where MCF comes in?</p>\n\n<p>One way you might be able to use the knowledge is to assign the second most confident label to the most uncertainly labeled images in class_1, until you've moved the 5% to other classes (1/3 of class_1 labels). The thought is: \"while we know that 5% of these labels are wrong, which 5% is it? Let's assume the least confident predictions are the incorrect ones.\" One could repeat this for each class until hitting a threshold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 280418,
          "author_name": "ceperaang",
          "author_url": "",
          "post_date": "02/09/2018 23:32:14",
          "content": "<p>Yes, that's most basic \"greedy\" algorithm you may use. There are more complicated schemes possible. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "280105": "First of all, I would like to thank to Kaggle to organize this challenge, a big thank to all people especially Andres Torrubia, and of course many thanks to my awesome team mates. In this thread I would like to share one idea/improvement that were not mentioned by top teams.\n\nThe idea is to apply a post-processing step to balance number of images for each camera type. The input is the actual probabilities, the number of images for each group and the output is the final prediction. The optimal objective is the sum of all (probability * assignment), where assignment can be zero or one.\n\nWhat is the number of images for each camera? We discovered that the proportion of different type of images are almost equal. Even more, the number of unalt images and manip images are the same as well. And we also discovered that the public test and private test set are split by file timestamp.\n\nThe optimization code is quite simple and it runs within a second, you can have a look at one of our first version at https://www.kaggle.com/golubev/simple-example-mcf-balance-improve-0-976-to-0-984",
    "280342": "I also did post-processing to equalize number of images for each camera, but at the end I have determined I overfitted the public set and my regular inference worked better on the private test set.\n\nHow did you find about about file timestamp? How much did you have to probe the public LB?",
    "280362": "Andres. I probed the LB quite a lot during early phase of the competition to discover the distribution of images as I already knew that it was possible to optimize the predictions. Thanks to your code we had a pretty good result and we decided to probe the LB in the other way to get submissions with zero accuracy.\n\nOn the subject of the LB split, I sorted the test files by timestamp and made an accumulate count for both manip and unalt images. The plot already shows it clearly.  ![Number of manip/unalt images by timestamp ][1] \n\nWe even verified that by giving the best predictions to the first 36% images and the wrong predictions to the rest. And the public score for this new submission did not change.\n\n\n  [1]: https://image.ibb.co/b0G73x/unalt_manip.png",
    "280410": "Here's my attempt at understanding your technique, please correct me:\n\nThe gist: Use an estimated test set class distribution to evaluate your predictions, and use that evaluation to improve your predictions. For example, if the distribution is 10% per class, and you're predicting 15% of the images as class_1, you know at least 5% of the predictions are wrong. Use this knowledge to improve your predictions.\n\nIf I'm correct so far, I don't know how you're using the knowledge. Is that where MCF comes in?\n\nOne way you might be able to use the knowledge is to assign the second most confident label to the most uncertainly labeled images in class_1, until you've moved the 5% to other classes (1/3 of class_1 labels). The thought is: \"while we know that 5% of these labels are wrong, which 5% is it? Let's assume the least confident predictions are the incorrect ones.\" One could repeat this for each class until hitting a threshold.",
    "280418": "Yes, that's most basic \"greedy\" algorithm you may use. There are more complicated schemes possible."
  },
  "source": "meta"
}