{
  "id": 56079,
  "title": "image-top-1",
  "url": "/competitions/avito-demand-prediction/discussion/56079",
  "author_name": "",
  "post_date": "2018-05-05T17:03:47.083525600Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>What does <strong><em>image-top-1</em></strong> column represent? The description of data says <em>\"Avito's classification code for the image\"</em>. I really didn't get the meaning.</p>",
  "messages": [
    {
      "id": "323607",
      "postDate": "05/05/2018 17:03:47",
      "content": "<p>What does <strong><em>image-top-1</em></strong> column represent? The description of data says <em>\"Avito's classification code for the image\"</em>. I really didn't get the meaning.</p>",
      "rawMarkdown": "What does ***image-top-1*** column represent? The description of data says *\"Avito's classification code for the image\"*. I really didn't get the meaning.",
      "votes": null
    },
    {
      "id": "323627",
      "postDate": "05/05/2018 18:10:41",
      "content": "<p>I suppose it means that Avito automatically classifies images into classes. image_top_1 represents these classes.</p>",
      "rawMarkdown": "I suppose it means that Avito automatically classifies images into classes. image_top_1 represents these classes.",
      "votes": null
    },
    {
      "id": "323630",
      "postDate": "05/05/2018 18:13:18",
      "content": "<p>Oh I see.Thanks.</p>",
      "rawMarkdown": "Oh I see.Thanks.",
      "votes": null
    },
    {
      "id": "323640",
      "postDate": "05/05/2018 18:37:35",
      "content": "<p>Would still be nice to know how it's derived and what it means exactly (e.g., is it a quality score)?</p>",
      "rawMarkdown": "Would still be nice to know how it's derived and what it means exactly (e.g., is it a quality score)?",
      "votes": null
    },
    {
      "id": "323987",
      "postDate": "05/06/2018 20:58:30",
      "content": "<p>For me, it severely decreased my 5-fold valid and public RMSE score by encoding it as a categorical variable... I think it is a numerical metric of the images' quality as opposed to a classification of the images' quality.</p>",
      "rawMarkdown": "For me, it severely decreased my 5-fold valid and public RMSE score by encoding it as a categorical variable... I think it is a numerical metric of the images' quality as opposed to a classification of the images' quality.",
      "votes": null
    },
    {
      "id": "323999",
      "postDate": "05/06/2018 22:04:16",
      "content": "<p>Encoding it as categorical increased my score in CV and leaderboard though, haha!</p>",
      "rawMarkdown": "Encoding it as categorical increased my score in CV and leaderboard though, haha!",
      "votes": null
    },
    {
      "id": "324282",
      "postDate": "05/07/2018 14:17:36",
      "content": "<p>Is This competition is good ?</p>",
      "rawMarkdown": "Is This competition is good ?",
      "votes": null
    },
    {
      "id": "324372",
      "postDate": "05/07/2018 16:27:17",
      "content": "<p>Yes, it is quite interesting thanks to a lot of different types of data.</p>",
      "rawMarkdown": "Yes, it is quite interesting thanks to a lot of different types of data.",
      "votes": null
    },
    {
      "id": "327540",
      "postDate": "05/11/2018 19:58:11",
      "content": "<p>Seems to me it is a classification of the image done by Avito. I haven't yet looked at many of them but images of good quality could range from 50 to 2000 so it doesn't seem to be a quality score. \nAlso looking at a few examples like that</p>\n\n<pre><code>train[train['image_top_1']==30].groupby(['category_name'])['item_id'].count()\ntrain[train['image_top_1']==2000].groupby(['category_name'])['item_id'].count()\n</code></pre>\n\n<p>Shows very distinct category distribution for a specific image_top_1 number.</p>\n\n<p>So far treating it as class makes the most sense to me. But it seems that there is some sort of grouping in the classes. Like image top 1, 2, 3 are all very similar distributions. So this may be why using it as numerical can actually work in a tree. You can create a heatmap reflecting that some categories get clustered in a image_top_1 range.  It's not perfect but the pattern is there.</p>",
      "rawMarkdown": "Seems to me it is a classification of the image done by Avito. I haven't yet looked at many of them but images of good quality could range from 50 to 2000 so it doesn't seem to be a quality score. \nAlso looking at a few examples like that\n\n    train[train['image_top_1']==30].groupby(['category_name'])['item_id'].count()\n    train[train['image_top_1']==2000].groupby(['category_name'])['item_id'].count()\n\nShows very distinct category distribution for a specific image_top_1 number.\n\nSo far treating it as class makes the most sense to me. But it seems that there is some sort of grouping in the classes. Like image top 1, 2, 3 are all very similar distributions. So this may be why using it as numerical can actually work in a tree. You can create a heatmap reflecting that some categories get clustered in a image_top_1 range.  It's not perfect but the pattern is there.",
      "votes": null
    },
    {
      "id": "329906",
      "postDate": "05/17/2018 14:21:42",
      "content": "<p>I write a kernel about this，<a href=\"https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook\">https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook</a></p>",
      "rawMarkdown": "I write a kernel about this，https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 323627,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "05/05/2018 18:10:41",
      "content": "<p>I suppose it means that Avito automatically classifies images into classes. image_top_1 represents these classes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 323630,
          "author_name": "sukhadj",
          "author_url": "",
          "post_date": "05/05/2018 18:13:18",
          "content": "<p>Oh I see.Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323640,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/05/2018 18:37:35",
          "content": "<p>Would still be nice to know how it's derived and what it means exactly (e.g., is it a quality score)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323987,
          "author_name": "matthewa313",
          "author_url": "",
          "post_date": "05/06/2018 20:58:30",
          "content": "<p>For me, it severely decreased my 5-fold valid and public RMSE score by encoding it as a categorical variable... I think it is a numerical metric of the images' quality as opposed to a classification of the images' quality.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323999,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/06/2018 22:04:16",
          "content": "<p>Encoding it as categorical increased my score in CV and leaderboard though, haha!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 324282,
      "author_name": "karthikbd",
      "author_url": "",
      "post_date": "05/07/2018 14:17:36",
      "content": "<p>Is This competition is good ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324372,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "05/07/2018 16:27:17",
          "content": "<p>Yes, it is quite interesting thanks to a lot of different types of data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 327540,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/11/2018 19:58:11",
      "content": "<p>Seems to me it is a classification of the image done by Avito. I haven't yet looked at many of them but images of good quality could range from 50 to 2000 so it doesn't seem to be a quality score. \nAlso looking at a few examples like that</p>\n\n<pre><code>train[train['image_top_1']==30].groupby(['category_name'])['item_id'].count()\ntrain[train['image_top_1']==2000].groupby(['category_name'])['item_id'].count()\n</code></pre>\n\n<p>Shows very distinct category distribution for a specific image_top_1 number.</p>\n\n<p>So far treating it as class makes the most sense to me. But it seems that there is some sort of grouping in the classes. Like image top 1, 2, 3 are all very similar distributions. So this may be why using it as numerical can actually work in a tree. You can create a heatmap reflecting that some categories get clustered in a image_top_1 range.  It's not perfect but the pattern is there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 329906,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "05/17/2018 14:21:42",
      "content": "<p>I write a kernel about this，<a href=\"https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook\">https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "323607": "What does ***image-top-1*** column represent? The description of data says *\"Avito's classification code for the image\"*. I really didn't get the meaning.",
    "323627": "I suppose it means that Avito automatically classifies images into classes. image_top_1 represents these classes.",
    "323630": "Oh I see.Thanks.",
    "323640": "Would still be nice to know how it's derived and what it means exactly (e.g., is it a quality score)?",
    "323987": "For me, it severely decreased my 5-fold valid and public RMSE score by encoding it as a categorical variable... I think it is a numerical metric of the images' quality as opposed to a classification of the images' quality.",
    "323999": "Encoding it as categorical increased my score in CV and leaderboard though, haha!",
    "324282": "Is This competition is good ?",
    "324372": "Yes, it is quite interesting thanks to a lot of different types of data.",
    "327540": "Seems to me it is a classification of the image done by Avito. I haven't yet looked at many of them but images of good quality could range from 50 to 2000 so it doesn't seem to be a quality score. \nAlso looking at a few examples like that\n\n    train[train['image_top_1']==30].groupby(['category_name'])['item_id'].count()\n    train[train['image_top_1']==2000].groupby(['category_name'])['item_id'].count()\n\nShows very distinct category distribution for a specific image_top_1 number.\n\nSo far treating it as class makes the most sense to me. But it seems that there is some sort of grouping in the classes. Like image top 1, 2, 3 are all very similar distributions. So this may be why using it as numerical can actually work in a tree. You can create a heatmap reflecting that some categories get clustered in a image_top_1 range.  It's not perfect but the pattern is there.",
    "329906": "I write a kernel about this，https://www.kaggle.com/liuhdsgoal/about-image-top-1-is-a-classify-label/notebook"
  },
  "source": "meta"
}