{
  "id": 307390,
  "title": "Data misleading: For finding image of article_id, let change article_id from \"number\" to \"string\" type when reading dataframe",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/307390",
  "author_name": "Nguyentuananh",
  "post_date": "2022-02-14T03:20:21.880000",
  "votes": 17,
  "comment_count": 0,
  "views": 0,
  "content": "<p>For everyone who is looking for corresponding image for article_id, you will be misleading by finding corresponding subfolders of images by article_id viewed in dataframe. So i note problem, reason and solution here for any one who need helps.</p>\n<p><strong>misleading problem</strong> </p>\n<ol>\n<li>After showing \"articles.csv\" as a dataframe, we can see that article_id as 9-digits numbers: 108775015,  108775044,  108775051, …</li>\n<li>As note from data overview, images are placed in subfolders starting with the first three digits of the <code>article_id</code>. If you use above id for finding subfolders, images will be placed in \"108\" subfolder. But we have only folders with name started from \"010\" to \"090\". We don't have any folder with named \"108\".</li>\n<li>If you look at image name, you will see that all article_id is 10-digits numbers, not 9-digits numbers as we viewed in dataframe.</li>\n</ol>\n<p><strong>The reason</strong></p>\n<ul>\n<li>In <code>articles</code> raw csv file, all <code>article_id</code> is 10-digits numbers starting with 0. So when you read dataframe, pandas automatically converted this column to number, and skipped first <code>0</code> number in every values. So from 10-digits numbers in csv file, article_id is converted into 9-digits numbers as we see after viewing in dataframe. </li>\n<li>As example, all above id viewing in dataframe: 108775015,  108775044,  108775051 should be correct as 0108775015,  0108775044,  0108775051. If you want to find subfolders for corresponding image, you need to find in '010' named folder, not in '108' named folder.</li>\n</ul>\n<p><strong>Solution</strong></p>\n<ul>\n<li>Note: Changing type of <code>article_id</code> column from int to string doesn't help. After converting, it is still 9-digits numbers.</li>\n<li>Solution 1: convert  type of <code>article_id</code> column from int to string. Then add '0' to every rows.</li>\n<li>Solution 2: Specific dtype of of <code>article_id</code> column when reading csv: <code>pd.read_csv(r\"articles.csv\", dtype={'article_id': 'str'})</code></li>\n</ul>\n<p>Hope it help. Please upvote and comment if it helps you</p>",
  "messages": [
    {
      "id": 1689099,
      "postDate": "2022-02-14T03:20:21.880Z",
      "content": "<p>For everyone who is looking for corresponding image for article_id, you will be misleading by finding corresponding subfolders of images by article_id viewed in dataframe. So i note problem, reason and solution here for any one who need helps.</p>\n<p><strong>misleading problem</strong> </p>\n<ol>\n<li>After showing \"articles.csv\" as a dataframe, we can see that article_id as 9-digits numbers: 108775015,  108775044,  108775051, …</li>\n<li>As note from data overview, images are placed in subfolders starting with the first three digits of the <code>article_id</code>. If you use above id for finding subfolders, images will be placed in \"108\" subfolder. But we have only folders with name started from \"010\" to \"090\". We don't have any folder with named \"108\".</li>\n<li>If you look at image name, you will see that all article_id is 10-digits numbers, not 9-digits numbers as we viewed in dataframe.</li>\n</ol>\n<p><strong>The reason</strong></p>\n<ul>\n<li>In <code>articles</code> raw csv file, all <code>article_id</code> is 10-digits numbers starting with 0. So when you read dataframe, pandas automatically converted this column to number, and skipped first <code>0</code> number in every values. So from 10-digits numbers in csv file, article_id is converted into 9-digits numbers as we see after viewing in dataframe. </li>\n<li>As example, all above id viewing in dataframe: 108775015,  108775044,  108775051 should be correct as 0108775015,  0108775044,  0108775051. If you want to find subfolders for corresponding image, you need to find in '010' named folder, not in '108' named folder.</li>\n</ul>\n<p><strong>Solution</strong></p>\n<ul>\n<li>Note: Changing type of <code>article_id</code> column from int to string doesn't help. After converting, it is still 9-digits numbers.</li>\n<li>Solution 1: convert  type of <code>article_id</code> column from int to string. Then add '0' to every rows.</li>\n<li>Solution 2: Specific dtype of of <code>article_id</code> column when reading csv: <code>pd.read_csv(r\"articles.csv\", dtype={'article_id': 'str'})</code></li>\n</ul>\n<p>Hope it help. Please upvote and comment if it helps you</p>",
      "rawMarkdown": "For everyone who is looking for corresponding image for article_id, you will be misleading by finding corresponding subfolders of images by article_id viewed in dataframe. So i note problem, reason and solution here for any one who need helps.\n\n**misleading problem** \n1. After showing \"articles.csv\" as a dataframe, we can see that article_id as 9-digits numbers: 108775015,  108775044,  108775051, ...\n2. As note from data overview, images are placed in subfolders starting with the first three digits of the `article_id`. If you use above id for finding subfolders, images will be placed in \"108\" subfolder. But we have only folders with name started from \"010\" to \"090\". We don't have any folder with named \"108\".\n3. If you look at image name, you will see that all article_id is 10-digits numbers, not 9-digits numbers as we viewed in dataframe.\n\n**The reason**\n* In `articles` raw csv file, all `article_id` is 10-digits numbers starting with 0. So when you read dataframe, pandas automatically converted this column to number, and skipped first `0` number in every values. So from 10-digits numbers in csv file, article_id is converted into 9-digits numbers as we see after viewing in dataframe. \n* As example, all above id viewing in dataframe: 108775015,  108775044,  108775051 should be correct as 0108775015,  0108775044,  0108775051. If you want to find subfolders for corresponding image, you need to find in '010' named folder, not in '108' named folder.\n\n**Solution**\n* Note: Changing type of `article_id` column from int to string doesn't help. After converting, it is still 9-digits numbers.\n* Solution 1: convert  type of `article_id` column from int to string. Then add '0' to every rows.\n* Solution 2: Specific dtype of of `article_id` column when reading csv: `pd.read_csv(r\"articles.csv\", dtype={'article_id': 'str'})`\n\nHope it help. Please upvote and comment if it helps you",
      "votes": 17
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1689099": "For everyone who is looking for corresponding image for article_id, you will be misleading by finding corresponding subfolders of images by article_id viewed in dataframe. So i note problem, reason and solution here for any one who need helps.\n\n**misleading problem** \n1. After showing \"articles.csv\" as a dataframe, we can see that article_id as 9-digits numbers: 108775015,  108775044,  108775051, ...\n2. As note from data overview, images are placed in subfolders starting with the first three digits of the `article_id`. If you use above id for finding subfolders, images will be placed in \"108\" subfolder. But we have only folders with name started from \"010\" to \"090\". We don't have any folder with named \"108\".\n3. If you look at image name, you will see that all article_id is 10-digits numbers, not 9-digits numbers as we viewed in dataframe.\n\n**The reason**\n* In `articles` raw csv file, all `article_id` is 10-digits numbers starting with 0. So when you read dataframe, pandas automatically converted this column to number, and skipped first `0` number in every values. So from 10-digits numbers in csv file, article_id is converted into 9-digits numbers as we see after viewing in dataframe. \n* As example, all above id viewing in dataframe: 108775015,  108775044,  108775051 should be correct as 0108775015,  0108775044,  0108775051. If you want to find subfolders for corresponding image, you need to find in '010' named folder, not in '108' named folder.\n\n**Solution**\n* Note: Changing type of `article_id` column from int to string doesn't help. After converting, it is still 9-digits numbers.\n* Solution 1: convert  type of `article_id` column from int to string. Then add '0' to every rows.\n* Solution 2: Specific dtype of of `article_id` column when reading csv: `pd.read_csv(r\"articles.csv\", dtype={'article_id': 'str'})`\n\nHope it help. Please upvote and comment if it helps you"
  }
}