{
  "id": 312064,
  "title": "be careful about the type of 'article_id'",
  "url": "/competitions/h-and-m-personalized-fashion-recommendations/discussion/312064",
  "author_name": "",
  "post_date": "2022-03-10T07:06:34.867684400Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>use <code>pd.read_csv('.xxx.csv', dtype = {'article_id': str})</code> when you process dataframe contained article_id column.</p>",
  "messages": [
    {
      "id": "1717734",
      "postDate": "03/10/2022 07:06:34",
      "content": "<p>use <code>pd.read_csv('.xxx.csv', dtype = {'article_id': str})</code> when you process dataframe contained article_id column.</p>",
      "rawMarkdown": "use `pd.read_csv('.xxx.csv', dtype = {'article_id': str})` when you process dataframe contained article_id column.",
      "votes": null
    },
    {
      "id": "1718061",
      "postDate": "03/10/2022 13:22:08",
      "content": "<p>It is true that when we make a submission to Kaggle we need to format each <code>article_id</code> as a 10 digit string which includes the first digit zero. However when training out model (or processing our heuristic), storing the <code>article_id</code> as int32 uses 4 bytes instead of string which uses 10 bytes. So it's best to store as int32 while working with the data (i.e. dataframe groupbys, model training, etc), and then convert back to string before submission.</p>",
      "rawMarkdown": "It is true that when we make a submission to Kaggle we need to format each `article_id` as a 10 digit string which includes the first digit zero. However when training out model (or processing our heuristic), storing the `article_id` as int32 uses 4 bytes instead of string which uses 10 bytes. So it's best to store as int32 while working with the data (i.e. dataframe groupbys, model training, etc), and then convert back to string before submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1718061,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/10/2022 13:22:08",
      "content": "<p>It is true that when we make a submission to Kaggle we need to format each <code>article_id</code> as a 10 digit string which includes the first digit zero. However when training out model (or processing our heuristic), storing the <code>article_id</code> as int32 uses 4 bytes instead of string which uses 10 bytes. So it's best to store as int32 while working with the data (i.e. dataframe groupbys, model training, etc), and then convert back to string before submission.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1717734": "use `pd.read_csv('.xxx.csv', dtype = {'article_id': str})` when you process dataframe contained article_id column.",
    "1718061": "It is true that when we make a submission to Kaggle we need to format each `article_id` as a 10 digit string which includes the first digit zero. However when training out model (or processing our heuristic), storing the `article_id` as int32 uses 4 bytes instead of string which uses 10 bytes. So it's best to store as int32 while working with the data (i.e. dataframe groupbys, model training, etc), and then convert back to string before submission."
  },
  "source": "meta"
}