{
  "id": 206141,
  "title": "Why using merge with Questions_df fails",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206141",
  "author_name": "Jaideep",
  "post_date": "2020-12-23T11:56:02.270000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all,<br>\ni am having very tough time to understand some weird things..<br>\nbelow from my Test inference loop. It fails with in 2 minutes run of submission</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n\n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False)\n</code></pre>\n<p>This below thing works</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n\n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False,how='left)\n       test_df[test_df.part.isnull(),part]=1 \n</code></pre>\n<p>The below also works</p>\n<p><code>test_df['part]=test_df.map(dict)</code></p>",
  "messages": [
    {
      "id": 1123653,
      "postDate": "2020-12-23T11:56:02.270Z",
      "content": "<p>Hi all,<br>\ni am having very tough time to understand some weird things..<br>\nbelow from my Test inference loop. It fails with in 2 minutes run of submission</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n\n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False)\n</code></pre>\n<p>This below thing works</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n\n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False,how='left)\n       test_df[test_df.part.isnull(),part]=1 \n</code></pre>\n<p>The below also works</p>\n<p><code>test_df['part]=test_df.map(dict)</code></p>",
      "rawMarkdown": "Hi all,\ni am having very tough time to understand some weird things..\nbelow from my Test inference loop. It fails with in 2 minutes run of submission\n\n ```\nfor (test_df, sample_prediction_df) in iter_test:\n       \n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False)\n```\nThis below thing works\n ```\nfor (test_df, sample_prediction_df) in iter_test:\n       \n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False,how='left)\n       test_df[test_df.part.isnull(),part]=1 \n```\n\nThe below also works\n\n`test_df['part]=test_df.map(dict) `",
      "votes": 1
    },
    {
      "id": 1124461,
      "postDate": "2020-12-24T01:02:37.387Z",
      "content": "<p>By default merge is having a inner join, if you don't change that, you will miss on rows basically; And that's what breaks your code as row counts doesn't match.</p>",
      "rawMarkdown": "By default merge is having a inner join, if you don't change that, you will miss on rows basically; And that's what breaks your code as row counts doesn't match.",
      "replies": [
        {
          "id": 1124524,
          "postDate": "2020-12-24T02:47:33.333Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> see my response below </p>",
          "rawMarkdown": "@adityaecdrid see my response below "
        }
      ]
    },
    {
      "id": 1124428,
      "postDate": "2020-12-23T23:36:32.703Z",
      "content": "<p><code>how=left</code> is required to preserve the original row order.</p>",
      "rawMarkdown": "`how=left` is required to preserve the original row order.",
      "replies": [
        {
          "id": 1124523,
          "postDate": "2020-12-24T02:46:54.297Z",
          "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>  inner join  will lead to only content id = False rows as we joining with questions,  ,which  any ways we have to do, before submitting preds. So that is what making me scratch head </p>",
          "rawMarkdown": "@nyanpn  inner join  will lead to only content id = False rows as we joining with questions,  ,which  any ways we have to do, before submitting preds. So that is what making me scratch head \n"
        },
        {
          "id": 1124546,
          "postDate": "2020-12-24T03:32:54.280Z",
          "content": "<blockquote>\n  <p>inner join will lead to only content id = False rows as we joining with questions</p>\n</blockquote>\n<p>But the answers have values for lecture rows as well; you will have a shape mismatch when you will try updating your cache; Look into common failures tips!</p>",
          "rawMarkdown": ">inner join will lead to only content id = False rows as we joining with questions\n\nBut the answers have values for lecture rows as well; you will have a shape mismatch when you will try updating your cache; Look into common failures tips!"
        },
        {
          "id": 1124559,
          "postDate": "2020-12-24T03:53:25.383Z",
          "content": "<p>Hmm ,I get a bit i did thought about it but  here comes a conflict <br>\nIf I do<br>\n<code>test [col]= test.content I'd. Map (dict)</code> this works as it preserves rows goes with explanation. </p>\n<p><code>test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0])</code> <br>\nFails </p>",
          "rawMarkdown": "Hmm ,I get a bit i did thought about it but  here comes a conflict \nIf I do\n` test [col]= test.content I'd. Map (dict) ` this works as it preserves rows goes with explanation. \n\n` test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0]) ` \nFails \n\n\n"
        },
        {
          "id": 1124560,
          "postDate": "2020-12-24T03:53:25.383Z",
          "content": "<p>Hmm ,I get a bit i did thought about it but  here comes a conflict <br>\nIf I do<br>\n<code>test [col]= test.content I'd. Map (dict)</code> this works as it preserves rows goes with explanation. </p>\n<p><code>test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0])</code> <br>\nFails </p>\n<p>Moreover I do prev df with content id =False  before updating answers . If there is size mismatch issue then I should get the same during commit itself </p>",
          "rawMarkdown": "Hmm ,I get a bit i did thought about it but  here comes a conflict \nIf I do\n` test [col]= test.content I'd. Map (dict) ` this works as it preserves rows goes with explanation. \n\n` test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0]) ` \nFails \n\nMoreover I do prev df with content id =False  before updating answers . If there is size mismatch issue then I should get the same during commit itself \n"
        }
      ]
    },
    {
      "id": 1124342,
      "postDate": "2020-12-23T20:53:32.987Z",
      "content": "<p>I think we need to see the whole loop to understand.</p>",
      "rawMarkdown": "I think we need to see the whole loop to understand.",
      "replies": [
        {
          "id": 1124530,
          "postDate": "2020-12-24T02:57:31.013Z",
          "content": "<p><a href=\"https://www.kaggle.com/ahmeterdem\" target=\"_blank\">@ahmeterdem</a></p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\ntest_df=pd.merge(test_df,questions_df[['question_id','part']],  left_on='content_id',right_on='question_id',copy=False)\nIf Prev df is not none:\n......\nPrev df = copy test df \nTest df = testdf content type =0\n</code></pre>",
          "rawMarkdown": "\n\n@ahmeterdem\n```\nfor (test_df, sample_prediction_df) in iter_test:\ntest_df=pd.merge(test_df,questions_df[['question_id','part']],  left_on='content_id',right_on='question_id',copy=False)\n\nIf Prev df is not none:\n......\n\nPrev df = copy test df \n\nTest df = testdf content type =0\n```"
        },
        {
          "id": 1124873,
          "postDate": "2020-12-24T08:44:46.897Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1124892,
          "postDate": "2020-12-24T09:03:02.430Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1124461,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-12-24T01:02:37.387000",
      "content": "<p>By default merge is having a inner join, if you don't change that, you will miss on rows basically; And that's what breaks your code as row counts doesn't match.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1124524,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-24T02:47:33.333000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> see my response below </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1124428,
      "author_name": "nyanp",
      "author_url": "",
      "post_date": "2020-12-23T23:36:32.703000",
      "content": "<p><code>how=left</code> is required to preserve the original row order.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1124523,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-24T02:46:54.297000",
          "content": "<p><a href=\"https://www.kaggle.com/nyanpn\" target=\"_blank\">@nyanpn</a>  inner join  will lead to only content id = False rows as we joining with questions,  ,which  any ways we have to do, before submitting preds. So that is what making me scratch head </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124546,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-12-24T03:32:54.280000",
          "content": "<blockquote>\n  <p>inner join will lead to only content id = False rows as we joining with questions</p>\n</blockquote>\n<p>But the answers have values for lecture rows as well; you will have a shape mismatch when you will try updating your cache; Look into common failures tips!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124559,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-24T03:53:25.383000",
          "content": "<p>Hmm ,I get a bit i did thought about it but  here comes a conflict <br>\nIf I do<br>\n<code>test [col]= test.content I'd. Map (dict)</code> this works as it preserves rows goes with explanation. </p>\n<p><code>test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0])</code> <br>\nFails </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124560,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-24T03:53:25.383000",
          "content": "<p>Hmm ,I get a bit i did thought about it but  here comes a conflict <br>\nIf I do<br>\n<code>test [col]= test.content I'd. Map (dict)</code> this works as it preserves rows goes with explanation. </p>\n<p><code>test [col]= test.content I'd. Map (lambda x : df.loc[x,col][0])</code> <br>\nFails </p>\n<p>Moreover I do prev df with content id =False  before updating answers . If there is size mismatch issue then I should get the same during commit itself </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1124342,
      "author_name": "Ahmet Erdem",
      "author_url": "",
      "post_date": "2020-12-23T20:53:32.987000",
      "content": "<p>I think we need to see the whole loop to understand.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1124530,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-12-24T02:57:31.013000",
          "content": "<p><a href=\"https://www.kaggle.com/ahmeterdem\" target=\"_blank\">@ahmeterdem</a></p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\ntest_df=pd.merge(test_df,questions_df[['question_id','part']],  left_on='content_id',right_on='question_id',copy=False)\nIf Prev df is not none:\n......\nPrev df = copy test df \nTest df = testdf content type =0\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124873,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-24T08:44:46.897000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1124892,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-12-24T09:03:02.430000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1123653": "Hi all,\ni am having very tough time to understand some weird things..\nbelow from my Test inference loop. It fails with in 2 minutes run of submission\n\n ```\nfor (test_df, sample_prediction_df) in iter_test:\n       \n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False)\n```\nThis below thing works\n ```\nfor (test_df, sample_prediction_df) in iter_test:\n       \n        test_df=pd.merge(test_df,questions_df[['question_id','part']],\n                         left_on='content_id',right_on='question_id',copy=False,how='left)\n       test_df[test_df.part.isnull(),part]=1 \n```\n\nThe below also works\n\n`test_df['part]=test_df.map(dict) `",
    "1124461": "By default merge is having a inner join, if you don't change that, you will miss on rows basically; And that's what breaks your code as row counts doesn't match.",
    "1124428": "`how=left` is required to preserve the original row order.",
    "1124342": "I think we need to see the whole loop to understand."
  }
}