Skip to content
JournalsWorldThe Global Research Discovery Platform
Featured Dataset

Lending Club loan dataset for granting models

Lending Club offers peer-to-peer (P2P) loans through a technological platform for various personal finance purposes and is today one of the companies that dominate the US P2P lending market. The original dataset is publicly available on <a href="https://www.kaggle.com/datasets/wordsforthewise/len

👤
CreatorAriza-Garzón, Miller Janny
📅
Published2024-05-25
🔗
DOI10.5281/zenodo.11295916
📊
Downloads3,379
⚖️
Licensecc-by-4.0
File Size159.7 MB
Data TypeDataset
Published2024
Licensecc-by-4.0
Total Views8,487
Total Downloads3,379

Lending Club offers peer-to-peer (P2P) loans through a technological platform for various personal finance purposes and is today one of the companies that dominate the US P2P lending market. The original dataset is publicly available on Kaggle and corresponds to all the loans issued by Lending Club between 2007 and 2018. The present version of the dataset is for constructing a granting model, that is, a model designed to make decisions on whether to grant a loan based on information available at the time of the loan application. Consequently, our dataset only has a selection of variables from the original one, which are the variables known at the moment the loan request is made. Furthermore, the target variable of a granting model represents the final status of the loan, that are “default” or “fully paid”. Thus, we filtered out from the original dataset all the loans in transitory states. Our dataset comprises 1,347,681 records or obligations (approximately 60% of the original) and it was also cleaned for completeness and consistency (less than 1% of our dataset was filtered out).

TARGET VARIABLE

The dataset includes a target variable based on the final resolution of the credit: the default category corresponds to the event charged off and the non-default category to the event fully paid. It does not consider other values in the loan status variable since this variable represents the state of the loan at the end of the considered time window. Thus, there are no loans in transitory states. The original dataset includes the target variable “loan status”, which contains several categories (‘Fully Paid’, ‘Current’, ‘Charged Off’, ‘In Grace Period’, ‘Late (31-120 days)’, ‘Late (16-30 days)’, ‘Default’). However, in our dataset, we just consider loans that are either “Fully Paid” or “Default” and transform this variable into a binary variable called “Default”, with a 0 for fully paid loans and a 1 for defaulted loans.

EXPLANATORY VARIABLES

The explanatory variables that we use correspond only to the information available at the time of the application. Variables such as the interest rate, grade, or subgrade are generated by the company as a result of a credit risk assessment process, so they were filtered out from the dataset as they must not be considered in risk models to predict the default in granting of credit.

FULL LIST OF VARIABLES

Loan identification variables:

  • id: Loan id (unique identifier). 

  • issue_d: Month and year in which the loan was approved.

Quantitative variables:

  • revenue: Borrower’s self-declared annual income during registration. 

  • dti_n: Indebtedness ratio for obligations excluding mort

    📤 Share this page

    Found this useful? Share it with your network.

    ✓ Link copied! Paste it on ResearchGate / Academia.edu
📦
Lending Club loan dataset for granting models (Full Dataset)159.7 MB
⬇
📄
ReadmeVia DOI record
↗

Files are hosted on the source repository. Click download to access the full dataset.

Ariza-Garzón, Miller Janny (2024). Lending Club loan dataset for granting models. https://doi.org/10.5281/zenodo.11295916